A SLAM method for removing dynamic objects based on an RGBD sensor
Through the YOLOv5 network and pixel-level segmentation method based on RGBD sensor, the instability problem of traditional SLAM systems under dynamic target interference is solved, real-time and accurate dynamic target removal on the mobile robot side is achieved, and the stability and positioning accuracy of the SLAM system are improved.
Patent Information
- Application Number
- CN202111637308.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-12-29
AI Technical Summary
Traditional SLAM systems are unstable under dynamic target interference, have low accuracy, and are difficult to effectively apply in real environments.
Using an RGBD sensor-based method, the YOLOv5 target detection network is used to detect dynamic targets, combined with the pixel-level segmentation method, the dynamic target and static background are separated, and only ordinary image processors are used to remove dynamic target interference in real time.
It improves the accuracy and robustness of the SLAM system in a dynamic environment, can run in real time on the mobile robot side, reduces the amount of calculation, and has a good segmentation effect.
Smart Images

Figure CN114283198B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and mobile robot positioning, and particularly relates to a SLAM method for removing dynamic targets based on an RGBD sensor. Background Art
[0002] In recent years, computer vision and robots have become a hot research direction. Among them, the most fundamental research is the positioning problem of the robot itself. The Simultaneous Localization and Mapping (SALM) technology is widely used in robot positioning, navigation, and obstacle avoidance. There is a rich variety of robot sensors. Visual sensors have the advantages of low cost and large amount of information, and are widely used in SLAM technology. Among them, RGBD sensors can directly provide depth information, which can effectively reduce the amount of calculation. Traditional Simultaneous Localization and Mapping technology focuses on the research of ideal static scenarios without moving targets. In actual scenarios, dynamic targets can cause deviations in the calculation of the camera pose, resulting in inaccurate positioning of the entire visual SLAM system. The pure static assumption is not applicable in the real environment, and SLAM applicable to dynamic environments is a research hotspot of this technology.
[0003] With the development of deep learning technology, deep learning methods are widely used in object detection and semantic segmentation. However, most deep learning networks require a good Graphics Processing Unit (GPU) for acceleration support to achieve a real-time detection effect. Currently, many works are carried out based on semantic segmentation networks. Semantic segmentation requires a larger amount of calculation compared to object detection and is difficult to meet the real-time requirement. YOLOv5 is a new object detection network with high accuracy and good real-time performance. After detecting the target, pixel-level segmentation is also required to achieve the effect of removing dynamic targets. Summary of the Invention
[0004] In order to overcome the above-mentioned deficiencies of the prior art, the present invention provides a SLAM method for removing dynamic targets based on an RGBD sensor. The purpose of the present invention is to implement a SLAM system that can run in real time on a mobile device, eliminate the interference of pre-set dynamic target objects in the scene on the positioning and mapping of the mobile robot, and ensure the stability of the SLAM system.
[0005] The purpose of the present invention is achieved through the following technical solutions: A SLAM method for removing dynamic targets based on an RGBD sensor, the method comprising the following steps:
[0006] Step 1, obtain pictures of pre-set dynamic target categories to be removed, and construct a training sample set and a test data set;
[0007] Step 2: Use the training sample set and test data set constructed in Step 1 to train and test the object detection neural network YOLOv5, and obtain a trained YOLOv5 model;
[0008] Step 3: Use the object detection neural network YOLOv5 model trained in Step 2 to process the RGB images collected by the RGBD sensor, and judge whether there are dynamic targets in the images and the pixel coordinate positions where the dynamic target boxes are located;
[0009] Step 4: Combine the pixel coordinate positions of the dynamic target boxes obtained in Step 3 and the depth image information of the RGBD sensor to judge whether there are dynamic targets in the current frame of the depth image of the RGBD sensor. If there are dynamic targets, use the pixel-level segmentation method to process the same pixel area drawn within the dynamic target box in the depth image corresponding to the current frame to obtain a mask image. If there are no dynamic targets, directly set the mask image to be empty;
[0010] The pixel-level segmentation method specifically includes the following steps:
[0011] Step 4.1: For all dynamic target detection boxes in each frame of RGB image, within the dynamic target detection box area corresponding to the pixel positions in the depth map of the corresponding frame image, uniformly distribute and take the depth values of N pixel points and store them in the container Xi, where the value of i ranges from 1 to N;
[0012] Step 4.2: Use the absolute median deviation outlier algorithm for the obtained depth values of N pixel points to divide them into two sets composed of depth values, and then use weighted summation for the two sets to obtain the average pixel depth values of the two sets respectively. The set with the relatively smaller average pixel depth value is set as the dynamic target class depth value, and the set with the relatively larger average pixel depth value is set as the static background class depth value; within the dynamic target box area output by the object detection neural network YOLOv5 model, the dynamic target is closer to the camera optical center than the static background.
[0013] Step 4.3: For the dynamic target class depth value and the static background class depth value, use a distance-based clustering algorithm to cluster all pixels in the dynamic target detection box, and divide all pixels in the dynamic target detection box into the dynamic target class or the static background class. The obtained dynamic target class or static background class is a set of pixel points;
[0014] Step 4.4: Extract the pixel positions of the dynamic target classes in all the target detection boxes in the image, project the obtained pixel positions of the dynamic target classes onto an image of the same size as the original input image and output it in a binary form. Perform an OR operation on the pixel positions corresponding to the binary images generated by all the target detection boxes in the image, so that multiple images are superimposed and fused into a single mask image, and the mask image is also output in a binary form;
[0015] Step 5: Process the grayscale image in SLAM using the mask image. Specifically, perform an AND operation on the mask image output in binary form and the grayscale image in SLAM. Among them, the pixel values at the positions where the dynamic target classes are located will be 0, while the other parts of the grayscale image will remain unchanged. The part where the pixel values in the grayscale image are set to 0 is the dynamic target information in the image, and the unchanged part of the grayscale image is the static background information. Use the static background information for the positioning and map construction of the mobile robot.
[0016] Furthermore, in the above Step 1, after the test dataset labels are made, they are adjusted into the VOC dataset format for subsequent steps.
[0017] Furthermore, in the above Step 2, the simplest YOLOv5s network in the object detection neural network YOLOv5 is used for training. This object detection neural network has fewer parameters and faster running speed compared with other networks of the same type.
[0018] Furthermore, in the above Step 3, the pixel coordinate positions include the center point coordinates and the width and height of the target box, and both are normalized according to the size of the input image.
[0019] Furthermore, in the above Step 4.2, the absolute median deviation outlier algorithm can detect one or several values in the data that are significantly different from other values and eliminate them. The degree of difference is judged according to the parameter n in the absolute median deviation outlier algorithm. Use the algorithm to extract the average depth information of the dynamic target, and its formula is:
[0020]
[0021] The basic steps of this algorithm can be summarized as the following sub-steps:
[0022] Step 4.2.1: Calculate the median value \(X_{med}\) of all elements in \(X\); i in \(X\); median ;
[0023] Step 4.2.2: Calculate the absolute deviation of all elements from the median value, and then calculate the median value MAD of the relative absolute deviation of all elements;
[0024] Step 4.2.3: Determine the parameter n, and the parameter can be adjusted according to the algorithm formula; isolate the outlier data and store it in the container outlier, and store other data in the container normal. Then calculate the average depth value of all pixels in the container outlier and the average depth value of all pixels in the container normal, and compare the average depth values of the two. The one with the relatively smaller average pixel depth value is set as the dynamic target class depth value, and the one with the relatively larger average pixel depth value is set as the static background class depth value.
[0025] Further, in step 4.3, the distance-based clustering algorithm includes the following steps:
[0026] Step 4.3.1: First, set the clustering center points. In this step, set the depth values of two center points, which are the dynamic target class depth value and the static background class depth value obtained in step 4.2 respectively;
[0027] Step 4.3.2: Traverse all pixel points in the dynamic target detection frame, calculate the Euclidean distances from the depth value of each pixel point to the depth values of the two set clustering center points respectively, compare the magnitudes of the two distance values, and divide all pixel points in the dynamic target detection frame into the clustering center with the closer distance;
[0028] Step 4.3.3: Re-select the depth value of the clustering center, calculate the average pixel of the two sets of pixel points divided in step 4.3.2, and use it as the depth value of the clustering center point set in step 4.3.1, and re-perform the next round of iterative clustering;
[0029] Step 4.3.4: Loop steps 4.3.1 to 4.3.3 until at least one of the following conditions is met, then end the loop and output the pixel position information of the dynamic target class or the static background class in the dynamic target detection frame;
[0030] Condition 1: Until the difference between the average pixels of the two classes and the clustering center is less than the set value a;
[0031] Condition 2: The difference in the number of pixels belonging to the two classes in the target frame is greater than the set value b;
[0032] The parameters a and b are adjusted according to the actual application scenario.
[0033] Further, in step 5, the algorithm flow of the SLAM system is specifically as follows:
[0034] Step 5.1: Perform ORB feature extraction on the static image provided in step 4. When the number of extracted feature points exceeds the set value, the SLAM system is initialized;
[0035] Step 5.2: Estimate the motion pose of the camera by using the extracted ORB features in combination with the static background information of the previous frame. The RGBD camera provides depth information, and the PnP method is used to solve the camera pose.
[0036] Step 5.3: Minimize the reprojection error using Bundle Adjustment (BA) to optimize the local map.
[0037] Step 5.4: Use loop detection to optimize the pose and correct the drift error.
[0038] Further, Step 5.4 is divided into two parts, namely loop detection and loop correction. The loop detection first uses the Bag of Words (Bow) model for detection, and then calculates the similarity transformation through the Sim3 algorithm. The loop correction mainly performs graph optimization of the Closed-loop Fusion Essential Graph to achieve the effect of adjusting and correcting the error.
[0039] Advantages of the present invention: Aiming at the situation that traditional SLAM cannot overcome the interference of dynamic targets, the SLAM method based on RGBD sensors for removing dynamic targets proposed by the present invention effectively overcomes the disadvantages of traditional SLAM methods being unstable and having low accuracy under the interference of dynamic targets, and effectively improves the accuracy under the interference of dynamic targets. At the same time, a target detection framework of deep learning is introduced, making the method universal, and only requiring an ordinary image processor to support to achieve real-time effects. In terms of the pixel-level segmentation method, the method proposed by the present invention has a small amount of calculation and a good segmentation effect. Different from the mainstream method of using deep learning to process dynamic targets, the present invention has the advantages of small calculation amount, good real-time performance, and good effect of removing dynamic targets, which is beneficial to the deployment and application on mobile robots. It can effectively improve the accuracy and robustness of tracking using RGBD sensors, and improve the positioning and mapping accuracy of visual SLAM in dynamic scenarios. Brief Description of the Drawings
[0040] Figure 1 It is a structural schematic diagram of the SLAM method based on RGBD sensors for removing dynamic targets of the present invention;
[0041] Figure 2 It is a flowchart of the pixel-level segmentation method of the present invention;
[0042] Figure 3 It is a processing result diagram of YOLOv5 in a certain scenario in the embodiment of the present invention;
[0043] Figure 4 It is an effect diagram of uniformly sampling points in a certain scenario in the embodiment of the present invention;
[0044] Figure 5 It is a pixel-level segmentation result diagram in a certain scenario in the embodiment of the present invention;
[0045] Figure 6 Schematic diagram of the mask image in a certain scenario of the embodiment of the present invention;
[0046] Figure 7 Error comparison diagram of the motion trajectory in the xyz direction in the embodiment of the present invention;
[0047] Figure 8 Error comparison diagram of the motion trajectory in the rpy angle in the embodiment of the present invention. Specific implementation manner
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0049] Due to the static assumption of traditional simultaneous localization and mapping technology, it is difficult for visual SLAM technology to be widely applied in real scenarios. After removing the interference of dynamic targets, it can be more effectively deployed in real environments.
[0050] The object of the present invention is achieved through the following technical solutions: A SLAM method for removing dynamic targets based on an RGBD sensor, which provides RGB image information and depth image information by the RGBD sensor of a mobile robot, and uses an object detection network to perform dynamic object detection on the RGB image to obtain the position of the dynamic object in pixel coordinates. Using the depth image and the dynamic object coordinate position, pixel-level segmentation is performed on each dynamic object area to completely separate the dynamic object and the static background. The static background is used for the positioning and mapping of the mobile robot. This method removes the interference of dynamic targets with a small amount of calculation and can be applied to mobile robots in real time, which is beneficial to the stable operation of the SLAM system. As Figure 1 shown, this method includes the following steps:
[0051] Step 1: According to the actual application scenario, obtain the pictures of the dynamic targets that need to be removed preset, and construct a training sample set and a test data set. In this embodiment, the preset dynamic target is set as a person, and the pixel positions of the dynamic targets of the set class in the corresponding images are marked in its training sample set and test data set. A total of 17,125 relevant pictures containing the set targets are collected for the production of the data set. The Labelimg tool is used for the marking production of the data set, and after the marking production of the test data set, it is adjusted into the VOC data set format for subsequent steps.
[0052] Step 2: Use the constructed training sample set and test data set to train and test the YOLOv5 model of the object detection neural network, and obtain the trained YOLOv5 model. In the present invention, the simplest YOLOv5s model in YOLOv5 is used for training. The advantage of this network is that it is small in size and fast in speed. Among them, the ratio of the training sample set to the test sample set is 10:1. In the present invention, the network of YOLOv5 is not improved, and the open-source network framework is used. The results of its training are tested, and its detection accuracy can reach 85.1%. There is still room for improvement in its indicators in this network framework, and it can be specifically adapted to the specified application environment by adjusting the data set and network framework parameters. The ability to accurately detect dynamic targets is of great significance for the subsequent steps of the present invention.
[0053] Step 3: Use the trained YOLOv5 model to process the RGB image collected by the RGBD sensor, and judge whether there is a dynamic target in the image and the pixel coordinate position where the dynamic target box is located. The dynamic target box information output by its module can be expressed by the following formula:
[0054]
[0055] Among them, o j refers to the pre-set j types of dynamic targets in this frame of image, is the x coordinate of the center point of the i-th j-type target box in this frame of image, refers to the y coordinate of the center point of the i-th j-type target box in this frame of image, refers to the width and height of the coordinate box of the i-th j-type with the previous x and y as the center points in this frame of image. In the formula The 4 parameters, namely the center point coordinates and the width and height, have been normalized according to the size of the input image and can be applied to images of different sizes after scaling. The effect of the YOLO module is as Figure 3 shown.
[0056] Step 4: Combine the pixel coordinate position of the dynamic target box obtained in Step 3 and the depth image of the RGBD sensor to judge whether there is a dynamic target in the current frame. If there is a dynamic target, use the pixel-level segmentation method to process the same pixel area drawn within the dynamic target box in the depth image corresponding to the current frame to obtain a mask image. If there is no dynamic target, the mask image is directly set to be empty.
[0057] Further, as Figure 2 shown, in the said Step 4, the pixel-level segmentation method specifically includes the following sub-steps:
[0058] Step 4.1: For all the dynamic object detection boxes in each frame of RGB image, within the area of the dynamic object detection box corresponding to the pixel positions in the depth map of this frame of image, uniformly distribute and take the depth values of N pixel points, and store them in the container X i where the value of i ranges from 1 to N. When N is taken as 9, the effect is as shown in Figure 4 the figure below
[0059] Step 4.2: Use the absolute median deviation outlier algorithm for the obtained depth values of N pixel points to divide them into two sets composed of depth values, and then use weighted summation for the two sets respectively to obtain the average pixel depth values of the two sets, and compare the two average pixel depth values. Among them, the one with the relatively smaller average pixel depth value is set as the dynamic object class depth value, and the one with the relatively larger average pixel depth value is set as the static background class depth value. When the camera acquires images in an indoor environment, the distance between the position of the dynamic object and the background from the camera optical center is clearly defined in the depth image, where the depth value of the background is much larger than that of the object class. Within the dynamic object box area output by the object detection neural network in Step 3, the position of the dynamic object is closer to the camera optical center than the static background
[0060] The absolute median deviation outlier algorithm can detect one or several data values in the data that are significantly different from other values and eliminate them. The definition of the large difference is obtained from the experiments in the specific application scenarios of the present invention. The parameter n in the absolute median deviation outlier algorithm is used to define the size of the difference. In the experiments of the present invention, the parameter n is set to 1.1. As shown in Figure 3 the figure below, since most of the pixels in the object detection box belong to the dynamic object, their depth information is very close, while the depth information of the background is quite different from that of the detection object. This algorithm can be used to extract the average depth information of the dynamic object. The formula is as follows
[0061]
[0062] The basic steps of this algorithm can be summarized as the following sub-steps
[0063] Step 4.2.1: Calculate the median value X i of all elements in X median .
[0064] Step 4.2.2: Calculate the absolute deviation of all elements from the median value, and then calculate the median value MAD of all elements relative to the absolute deviation
[0065] Step 4.2.3: Determine the parameter n, and the data can be adjusted according to the MAD formula
[0066] In the present invention, the parameter n is adjusted to 1.1 according to the depth value situation adapting to the target scene. Meanwhile, only the outlier data is separated and stored in the container "outlier", and other data is stored in the container "normal", without adjusting the outlier data. Then, the average depth value of all pixels in the container "outlier" and the average depth value of all pixels in the container "normal" are calculated, and the average depth values of the two are compared. Among them, the one with a relatively smaller average pixel depth value is set as the dynamic target class depth value, and the one with a relatively larger average pixel depth value is set as the static background class depth value.
[0067] Step 4.3: For the dynamic target class depth value and the static background class depth value, use a distance-based clustering algorithm to cluster all pixels in the dynamic target detection box, and divide all pixels in the dynamic target box into the dynamic target class or the static background class. Here, both the dynamic target class and the static background class are sets of pixel points.
[0068] The distance-based clustering algorithm specifically includes the following sub-steps:
[0069] Step 4.3.1: First, set the clustering center points. In this step, the depth values of two center points are set, which are respectively the dynamic target class depth value and the static background class depth value obtained from Step 4.2.
[0070] Step 4.3.2: Traverse all pixel points in the dynamic target detection box, calculate the Euclidean distances from the depth value of each pixel point to the depth values of the two set clustering center points respectively, compare the magnitudes of the two distance values, and divide all pixel points in the dynamic target detection box into the clustering center with a closer distance.
[0071] Step 4.3.3: Re-select the depth value of the clustering center, calculate the average pixels of the two sets of pixel points divided in Step 4.3.2, and use it as the depth value of the clustering center point set in Step 4.3.1, and re-perform the next round of iterative clustering.
[0072] Step 4.3.4: Loop through Step 4.3.1 to Step 4.3.3 until at least one of the following conditions is met, then end the loop and output the pixel position information of the dynamic target class or the static background class in the dynamic target detection box.
[0073] Condition 1: The difference between the average pixels of the two classes and the clustering center is less than the set value a. In the present invention, a is set to 1.
[0074] Condition 2: The difference in the number of pixels belonging to the two classes in the target box is greater than the set value b. In the present invention, b is set to 1.7.
[0075] Step 4.4: Extract the object class pixel positions of all object detection boxes in this frame of image, project the obtained dynamic object class pixel positions onto an image of the same size as the original input image and output it in a binary form. Perform an OR operation on the pixel positions corresponding to the binary images generated by all object detection boxes in this frame of image, so that multiple images are superimposed and fused into a mask image output in a binary form, as Figure 6 shown;
[0076] Figure 5 The display is the effect after the corresponding Figure 4 dynamic object detection box is processed by the pixel-level segmentation method.
[0077] Step 5: After processing the grayscale image in the SLAM system with the mask image, the specific method is to perform an AND operation on the mask image output in a binary form and the grayscale image in SLAM. Among them, the pixel values at the positions where the dynamic object classes are located will be 0, while the other parts of the grayscale image will not change. The part where the pixel values in the grayscale image are set to 0 is the dynamic object information in the image, and the unchanged part of the grayscale image is the static background information. Use the static background information for the positioning and map construction of the mobile robot. In the present invention, combined with the ORB-SLAM2 system, the algorithm flow of the SLAM system can be summarized into the following sub-steps:
[0078] Step 5.1: Perform ORB feature extraction on the static image provided in Step 4. When the number of extracted feature points exceeds the set value, the SLAM system is initialized.
[0079] Step 5.2: Use the extracted ORB features combined with the previous frame of static background information to estimate the motion pose of the camera. The RGBD camera provides depth information, so the PnP method can be directly used to solve the camera pose.
[0080] PNP is a method of estimating the camera motion by projecting the matching points from the three-dimensional space onto the image plane and calculating the error with the observed data. This method is also called the reprojection error. The analytical PNP method only uses a small number of matching pairs to estimate the relative motion, and then optimizes the camera motion through a non-linear optimization method, which can effectively reduce the calculation amount.
[0081] Step 5.3: Minimize the reprojection error through a non-linear optimization method, that is, use the bundle adjustment method (BA) to optimize the local map.
[0082] Step 5.4. Optimize the pose using loop detection to correct the drift error. This process is divided into two parts, namely closed-loop detection and closed-loop correction. For closed-loop detection, first use the Bag of Words (Bow) model for detection, and then calculate the similarity transformation through the Sim3 algorithm. For closed-loop correction, mainly perform graph optimization of the closed-loop fusion of the Essential Graph to achieve the effect of adjusting and correcting the error.
[0083] The key point of the present invention is to use the new YOLOv5 algorithm to detect preset dynamic objects in the scene image, and then use the pixel-level segmentation method to extract the dynamic target mask image; perform feature point matching on the static background and then apply it to the ORB-SLAM2 system.
[0084] In this embodiment, the performance of the algorithm is evaluated on the publicly available TUM dataset. A total of 5 image sequences in the TUM dataset are used, including fr1_xyz, fr3_sitting_static, fr3_walking_halfsphere, fr3_walking_rpy, and fr3_walking_xyz. Among them, fr1 or fr3 represents different configuration files required to run the image sequence, the xyz suffix represents the movement of the camera in the x-y-z three-axis directions, the rpy suffix represents the rotation of the camera in the r-p-y three direction angles, the halfsphere suffix represents that the camera has an additional arc movement in space on the basis of xyz and rpy, the static suffix represents that the camera is relatively stationary, the sitting suffix represents the image sequence under low dynamics, and the walking suffix represents the high-dynamics image sequence.
[0085] The comparison results between the algorithm of this application and other algorithms are shown in Table 1.
[0086] Table 1. Comparison table of the positioning accuracy of the present invention
[0087]
[0088] Comparing the data in Table 1, for the high-dynamics image sequence, the positioning accuracy of the present invention has a huge improvement compared to the original ORB-SLAM2 system. At the same time, in the low-dynamics and pure static image sequences, the positioning accuracy of the present invention can achieve the same effect as that of the ORB-SLAM2.
[0089] Figure 7 and Figure 8 , intuitively demonstrates the effect achieved by the present invention. In the figure, groudtruth represents the true trajectory, CameraTrajectory_orb represents the trajectory of the ORB-SLAM2 system, and CameraTrajectory_my represents the trajectory of the present invention. Figure 7is the error of the trajectories of the three in the x-y-z axes directions. Figure 8 is the error of the trajectories of the three in the r-p-y three direction angles. As can be seen from the figure, in the dynamic environment, the positioning accuracy of the original system is inaccurate, with a huge drift error, and there is also a situation of tracking failure during the period from 63 seconds to 68 seconds. However, the present invention can accurately track the real trajectory in the dynamic environment, quickly reposition to find the pose, and better remove the influence of dynamic targets on the SLAM system.
[0090] In terms of running time cost, the present invention uses the simplest YOLOv5s model in YOLOv5 for training. The advantage of this network is that it is small in size and fast in speed. Only a general GPU is needed to accelerate the operation, and YOLOv5-s can achieve real-time effects. If the GPU performance is good, the detection speed of the YOLOv5-s model only needs 2.5ms per frame, that is, the detection speed is 400 frames per second.
[0091] As shown in Table 2, on the CPU i7-7700, the average running time of the ORB-SLAM2 system is about 30 milliseconds per frame, that is, the running speed is about 33 frames per second. Although the present invention adds a pixel-level segmentation method to process the image, the time cost of this method is extremely small, and its total time cost is basically the same as that of the original system. However, the original ORB-SLAM2 is not robust in the dynamic environment, while the present invention can remove the influence of dynamic targets on the SLAM system in the dynamic environment, not only has high accuracy in the dynamic environment, but also can run on the central processing unit (CPU) in real time.
[0092] Table 2. Comparison table of the time cost of the present invention
[0093]
[0094] The above are only the preferred embodiments of the present invention. Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make many possible changes and modifications to the technical solution of the present invention by using the methods and technical contents disclosed above, or modify it into an equivalent embodiment with equivalent changes. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the technical solution of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A SLAM method for removing dynamic objects based on an RGBD sensor, characterized in that, The method includes the following steps: Step 1: Obtain the pictures that need to pre-set the removal of dynamic target categories, and construct a training sample set and a test data set; Step 2: Use the training sample set and the test data set constructed in Step 1 to train and test the target detection neural network YOLOv5, and obtain the trained YOLOv5 model; Step 3: Use the target detection neural network YOLOv5 model trained in Step 2 to process the RGB images collected by the RGBD sensor, and judge whether there are dynamic targets in the images and the pixel coordinate positions where the dynamic target boxes are located; Step 4: Combine the pixel coordinate positions of the dynamic target boxes obtained in Step 3 and the depth image information of the RGBD sensor to judge whether there are dynamic targets in the current frame of the depth image of the RGBD sensor. If there are dynamic targets, use the pixel-level segmentation method to process the same pixel area drawn within the dynamic target box in the depth image corresponding to the current frame to obtain a mask image. If there are no dynamic targets, directly set the mask image to be empty; The pixel-level segmentation method specifically includes the following steps: Step 4.1: For all dynamic object detection boxes in each frame of RGB image, within the dynamic object detection box area corresponding to the pixel positions in the depth map of this frame image, uniformly distribute and take the depth values of N pixel points, and store them in container X i where the value of i ranges from 1 to N; Step 4.2: Use the absolute median deviation outlier algorithm for the depth values of the obtained N pixel points, divide them into two sets composed of depth values, and then use weighted summation for the two sets to obtain the average pixel depth values of the two sets respectively. Among them, the set with the smaller average pixel depth value is set as the dynamic target class depth value, and the set with the larger average pixel depth value is set as the static background class depth value; within the dynamic target box area output by the target detection neural network YOLOv5 model, the dynamic target is closer to the camera optical center than the static background; the absolute median deviation outlier algorithm detects one or several numerical values in the data that are significantly different from other numerical values and eliminates them. The size of the difference is judged according to the parameter n in the absolute median deviation outlier algorithm. Use the absolute median deviation outlier algorithm to extract the average depth information of the dynamic target, and its formula is: The absolute median deviation outlier algorithm includes the following sub-steps: Step 4.2.1, calculate X i the median value X of all elements in median ; Step 4.2.2: Calculate the absolute deviation of all elements from the median, and then calculate the median MAD of all elements' relative absolute deviations; Step 4.2.3: Determine the parameter n and adjust the parameter according to the absolute median deviation outlier algorithm formula; separate the outlier data and store it in the container outlier, and store other data in the container normal. Then calculate the average depth value of all pixels in the container outlier and the average depth value of all pixels in the container normal, and compare the average depth values of the two. Among them, the set with the smaller average pixel depth value is set as the dynamic target class depth value, and the set with the larger average pixel depth value is set as the static background class depth value; Step 4.3: For the dynamic target class depth value and the static background class depth value, use a distance-based clustering algorithm to cluster all pixels in the dynamic target detection box, and divide all pixels in the dynamic target box into the dynamic target class or the static background class. The obtained dynamic target class or static background class is a set of pixel points; Step 4.4: Extract the pixel positions of the dynamic object classes in all target detection boxes in the image, project the obtained pixel positions of the dynamic object classes onto an image of the same size as the original input image and output it in a binary form. Perform an OR operation on the pixel positions corresponding to the binary images generated by all target detection boxes in the image to superimpose and fuse multiple images into a single mask image, and the mask image is also output in a binary form; Step 5: Process the grayscale image in SLAM using the mask image. Specifically, perform an AND operation on the mask image output in binary form and the grayscale image in SLAM. The pixel values at the positions where the dynamic object classes are located will be 0, while the other parts of the grayscale image will remain unchanged. The part where the pixel values in the grayscale image are set to 0 is the dynamic object information in the image, and the unchanged part of the grayscale image is the static background information. Use the static background information for the positioning and map construction of the mobile robot.
2. The SLAM method for removing dynamic objects based on an RGBD sensor according to claim 1, wherein, In the above Step 1, after the test dataset labels are made, they are adjusted to the VOC dataset format for subsequent steps.
3. A SLAM method for removing dynamic objects based on an RGBD sensor according to claim 1, characterized in that, In the above Step 2, the simplest YOLOv5s network in the object detection neural network YOLOv5 is used for training. This object detection neural network has fewer parameters and faster running speed compared to other networks of the same type.
4. A SLAM method for removing dynamic objects based on an RGBD sensor according to claim 1, characterized in that, In the above Step 3, the pixel coordinate positions include the center point coordinates and the width and height of the target box, and both are normalized according to the size of the input image.
5. A SLAM method for removing dynamic objects based on an RGBD sensor according to claim 1, characterized in that, In the above Step 4.3, the distance-based clustering algorithm includes the following steps: Step 4.3.1: First, set the center points of the clustering. In this step, set the depth values of two center points, which are the depth values of the dynamic object class and the static background class obtained from Step 4.2 respectively; Step 4.3.2: Traverse all pixel points in the dynamic target detection box, calculate the Euclidean distances from the depth values of each pixel point to the depth values of the two set clustering center points respectively, compare the sizes of the two distance values, and divide all pixel points in the dynamic target detection box into the closer clustering center; Step 4.3.3: Re-select the depth values of the clustering center points, calculate the average pixels of the two sets of pixel points divided in Step 4.3.2 as the depth values of the clustering center points set in Step 4.3.1, and perform the next round of iterative clustering; Step 4.3.4: Loop through Step 4.3.1 to Step 4.3.3 until at least one of the following conditions is met, then end the loop and output the pixel position information of the dynamic object class or the static background class in the dynamic target detection box; Condition 1: Until the difference between the average pixels of the two classes and the clustering center is less than the set value a; Condition 2: The difference in the number of pixels belonging to the two classes in the target box is greater than the set value b; The parameters a and b are adjusted according to the actual application scenario.
6. A SLAM method for removing dynamic objects based on an RGBD sensor according to claim 1, characterized in that, In Step 5, the specific steps of the algorithm flow of the SLAM system are as follows: Step 5.1: Perform ORB feature extraction on the static image provided in Step 4. When the number of extracted feature points exceeds the set value, the SLAM system is initialized; Step 5.2: Estimate the motion pose of the camera by using the extracted ORB features combined with the static background information of the previous frame. The RGBD camera provides depth information, and the PnP method is used to solve the camera pose; Step 5.3: Use Bundle Adjustment (BA) to minimize the reprojection error and optimize the local map; Step 5.4: Use loop detection to optimize the pose and correct the drift error.
7. A SLAM method for removing dynamic objects based on an RGBD sensor according to claim 6, wherein, The said Step 5.4 is divided into two parts, namely loop detection and loop correction; the loop detection first uses the Bag of Words (Bow) model for detection, and then calculates the similarity transformation through the Sim3 algorithm; The said loop correction includes the graph optimization of the Essential Graph for loop fusion to achieve the effect of adjusting and correcting errors.
Citation Information
Patent Citations
Point cloud based indoor dynamic scene SLAM (Simultaneous Location and Mapping) method and system
CN106056643A
Environment data sharing platform, unmanned aerial vehicle and positioning method and system
CN106931963A