Dynamic obstacle identification and elimination method based on SLAM research
Through the improved Mask R-CNN network and iterative Gaussian algorithm, combined with IMU and depth data for self-motion compensation, the problem of dynamic obstacle occlusion in dynamic environments is solved, and dynamic obstacle recognition and removal with high accuracy and robustness are achieved, ensuring the stability and accuracy of three-dimensional map construction.
Patent Information
- Application Number
- CN202510266049.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-27
AI Technical Summary
In a dynamic environment, when a robot constructs a three-dimensional map, the problem of dynamic obstacle occlusion is a key challenge, resulting in reduced accuracy of the map, and some structures of the static environment are misidentified or map errors are generated.
A dynamic obstacle recognition and culling method based on SLAM research is adopted, and the improved Mask R-CNN network is constructed, combined with IMU data and depth data for self-motion compensation, and the iterative Gaussian algorithm is used to generate and update masks to achieve accurate identification and culling of dynamic obstacles.
It improves the accuracy of dynamic object detection and culling, enhances the robustness of the system, optimizes the computing efficiency and real-timeness, and ensures the stability and accuracy of static map construction.
Smart Images

Figure CN120220106A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the technical field of dynamic obstacle recognition and removal, and particularly to a method for dynamic obstacle recognition and removal based on SLAM research. Background Art
[0002] With the rapid development of robot technology in recent years, the Simultaneous Localization and Mapping (SLAM) technology has become the core support in the fields of robot autonomous navigation and environmental perception. This technology aims to enable a robot to accurately locate its own position and precisely construct a map of the surrounding environment while moving in an unknown environment.
[0003] Currently, the methods for dynamic obstacle recognition and removal based on SLAM mainly include techniques such as sensor detection, motion estimation, graph optimization, time series and data association, and semantic segmentation and object recognition methods.
[0004] The sensor-based detection methods include the combination of lidar and visual sensors. By comparing the changes in lidar point clouds over consecutive time periods, moving objects can be detected. In addition, deep learning methods, such as convolutional neural networks (CNNs), have also shown good results in the recognition of dynamic objects. The combination of lidar and visual sensors is a common technical path. Using lidar point cloud data and image information, by comparing the changes over consecutive time periods, the presence of dynamic objects can be detected.
[0005] In a dynamic environment, when a robot constructs a three-dimensional map, the problem of dynamic obstacle occlusion is a key challenge. Dynamic obstacles not only affect the accuracy of the map but may also cause some structures in the static environment to be mislabeled or map errors to occur. This is because the presence and movement of dynamic objects may occlude some static environmental features, making it impossible for the SLAM system to accurately reconstruct the complete three-dimensional environment. Although the existing methods for dynamic obstacle recognition and removal based on SLAM have made some progress in dynamic scenarios, there are still some significant deficiencies in solving the problem of dynamic obstacle occlusion.
[0006] Specifically, these deficiencies are mainly reflected in several aspects: 1) Due to the occlusion of dynamic obstacles, some sensor data may be lost, resulting in the robot being unable to obtain complete environmental information. This information loss will affect the recognition accuracy of dynamic obstacles and further affect the accuracy of map construction.
[0007] 2) Data synchronization and fusion issues also exacerbate the challenges brought by dynamic obstacle occlusion. When a multi-sensor system processes data from different sensors, errors may occur in the information fusion process due to differences in time synchronization, perception range or update frequency, further affecting the accurate identification and removal of dynamic obstacles.
[0008] 3) Inaccurate estimation of the motion trajectory of dynamic objects is also a prominent problem. In the case of occlusion, traditional motion estimation algorithms (such as Kalman filtering, particle filtering, etc.) may not be able to correctly track the motion of dynamic objects, especially in the case of high-speed or irregular motion, resulting in estimation deviation of the object position, thus affecting the removal of dynamic obstacles and the accuracy of the 3D map.
[0009] Therefore, it is necessary to improve one or more problems existing in the above-mentioned related technical solutions.
[0010] It should be noted that this section is intended to provide background or context for the technical solutions of the present disclosure stated in the claims. The description herein is not admitted to be prior art by virtue of being included in this section. Summary of the invention
[0011] The purpose of the embodiments of the present disclosure is to provide a method for dynamic obstacle recognition and removal based on SLAM research, thereby overcoming one or more problems caused by the limitations and defects of related technologies at least to a certain extent.
[0012] According to a first aspect of an embodiment of the present disclosure, a method for identifying and removing dynamic obstacles based on SLAM research is provided, the method comprising: Constructing an improved Mask R-CNN network, acquiring an original image, and inputting the original image into the improved Mask R-CNN network; wherein a mask culling branch is added to the FCN layer of the improved Mask R-CNN network; According to the IMU data and depth data, the camera's self-motion parameters are calculated, and the predicted trajectory of the object in the original image is reversely compensated according to the camera's self-motion parameters; The original image is input into the improved Mask R-CNN network for feature extraction, and the mask is generated and updated in combination with the iterative Gaussian algorithm; Dynamic obstacles are removed according to the mask to obtain an effect image of dynamic obstacle removal.
[0013] Furthermore, the improved Mask R-CNN network includes: The backbone network ResNet50, FPN network, region projection network RPN, RoIAlign layer and full convolution network FCN are sequentially connected in series; among them, the full convolution network FCN includes mask elimination branch and Mask branch.
[0014] Furthermore, the loss function of the RoIAlign layer is as follows:
[0015] wherein, is the classification loss, is the bounding box, is the mask branch loss, is the mask rejection branch loss.
[0016] Furthermore, in the step of calculating the self-motion parameters of the camera according to the IMU data and the depth data, and performing backward compensation on the predicted trajectory of the object in the original image according to the self-motion parameters of the camera, it includes: Obtain the motion information in the IMU data of the camera or sensor in real time, and correct the motion trajectory of the object in combination with the state estimation of the object in the depth data; wherein, the motion information includes the rotation matrix and the translation vector ; By continuously updating the motion information of the camera or sensor, continuously adjust the predicted position of the object, so as to eliminate the error caused by the self-motion of the device.
[0017] Furthermore, in the step of inputting the original image into the improved Mask R-CNN network for feature extraction and generating and updating the mask in combination with the iterative Gaussian algorithm, it includes: Input the original image into the backbone network ResNet50 and the FPN network for feature extraction to obtain the feature map; Input the feature map into the region projection network RPN to determine the anchor boxes to obtain the candidate regions; Input the candidate regions and the feature map into the RoIAlign layer, and use the bilinear interpolation method to map the candidate regions onto the feature map for alignment; Input the aligned feature map into the Mask branch of the fully convolutional network FCN to generate the binary mask; Perform dynamic background reconstruction according to the binary mask, and introduce the iterative Gaussian algorithm for optimization to obtain the mask of the dynamic obstacle.
[0018] Furthermore, in the step of inputting the feature map into the region projection network RPN to determine the anchor boxes to obtain the candidate regions, it includes: Input the feature map into the region projection network RPN to generate pixel anchor boxes of different sizes and aspect ratios; Use the Softmax method to classify each pixel anchor box, and calculate the foreground probability and background probability of each pixel anchor box; Perform bounding box regression based on the foreground probability and background probability of each pixel anchor box to obtain an initial regression; Modify the pixel anchor box using the regression prediction value, sort the foreground anchor box results according to the foreground probability and background probability, and use the non-maximum suppression algorithm to delete pixel anchor boxes below the threshold or with a high overlap rate to obtain the candidate region RoI; wherein, the candidate region contains dynamic obstacles or static environmental features.
[0019] Further, when inputting the candidate region and the feature map into the RoIAlign layer and using the bilinear interpolation method to map the candidate region onto the feature map for alignment, it includes: The Region Proposal Network RPN generates a series of target regions through the anchor box mechanism and filters them according to foreground and background information; The input image where the candidate region is located extracts features through ResNet50 to obtain a high-dimensional feature map; The RoIAlign layer converts the coordinates of the candidate region in the original image into the feature map coordinate space and calculates the floating-point coordinates of the candidate region; wherein, the candidate region is divided into grid cells of a fixed size; For each sampling point of the grid cell, use the bilinear interpolation method to calculate the corresponding feature value from the feature map; Each candidate region is mapped to a feature map region of a fixed size to ensure consistent spatial alignment accuracy for candidate regions of different scales in subsequent classification and segmentation tasks.
[0020] Further, when inputting the aligned feature map into the Mask branch of the fully convolutional network FCN to generate a binary mask, it includes: The feature map processed by the RoIAlign layer is input into the Mask branch network, and the depth features of the target region are extracted through the fully convolutional network FCN; wherein, an upsampling operation is performed on the feature map output by the RoIAlign layer; The fully convolutional network FCN extracts multi-scale features through convolutional layers of different depths, thereby ensuring that the spatial dimensions of the input and output remain consistent and avoiding the loss of spatial information caused by the reduction of the feature map size; Calculate the object probability of each pixel and generate a binary mask; wherein, the segmentation of the binary mask is a binary classification task, and the sigmoid activation function is applied to each pixel to map the output value to between [0,1], indicating the probability that the pixel belongs to the object of this category.
[0021] Further, when performing dynamic background reconstruction based on the binary mask and introducing the iterative Gaussian algorithm for optimization to obtain the mask of the dynamic obstacle, it includes: Generate a pixel-level binary mask of dynamic objects using the segmentation branch of Mask R-CNN; Use a fully convolutional network (FCN) to perform multi-layer convolution on the feature map of the candidate region to extract high-dimensional semantic information, and combine it with the classification branch of Mask R-CNN to identify dynamic objects; Use semantic segmentation to further verify the dynamic object category to improve the accuracy of rejection; Through pixel-level logical operations, apply the binary mask to the original image to perform a region growing algorithm on the depth map, expand the dynamic object region, and ensure the integrity of rejection; Model the background through an iterative Gaussian algorithm to distinguish dynamic objects from static backgrounds and make background reconstruction more accurate; Fill the rejected area with background information from historical frames; Use the optical flow method to predict the pixel motion trend of the front and back frames to supplement the occluded area.
[0022] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: In the embodiments of the present disclosure, through the above dynamic obstacle recognition and rejection method based on SLAM research, on the one hand, self-motion compensation solves the influence of camera or sensor self-motion on the prediction accuracy of dynamic objects. By real-time estimating the camera motion and compensating the object prediction results, the accuracy of dynamic object detection and rejection is greatly improved. Especially in dynamic scenarios, it is particularly applicable to fields such as unmanned driving and robot navigation. On the other hand, the fusion of the improved Mask R-CNN and the iterative Gaussian algorithm provides a solution for multi-algorithm collaboration. The improved Mask R-CNN accurately segments dynamic objects, while the iterative Gaussian algorithm is used to accurately distinguish the background from dynamic objects, effectively improving the robustness of the system. Combining semantic segmentation and depth map region growing technology, this method can accurately identify and reject dynamic objects, especially in complex interaction and occlusion situations, optimizing the computational efficiency and real-time performance. It can predict the future position of objects based on their motion trajectories, achieve early rejection of dynamic obstacles, and ensure the stability and accuracy of static map construction. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 Show a dynamic obstacle recognition and rejection method based on SLAM research in an exemplary embodiment of the present disclosure; Figure 2 A block diagram of a system corresponding to a dynamic obstacle recognition and removal method based on SLAM research in an exemplary embodiment of the present disclosure is shown; Figure 3 A schematic diagram showing the principle of motion compensation in an exemplary embodiment of the present disclosure is shown; Figure 4 An improved Mask R-CNN network architecture diagram in an exemplary embodiment of the present disclosure is shown; Figure 5 A dotted grid feature map in an exemplary embodiment of the present disclosure is shown; Figure 6 A flowchart showing the mask segmentation process in an exemplary embodiment of the present disclosure is shown; Figure 7 A specific process of the removal algorithm in an exemplary embodiment of the present disclosure is shown; Figure 8 The dynamic feature removal effect in an exemplary embodiment of the present disclosure is shown. Detailed implementation manners
[0025] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.
[0026] In addition, the accompanying drawings are only schematic illustrations of the embodiments of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities.
[0027] In this example embodiment, a dynamic obstacle recognition and removal method based on SLAM research is provided. Referring to Figure 1 as shown in, the dynamic obstacle recognition and removal method based on SLAM research may include: step S101 to step S104.
[0028] Step S101: Construct an improved Mask R-CNN network, obtain an original image, and input the original image into the improved Mask R-CNN network; wherein, a mask removal branch is added to the FCN layer of the improved Mask R-CNN network; Step S102: Calculate the self-motion parameters of the camera according to the IMU data and the depth data, and perform reverse compensation on the predicted trajectory of the object in the original image according to the self-motion parameters of the camera; Step S103: Input the original image into the improved Mask R-CNN network for feature extraction, and combine the iterative Gaussian algorithm to generate and update the mask; Step S104: Remove the dynamic obstacles according to the mask to obtain the effect image after removing the dynamic obstacles.
[0029] Through the above method for dynamic obstacle recognition and removal based on SLAM research, on the one hand, self-motion compensation solves the influence of camera or sensor self-motion on the prediction accuracy of dynamic objects. By real-time estimating the camera motion and compensating the object prediction results, the accuracy of dynamic object detection and removal is greatly improved. Especially in dynamic scenes, it is especially applicable to fields such as unmanned driving and robot navigation. On the other hand, the fusion of the improved Mask R-CNN and the iterative Gaussian algorithm provides a solution for multi-algorithm cooperation. The improved Mask R-CNN accurately segments dynamic objects, while the iterative Gaussian algorithm is used to accurately distinguish the background from dynamic objects, effectively improving the robustness of the system. Combining semantic segmentation and depth map region growing technology, this method can accurately identify and remove dynamic objects, especially in complex interaction and occlusion situations, optimizing the computational efficiency and real-time performance. It can predict the future position of objects according to their motion trajectories, realize the early removal of dynamic obstacles, and ensure the stability and accuracy of static map construction.
[0030] Next, reference will be made to Figures 1 to 8 to explain each step of the above method for dynamic obstacle recognition and removal based on SLAM research in the present exemplary embodiment in more detail.
[0031] In a specific embodiment, in order to solve the challenges brought by dynamic obstacle occlusion, insufficient accuracy and robustness of dynamic obstacle recognition, and feature matching problems in SLAM mapping during the process of a robot constructing a three-dimensional map in a dynamic scene. Specifically, this research proposes an innovative dynamic obstacle recognition and removal solution to enhance the performance of the SLAM system in complex dynamic environments, ensuring the accuracy, stability, and real-time performance of the three-dimensional map. It improves the generalization ability of the model, refines the detection and recognition of dynamic obstacles, and improves the problem of insufficient feature matching in the SLAM mapping process in dynamic scenes. The main research objectives are as follows: 1. Solve the problem of dynamic obstacle occlusion: The primary objective of this application is to propose an effective method for dynamic obstacle extraction and behavior prediction aiming at the influence of dynamic obstacles on the robot's three-dimensional map construction in a dynamic environment. By improving data acquisition, fusion, and analysis methods, the interference of dynamic obstacle occlusion on the features of the static environment is reduced, thereby improving the accuracy of map construction and avoiding misidentification or errors in the static part.
[0032] 2. Improve the accuracy and robustness of dynamic obstacle recognition: An advanced motion compensation method is proposed to improve the efficiency and accuracy of dynamic obstacle detection. To enhance the navigation and positioning performance of the robot in a dynamic environment, the dynamic obstacle recognition algorithm is optimized by adding a dynamic obstacle removal module to the target detection algorithm, strengthening the accuracy and robustness of the system. Through sensor data fusion, deep learning algorithms, and precise motion estimation, dynamic obstacles can be accurately identified and removed in real time in a complex dynamic environment. Even when dynamic obstacles are occluded or moving, the recognition accuracy can be effectively maintained, reducing system misjudgment and missed detection.
[0033] 3. Solution for insufficient feature matching in SLAM mapping: In a dynamic scenario, an important problem faced by the SLAM algorithm is insufficient feature matching, especially in the presence of dynamic obstacles. A feature matching optimization solution based on dynamic obstacle recognition and removal is proposed, which can effectively handle false matches or omissions caused by dynamic objects during the map construction process, improving the feature matching accuracy and map construction quality of the SLAM system.
[0034] In a specific embodiment, to address the dynamic interference and occlusion problems faced by the robot during mapping in a dynamic environment, a dynamic target module is added by improving the Mask R-CNN network. This method aims to enhance the mapping accuracy and robustness of the robot in a dynamic scenario, especially for the detection and removal process of dynamic obstacles. Traditional target detection algorithms can effectively identify objects but still have deficiencies in removing dynamic obstacles, especially when dealing with moving obstacles. The improved Mask R-CNN network can not only perform accurate dynamic target detection but also add a dynamic obstacle removal module (i.e., the mask removal branch), which can effectively remove dynamic obstacles affecting mapping. In this way, when the robot performs real-time SLAM mapping, it can better handle the occlusion problem of moving obstacles, thus ensuring the accuracy of map construction and the stability of the system.
[0035] Figure 2Shows the main framework of this application. The overall process can be divided into two core parts: camera motion compensation and dynamic object detection and removal. In terms of camera motion compensation, this application integrates IMU (Inertial Measurement Unit) and depth data, and designs an advanced motion compensation algorithm that can effectively filter out the rotation and translation errors caused by the camera's self-motion, thereby improving the stability and accuracy of mapping. For the handling of dynamic obstacles, the Mask R-CNN network is used for the recognition and classification of dynamic objects. During the object detection process, an iterative Gaussian fitting algorithm is introduced to more accurately extract the contours of dynamic obstacles. This process effectively avoids the interference of dynamic obstacles to map construction, ensuring the real-time performance and accuracy of the SLAM system. Through semantic segmentation technology, dynamic obstacles are accurately segmented and removed, further optimizing the quality of the map and the robustness of the SLAM system. It can significantly improve the performance of the SLAM system in a complex dynamic environment, providing a more effective solution for real-time mapping tasks. The specific research steps are as follows: I. Extraction and Prediction of Dynamic Features The extraction of dynamic features and motion prediction are the core parts of dynamic object detection. Especially in a complex dynamic environment, how to accurately identify and predict the motion trajectory of an object is crucial for the removal of dynamic obstacles. In this part, the iterative Gaussian algorithm is combined with the object detection network (Mask R-CNN) to achieve the effective extraction and motion prediction of dynamic objects, thereby improving the accuracy of dynamic object removal. The extraction of dynamic objects depends on the modeling of the scene background and the detection of the foreground. By combining self-motion compensation, deep learning segmentation, iterative Gaussian algorithm, and motion prediction model, the system can accurately extract the dynamic object features and predict the future position and trajectory of the object, effectively reducing the interference of dynamic obstacles to static map construction. It is mainly described from the following two points: 1. Extraction of Dynamic Features The extraction of dynamic features first depends on the analysis of the object's motion pattern. In a complex dynamic environment, the motion of an object is usually affected by different factors, including the object's own motion law, the self-motion of the camera or sensor, etc. To accurately extract the dynamic features of an object, it is first necessary to eliminate the errors caused by the self-motion of the camera or sensor through self-motion compensation technology. By estimating the rotation and translation of the camera in real time, the system can effectively adjust the motion trajectory prediction of the object, ensuring the accurate extraction of motion features.
[0036] During the detection of dynamic objects, the Mask R-CNN network plays a crucial role. Through fine-grained instance segmentation, Mask R-CNN can generate a binary mask for each dynamic object, thus accurately identifying the position and shape of the object. Through this mask, the dynamic features of the object in the image can be extracted, including key parameters such as position, shape, and speed. In addition, the iterative Gaussian algorithm is also applied to background modeling to help the system distinguish dynamic objects from static backgrounds, further improving the accuracy of dynamic feature extraction. The iterative Gaussian algorithm is an application and extension of the Gaussian mixture model. It uses multiple Gaussian distributions to model the dynamic background and continuously updates the parameters of the Gaussian distributions through iteration to adapt to the dynamic environment. Specifically, the iterative Gaussian algorithm mainly performs dynamic object extraction through the following steps: 1) Initialization stage: Assign multiple Gaussian distributions to each pixel. Usually, the Gaussian distribution parameters of each pixel are initialized by learning the pixel values in historical image frames, including the mean, variance, and weight.
[0037] 2) Update stage: Whenever a new image frame arrives, the algorithm calculates the matching degree between the observed value of each pixel in this frame and its corresponding Gaussian distribution. If the current pixel value matches well with the mean of a certain Gaussian distribution (i.e., the distance between the pixel value and the mean is small), then this Gaussian distribution updates its mean and variance parameters; if the matching degree is low, a new Gaussian distribution will be generated, and the observed value of this pixel will be classified as the foreground.
[0038] 3) Background and foreground separation: By setting a threshold, it is judged whether each pixel belongs to the background or the foreground. When the value of a pixel matches well with multiple Gaussian components of the background, it is considered that this pixel belongs to the background; if the pixel value matches poorly with the Gaussian distribution in the background model, it is considered that this pixel belongs to the dynamic object (foreground).
[0039] 4) Dynamic object extraction: When the pixels of an object continuously change in multiple frames of images, the GMM can detect these changes and distinguish them as dynamic objects. In this process, the iterative Gaussian algorithm continuously optimizes the background modeling, enabling the system to adapt to changes in lighting, seasons, or the environment, thus effectively detecting real dynamic objects.
[0040] In a dynamic environment, the observed value of each pixel can be represented by multiple Gaussian distributions. Specifically, the observed value of each pixel is composed of the weighted sum of multiple Gaussian components: (1) where is the probability of observing the pixel at time is the weight of the n-th Gaussian distribution, and are the mean and variance of the n-th Gaussian distribution respectively, and is the number of Gaussian components. In this way, it is possible to adapt to different background information of each pixel according to the change of the scene, and improve the recognition accuracy of dynamic objects.
[0041] Combining the high-precision segmentation results of Mask R-CNN with the iterative Gaussian algorithm can simultaneously achieve the accurate segmentation of dynamic objects and the accurate modeling of the background. Mask R-CNN provides the accurate contour of each dynamic object, while the iterative Gaussian algorithm is responsible for identifying and adapting to background changes, further improving the accuracy of dynamic object extraction. Under this combination, the system can accurately extract dynamic objects, maintain the stability and consistency of the background, thus optimizing the dynamic object removal process and ensuring the accuracy of the static map.
[0042] 2. Motion Behavior Prediction The action prediction of dynamic objects plays a key role in avoiding their impact on the construction of static maps, and accurate action prediction is the core of it. When performing action prediction, the extraction of moving objects is an important pre-step. Generally, the contour of moving objects can be extracted completely and accurately. However, when the scale of moving objects in the image plane is too small, the region of interest (ROI) may not converge. In response to this special situation, the method of finding connected components is used to determine the region most likely to be the moving target in the un-fused ROI. After such an operation, the accurate extraction of the moving target can be achieved in most cases.
[0043] In terms of ego-motion compensation, previous algorithms have problems of high computational requirements or insufficient accuracy. To solve this problem, the proposed algorithm performs ego-motion compensation by fusing depth and IMU data, which takes into account computational efficiency, accuracy, and robustness. The basic principle of ego-motion compensation is to estimate the motion of the camera or sensor (such as rotation and displacement), and then make corresponding adjustments to the prediction results of the object. It includes rotation and translation, where rotation is represented by the rotation matrix and translation is represented by the translation vector . The initial position of the object in the camera coordinate system is , and the new position after the camera motion can be calculated by the formula , where the rotation matrix is a matrix that satisfies ( A 3×3 matrix (where is the identity matrix). During the prediction process, the system will obtain the motion information of the camera or sensor in real time, and combine it with the state estimation of the object to correct the motion trajectory of the object. By continuously updating the motion parameters of the camera or sensor, that is, the rotation matrix and the translation vector Figure 3 , the predicted position of the object is continuously adjusted to eliminate the errors caused by the movement of the device itself, as shown in . Both the rotation and translation of the event camera are considered in
[0044] II. Improvement of Mask R-CNN Dynamic Object Detection The detection of dynamic objects is achieved by using Mask R-CNN to perform instance segmentation on the target while combining motion compensation and the feature elimination of dynamic obstacles. By combining with an IMU sensor or other motion estimation algorithms, it is identified which objects are dynamic. These dynamic objects will be reflected in the generated mask. The mask branch is a small fully convolutional network (FCN) that predicts the segmentation mask of each RoI at the pixel level. Considering the Faster R-CNN framework, the implementation and training of Mask R-CNN are relatively simple and can adapt to various flexible architecture designs. In addition, the mask branch only adds a small amount of computational overhead, so the system can run quickly and conduct efficient experiments. Although Mask R-CNN is essentially an extension of Faster R-CNN, the construction of the mask branch is crucial for obtaining high-quality segmentation results. By improving Mask R-CNN, a new mask elimination branch is added. This improvement enables the effective elimination of the masks of dynamic obstacles after generating the binary mask, providing support for the final frame image restoration, thereby constructing a more accurate static 3D map. The improved network architecture is shown in Figure 4 . This network architecture mainly uses the backbone network ResNet50 of Mask R-CNN. It constructs hierarchical residual class connections to obtain more fine-grained features. FPN is integrated with the backbone network, and the backbone integrates three channels and extracts feature information of various sizes.
[0045] In object detection and segmentation tasks, the recognition and extraction of image features are crucial, which directly affect the accuracy of detection and segmentation. To extract richer and more expressive features, ResNet50 introduces hierarchical residual connections to obtain more fine-grained features, thereby enhancing the network's ability to capture information at different scales. To further improve the processing ability of multi-scale features, the Feature Pyramid Network (FPN) is integrated with the backbone network, and the backbone network extracts multi-scale feature information through multiple channels to better handle objects of different sizes. The RPN first generates anchor boxes of different sizes and aspect ratios. The softmax method is used to determine the anchor boxes. After feature extraction, the generated feature map will be input into the Region Proposal Network (RPN). The task of the RPN is to perform object detection by generating anchor boxes with different sizes and aspect ratios. The size and aspect ratio of each anchor box can be expressed as: (2) where represents the center coordinates of the anchor box, and represent the width and height of the anchor box. The RPN classifies each anchor box using the Softmax method and calculates the probability that each pixel anchor box belongs to the foreground or background, which is expressed as: (3) where, and represent the scores of the foreground and background parts of the anchor box respectively. Through this step, the RPN can classify the anchor boxes and provide input for subsequent bounding box regression. Bounding box regression is performed through the RPN (Region Proposal Network) to adjust the positions of the generated anchor boxes. The regression target for each anchor box is to precisely adjust the anchor box through the predicted offsets (such as changes in position, width, and height). The regression target for each anchor box can be expressed as to better cover the boundaries of the actual object. The regression error is calculated through the smooth loss. This loss function can effectively measure the difference between the predicted box and the ground truth box and help optimize the regression process. Through regression, the RPN can precisely adjust the positions of the anchor boxes and further improve the accuracy of object detection. Initial regression is obtained through bounding box regression, and then the anchor boxes are modified using the regression prediction values. The foreground anchor box results are sorted according to the scores, and boxes with scores below the threshold or high overlap rates are removed using the non-maximum suppression algorithm.
[0046] After generating multiple anchor boxes, the RPN uses the non-maximum suppression (NMS) algorithm to remove redundant anchor boxes. First, the RPN sorts the anchor boxes according to the score of each anchor box and calculates the overlap (IoU) between the anchor boxes. If the overlap between the anchor boxes exceeds the set threshold, the anchor box with a lower score will be deleted. The core idea of NMS is to retain the anchor boxes with higher scores and less overlap with other boxes to reduce the interference of redundant boxes and ensure that the finally output anchor boxes are the optimal foreground regions.
[0047] Through the above steps, the RPN can generate high-quality object region proposals, providing accurate inputs for subsequent tasks such as object classification, segmentation, and dynamic obstacle removal. Finally, a feature map is obtained. The feature map and the remaining anchor boxes in the RPN are sent to RoIAlign. RoIAlign is a key module in Mask R-CNN. It solves the problem of accuracy loss caused by quantization operations in traditional convolutional neural networks by removing the quantization step and using the bilinear interpolation method to accurately map the features in each RoI to the feature map, ensuring pixel-level alignment. Figure 5 The dashed grid in the figure represents the feature map processed by the Region Proposal Network (RPN), the solid line represents the RoI (with 2×2 bins in this example), and the dots represent the 4 sampling points in each bin. RoIAlign calculates the value of each sampling point through bilinear interpolation from the nearby grid points on the feature map. No quantization is performed on any coordinates involved in the RoI, its bins, or the sampling points. RoIAlign maps each RoI to the original feature map and calculates the feature values of each pixel point within the RoI region. Different from traditional pooling operations, RoIAlign does not map the input features to a fixed grid but uses bilinear interpolation to more accurately obtain the features at each position. RoIAlign removes the quantization process, uses the bilinear interpolation algorithm to map the remaining anchor boxes in the RPN to the feature map, effectively avoids the problem of pixel mismatch, and performs pixel-level instance segmentation tasks.
[0048] III. Mask Update of RoIAlign Layer Traditional instance segmentation tasks are highly challenging due to their high accuracy requirements. Instance segmentation not only requires accurately detecting all objects in the image but also precisely segmenting each object instance on this basis, making the task more complex. Instance segmentation combines classic computer vision techniques in object detection, and its core goal is to classify each object and localize the object through bounding boxes and semantic information. This process is particularly important for handling dynamic objects because the shape, position, and occlusion of dynamic obstacles may change over time.
[0049] To address this issue, Mask R-CNN, as an extension of Faster R-CNN, introduces a new branch to predict the segmentation mask for each region of interest (RoI). This mask branch works in parallel with the original classification and bounding box regression branches, enabling the model to perform pixel-level object segmentation while detecting objects, thus more accurately identifying and distinguishing objects. This extension plays an important role in dynamic obstacle detection, being able to effectively classify dynamic objects and provide a better foundation for subsequent semantic segmentation. In this way, Mask R-CNN not only enhances the object detection ability but also provides stronger robustness and accuracy for dealing with dynamic obstacles and complex scenarios.
[0050] In the improved Mask R-CNN, the multi-task loss for each sampled RoI is defined during the training process to ensure that the model can effectively perform classification, bounding box regression, mask prediction, and dynamic obstacle removal. Specifically, the loss function can be expressed as: (4) Where the classification loss and the bounding box are defined to use the cross-entropy loss and smooth loss as in the original Mask R-CNN. The mask branch is optimized through binary cross-entropy to ensure that the mask for each RoI can accurately represent the shape of the object. The dynamic obstacle removal loss is a newly added module designed to avoid these dynamic obstacles from interfering with map construction and pose estimation by judging and removing the influence area of dynamic objects in each mask. The output of the mask branch for each RoI is a two-dimensional output, where is the number of target classes, is the resolution of the mask, and each class corresponds to a binary mask representing the position and shape of that class in the image. For each pixel, the value of the mask is predicted through the sigmoid function, and the output value ranges between 0 and 1, indicating the probability that the pixel belongs to the object. This way provides pixel-level spatial information for each object, enabling the precise segmentation of the object's boundary.
[0051] Through the collaborative work with the mask branch, dynamic obstacles can be accurately identified and effectively removed, thus ensuring that the robot can maintain high-precision mapping and localization in a dynamic scenario. The feature map processed by the region proposal network (RPN) and the filtered anchor boxes will be passed to the RoIAlign module.
[0052] This method ensures that the feature map of each RoI can be accurately aligned, avoiding the pixel mismatch problem caused by quantization error, thereby improving the accuracy and effect of instance segmentation. Based on the precisely aligned feature map output by RoIAlign, the network further passes the feature map to two branches. The first branch is the mask branch network, which uses a fully convolutional network (FCN) to predict and generate the segmentation mask of the object. By generating a binary mask, this branch can accurately identify the shape of each object and remove the mask of dynamic obstacles from the image to prevent them from interfering with the final mapping result.
[0053] 4. Mask generation and segmentation removal After generating the binary mask, a series of candidate regions (RoIs) are generated by improving the RPN of Mask R-CNN. These regions usually contain the potential locations of the target objects. Subsequently, the RoIAlign module accurately maps these candidate regions to the feature map to avoid the errors caused by traditional quantization operations. The goal of RoIAlign is to accurately align the feature map of the RoI region. It performs pixel-level alignment through bilinear interpolation method, ensuring that the fine-grained features of the target region in the feature map can be preserved.
[0054] After RoIAlign, each RoI in the image is passed to the Mask branch network for mask generation. The Mask network generates binary masks based on the Fully Convolutional Network (FCN). Each mask is used to represent the position and shape of a target object in the image and is predicted by the pixel-level sigmoid activation function. Specifically, the mask is , for each pixel For example, the value of the mask represents the probability that the pixel belongs to the target object. The mask prediction formula is as follows: (5) in, express The probability of belonging to the target object, is the sigmoid activation function, is the network weight, is extracted from the feature map and the pixel Related features, is a paranoid term, and the generation process is as follows Figure 6 shown. Figure 6 It mainly demonstrates two processes: outputting the original image to the network for feature extraction and dynamic object detection, and then combining the iterative Gaussian algorithm to generate and update the mask.
[0055] During the dynamic background modeling process, an iterative Gaussian algorithm is introduced to optimize the background reconstruction. It is used to model the background of the scene and can effectively distinguish static background and dynamic objects in a dynamic environment. The Gaussian mixture model can adapt to changes in the scene over time by continuously updating the distribution parameters of pixels, accurately distinguishing the foreground (dynamic objects) from the background. The segmentation branch of Mask R-CNN combined with the iterative Gaussian algorithm can accurately identify the contours of dynamic objects, thereby generating a pixel-level mask to separate these dynamic objects from the background. This fine segmentation mask provides a basis for subsequent background restoration, facilitating the system to accurately restore the occluded background information after removing dynamic objects.
[0056] The dynamic background reconstruction in the research uses the segmentation results of Mask R-CNN for feature extraction and screening, and combines semantic segmentation with multi-view geometric information to ensure the integrity and authenticity of the background after removing dynamic objects. Using the above method, moving objects can be better removed. The segmentation network can better remove prior dynamic objects, but has limited effect on distinguishing non-prior dynamic objects in the scene, and a multi-view geometric algorithm is used for further processing. The proposed algorithm flow for removing dynamic objects is as Figure 7 shown. Dynamic points are detected based on the motion relationship between key frames. After the dynamic points are detected, they are divided into dynamic points with semantic information and dynamic points without semantic information by judging whether they have semantic labels; for the feature points with semantic information, semantic contour search is performed on the semantic map, and for the dynamic points without semantic information, region growing is performed on the depth map, making full use of semantic information to reduce the number of region growing seed points and improve the operating efficiency of the system. The masks of dynamic objects without semantics and dynamic objects with semantic information are fused to obtain a complete mask of dynamic objects.
[0057] The next step is to perform the removal of dynamic obstacles according to the mask. To remove dynamic obstacles, first, it is necessary to use the classification branch of Mask R-CNN to determine which objects are dynamic. The masks of these dynamic objects will be identified and marked. Then, using logical operations (such as binary and multiplication operations), the pixels in the dynamic object region will be set to zero and removed from the image. Specifically, for the generated mask and the image , the removal process is: (6) is the binary mask of the dynamic object, and the value 1 represents the dynamic object region. By multiplying the image with pixel by pixel, the dynamic object region will be deleted from the dynamic object, and the final image . An RGB-D camera is used to obtain RGB images and depth images. The RGB images are processed by a semantic segmentation network to obtain pixel-level semantic information, and the semantic information is used to remove the feature points on the prior dynamic objects in the images. The feature points corresponding to non-prior dynamic objects are further detected by a multi-view geometry algorithm, and the detection results of the multi-view geometry and the semantic segmentation network are cross-validated to obtain a complete dynamic region. After filtering out the feature points on the dynamic region, it enters the tracking thread to obtain a more accurate pose. The effect of dynamic feature removal is as Figure 8 shown.
[0058] Through the above method for dynamic obstacle recognition and removal based on SLAM research, on the one hand, self-motion compensation solves the influence of the self-motion of the camera or sensor on the prediction accuracy of dynamic objects. By real-time estimating the camera motion and compensating the object prediction results, the accuracy of dynamic object detection and removal is greatly improved. Especially in dynamic scenes, it is particularly applicable to fields such as unmanned driving and robot navigation. On the other hand, the fusion of the improved Mask R-CNN and the iterative Gaussian algorithm provides a solution for multi-algorithm cooperation. The improved Mask R-CNN accurately segments dynamic objects, while the iterative Gaussian algorithm is used to accurately distinguish the background from dynamic objects, effectively improving the robustness of the system. Combining semantic segmentation and depth map region growing technology, this method can accurately identify and remove dynamic objects, especially in complex interaction and occlusion situations, optimizing the computational efficiency and real-time performance. It can predict the future position according to the motion trajectory of the object, realize the early removal of dynamic obstacles, and ensure the stability and accuracy of static map construction.
[0059] In summary, this method has strong patentability and application prospects in dynamic object detection, background restoration, and real-time performance, and is particularly applicable to complex scenes in high-dynamic environments.
[0060] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0061] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.
Claims
1. A dynamic obstacle recognition and elimination method based on SLAM research, characterized in that: The method includes: Constructing an improved Mask R-CNN network, acquiring an original image, and inputting the original image into the improved Mask R-CNN network; wherein a mask culling branch is added to the FCN layer of the improved Mask R-CNN network; According to the IMU data and depth data, the camera's self-motion parameters are calculated, and the predicted trajectory of the object in the original image is reversely compensated according to the camera's self-motion parameters; The original image is input into the improved Mask R-CNN network for feature extraction, and the mask is generated and updated in combination with the iterative Gaussian algorithm; Dynamic obstacles are removed according to the mask to obtain an effect image of dynamic obstacle removal.
2. The method for identifying and removing dynamic obstacles based on SLAM research according to claim 1 is characterized in that: The improved Mask R-CNN network includes: The backbone network ResNet50, FPN network, region projection network RPN, RoIAlign layer and full convolution network FCN are sequentially connected in series; among them, the full convolution network FCN includes mask elimination branch and Mask branch.
3. The dynamic obstacle recognition and elimination method based on SLAM research according to claim 2 is characterized in that: The loss function of the RoIAlign layer is: in, is the classification loss, is the bounding box, is the mask branch loss, is the mask culling branch loss.
4. The method for identifying and removing dynamic obstacles based on SLAM research according to claim 3 is characterized in that: The step of calculating the self-motion parameters of the camera according to the IMU data and the depth data, and performing reverse compensation on the predicted trajectory of the object in the original image according to the self-motion parameters of the camera includes: Obtain motion information from camera or sensor IMU data in real time, and correct the object's motion trajectory by combining the object's state estimation in depth data; the motion information includes the rotation matrix and translation vectors ; By continuously updating the motion information of the camera or sensor, the predicted position of the object is continuously adjusted to eliminate errors caused by the movement of the device itself.
5. The method for identifying and removing dynamic obstacles based on SLAM research according to claim 4 is characterized in that: The original image is input into the improved Mask R-CNN network for feature extraction, and the mask generation and update steps are combined with the iterative Gaussian algorithm, including: The original image is input into the backbone network ResNet50 and FPN network for feature extraction to obtain the feature map; The feature map is input into the region projection network (RPN) to determine the anchor box to obtain the candidate region. The candidate region and feature map are input into the RoIAlign layer, and the candidate region is mapped to the feature map for alignment using the bilinear interpolation method; Input the aligned feature map into the Mask branch in the fully convolutional network (FCN) to generate a binary mask; The dynamic background is reconstructed according to the binary mask, and the iterative Gaussian algorithm is introduced for optimization to obtain the mask of the dynamic obstacle.
6. The method for identifying and removing dynamic obstacles based on SLAM research according to claim 5 is characterized in that: The step of inputting the feature map into the region projection network RPN to determine the anchor box to obtain the candidate region includes: The feature map is input into the region projection network RPN to generate pixel anchor boxes of different sizes and aspect ratios; Use the Softmax method to classify each pixel anchor box and calculate the foreground probability and background probability of each pixel anchor box; Perform bounding box regression based on the foreground probability and background probability of each pixel anchor box to obtain the initial regression; The pixel anchor box is modified using the regression prediction value, the foreground anchor box results are sorted according to the foreground probability and background probability, and the pixel anchor boxes with values below the threshold or with high overlap are deleted using the non-maximum suppression algorithm to obtain the candidate region RoI; the candidate region contains dynamic obstacles or static environmental features.
7. The method for identifying and eliminating dynamic obstacles based on SLAM research according to claim 6 is characterized in that: The candidate region and the feature map are input into the RoIAlign layer, and the candidate region is mapped to the feature map using the bilinear interpolation method for alignment, including: The region projection network (RPN) generates a series of target regions through the anchor box mechanism and screens them based on foreground and background information. The input image where the candidate area is located is extracted through ResNet50 to obtain a high-dimensional feature map; The RoIAlign layer converts the coordinates of the candidate region in the original image into the feature map coordinate space and calculates the floating point coordinates of the candidate region; the candidate region is divided into grid units of fixed size; For each sampling point of the grid cell, the corresponding eigenvalue is calculated from the feature map using the bilinear interpolation method; Each candidate region is mapped to a fixed-size feature map region, ensuring that candidate regions of different scales have consistent spatial alignment accuracy in subsequent classification and segmentation tasks.
8. The method for identifying and removing dynamic obstacles based on SLAM research according to claim 7 is characterized in that: The aligned feature map is input into the Mask branch in the fully convolutional network FCN. The steps to generate a binary mask include: The feature map processed by the RoIAlign layer is input into the Mask branch network, and the deep features of the target area are extracted through the fully convolutional network FCN. Among them, the feature map output by the RoIAlign layer is upsampled; The fully convolutional network (FCN) extracts multi-scale features through convolutional layers of different depths, thereby ensuring that the spatial dimensions of input and output remain consistent and avoiding the loss of spatial information due to the reduction of feature map size; Calculate the object probability of each pixel and generate a binary mask; the segmentation of the binary mask is a two-classification task. The sigmoid activation function is applied to each pixel, and the output value is mapped to [0,1], indicating the probability that the pixel belongs to the object of this category.
9. The method for identifying and removing dynamic obstacles based on SLAM research according to claim 6 is characterized in that: The steps of reconstructing the dynamic background according to the binary mask and introducing the iterative Gaussian algorithm for optimization to obtain the mask of the dynamic obstacle include: Use the Mask R-CNN segmentation branch to generate pixel-level binary masks of dynamic objects; The fully convolutional network (FCN) is used to perform multi-layer convolution on the feature map of the candidate area to extract high-dimensional semantic information, and combined with the classification branch of Mask R-CNN to identify dynamic objects; Semantic segmentation is used to further verify the category of dynamic objects to improve the accuracy of elimination; Through pixel-level logical operations, a binary mask is applied to the original image to perform a region growing algorithm on the depth map to expand the dynamic object area and ensure the completeness of the culling; The background is modeled through the iterative Gaussian algorithm to distinguish dynamic objects from static backgrounds, making background reconstruction more accurate; Fill the removed area with background information from the historical frame; The optical flow method is used to predict the pixel motion trend of the previous and next frames to supplement the occluded areas.
Citation Information
Cited By
Automatic driving obstacle recognition method for complex scenic spot road scene
CN120913178A
Multi-degree-of-freedom mechanical arm obstacle avoidance path planning method based on three-dimensional reconstruction
CN121374660A
Lightweight dynamic feature point elimination method based on deep learning
CN121937726A