Dynamic removal-based mapping and path planning method and device, equipment, medium and product
By adopting dynamic removal-based mapping and path planning methods in micro quadruped robots, using deep learning and visual odometer technology to build a static environment map and perform path planning, the problem of insufficient perception and path planning capabilities of robots in dynamic environments is solved, and efficient and accurate autonomous navigation is achieved.
Patent Information
- Application Number
- CN202510324814.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-19
AI Technical Summary
The lack of perception and path planning capabilities of micro-four-legged robots in dynamic environments leads to limited autonomous navigation capabilities in complex unstructured environments.
The mapping and path planning method based on dynamic removal is adopted. By obtaining the RGB image of the target environment, the deep learning algorithm is used to identify the object categories and set the dynamic degree value, the depth image is obtained in combination with the Depth-Anything algorithm, the 3D map is constructed through visual odometer, and dynamic points are eliminated to generate a global static environment map. Finally, the path planning algorithm is used to obtain the optimal planning path of the robot.
It realizes high-precision static environment map construction and optimal path planning in complex dynamic environments, and improves the robot's autonomous navigation ability and path planning efficiency in dynamic environments.
Smart Images

Figure CN119984282A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision and robot navigation, and in particular to a mapping and path planning method, device, equipment, medium and product based on dynamic removal. Background Art
[0002] With the rapid development of science and technology, micro quadruped robots have shown great application potential in the fields of environmental detection, unmanned reconnaissance and narrow space operations. These robots have shown unique advantages in complex environments with their flexible movement capabilities and unique structural characteristics.
[0003] However, there are relatively few studies on environmental perception and path planning for micro quadruped robots at home and abroad. Existing research on micro quadruped robots focuses more on improving the stability of their own motion, but is still weak in motion planning capabilities based on the environment, making it difficult to complete a complete exploration task through autonomous path planning. Traditional control methods usually rely heavily on the body's dynamics modeling, resulting in low system robustness and difficulty in adapting to high-dynamic motion requirements in unknown and complex environments. In addition, SLAM (Simultaneous Localization and Mapping) in a small space still faces many challenges. Most existing work assumes that the environment is static and ignores the interference caused by dynamic objects, which significantly affects the positioning accuracy in real environments, especially when there are moving objects, and even causes system failure.
[0004] In the existing technology, there are obvious defects in the perception and path planning capabilities of micro-quadruped robots in dynamic environments. The existence of this problem limits the autonomous navigation capabilities of micro-quadruped robots in complex unstructured environments and has become a key issue that needs to be urgently addressed. Summary of the invention
[0005] The purpose of this application is to provide a mapping and path planning method, device, equipment, medium and product based on dynamic removal, which can remove dynamic objects, build accurate static environment maps, and achieve optimal path planning in complex unstructured environments.
[0006] To achieve the above objectives, this application provides the following solutions:
[0007] In a first aspect, the present application provides a mapping and path planning method based on dynamic removal, comprising:
[0008] Obtain an RGB image of the target environment; the target environment is the environment where the robot is currently located;
[0009] Based on the RGB image, a deep learning algorithm is used to determine the category of each object in the RGB image; and the dynamic degree value of the corresponding pixel point in the RGB image is set according to the category of each object to obtain the category-dynamic degree labeled image at the current moment; the categories of the objects include humans, animals, vehicles and daily items;
[0010] Based on the RGB image, the Depth-Anything algorithm is used to obtain the depth image;
[0011] Based on RGB images, category-dynamic degree labeled images and depth images, each pixel is converted into point cloud data in three-dimensional space through visual odometry method to build a 3D map;
[0012] Based on the dynamic degree value of each point cloud data in the 3D map, the point cloud data of the dynamic points in the 3D map is determined and eliminated to obtain a global static environment map;
[0013] Based on the global static environment map, the path planning algorithm is used to obtain the optimal planning path of the robot in the target environment.
[0014] In a second aspect, the present application provides a mapping and path planning device based on dynamic removal, comprising:
[0015] The RGB image acquisition module is used to acquire the RGB image of the target environment; the target environment is the environment where the robot is located at the current moment;
[0016] The object category recognition and dynamic degree marking module is used to determine the category of each object in the RGB image using a deep learning algorithm based on the RGB image; and set the dynamic degree value of the corresponding pixel point in the RGB image according to the category of each object to obtain the category-dynamic degree marking image at the current moment; the categories of the objects include humans, animals, vehicles and daily items;
[0017] The depth image acquisition module is used to obtain the depth image based on the RGB image using the Depth-Anything algorithm;
[0018] 3D map construction module, which is used to convert each pixel into point cloud data in three-dimensional space through visual odometry method based on RGB images, category-dynamic degree labeled images and depth images to build 3D maps;
[0019] A dynamic point elimination module is used to determine and eliminate the point cloud data of dynamic points in the 3D map based on the dynamic degree value of each point cloud data in the 3D map, so as to obtain a global static environment map;
[0020] The path planning module is used to obtain the optimal planned path of the robot in the target environment based on the global static environment map and using the path planning algorithm.
[0021] In a third aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any of the above-mentioned methods of mapping and path planning based on dynamic removal.
[0022] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-mentioned methods for mapping and path planning based on dynamic removal.
[0023] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of any of the above-mentioned dynamic removal-based mapping and path planning methods.
[0024] According to the specific embodiments provided in this application, this application has the following technical effects:
[0025] The present application provides a mapping and path planning method, device, equipment, medium and product based on dynamic removal. By acquiring the RGB image of the target environment and using a deep learning algorithm to determine the category of each object in the image, the robot can accurately identify different objects in the surrounding environment, laying the foundation for subsequent dynamic degree evaluation. At the same time, the dynamic degree value of the corresponding pixel point is set according to the object category to obtain a category-dynamic degree marked image, which realizes the preliminary distinction of dynamic elements in the environment; the Depth-Anything algorithm is used to obtain a depth image, and the RGB image, the category-dynamic degree marked image and the depth image are combined to construct a 3D map through the visual odometer method, solving the problem of mapping from a two-dimensional image to a three-dimensional space. It not only provides the three-dimensional structural information of the environment, but also retains the dynamic degree information of the object, making it possible to remove the dynamic points in the future. By determining and removing the point cloud data of dynamic points based on the dynamic degree value of each point cloud data in the 3D map, the problem of the influence of dynamic elements on the map accuracy is solved, ensuring that the final generated global static environment map only contains static elements, thereby improving the accuracy and reliability of the map. Based on the global static environment map, the path planning algorithm is used to obtain the optimal planning path of the robot in the target environment, achieving the efficiency and accuracy of robot navigation. The information of the static map is fully utilized to provide the robot with a clear and barrier-free navigation path. It effectively solves the challenges brought by the dynamic environment to the robot's mapping and path planning, and realizes the construction of a high-precision global static environment map and efficient path planning. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0027] Figure 1 A schematic diagram of a process flow of a mapping and path planning method based on dynamic removal provided in one embodiment of the present application;
[0028] Figure 2 A schematic diagram of functional modules of a mapping and path planning device based on dynamic removal provided in one embodiment of the present application;
[0029] Figure 3 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0030] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0031] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0032] In an exemplary embodiment, Figure 1 As shown, a mapping and path planning method based on dynamic removal is provided, including the following steps 101 to 106. Among them:
[0033] Step 101, obtaining an RGB image of a target environment; the target environment is the environment where the robot is located at the current moment.
[0034] Step 102, based on the RGB image, using a deep learning algorithm, determine the category of each object in the RGB image; and set the dynamic degree value of the corresponding pixel point in the RGB image according to the category of each object to obtain the category-dynamic degree labeled image at the current moment; the categories of the objects include humans, animals, vehicles and daily items.
[0035] Step 103: Based on the RGB image, a Depth-Anything algorithm is used to obtain a depth image.
[0036] Step 104, based on the RGB image, the category-dynamic degree labeled image and the depth image, each pixel is converted into point cloud data in three-dimensional space by a visual odometry method to construct a 3D map.
[0037] Step 105 , based on the dynamic degree value of each point cloud data in the 3D map, determine and remove the point cloud data of dynamic points in the 3D map to obtain a global static environment map.
[0038] Step 106: Based on the global static environment map, a path planning algorithm is used to obtain the optimal planned path of the robot in the target environment.
[0039] By implementing the above steps 101 to 106, the present application can significantly improve the accuracy of environmental modeling and the efficiency of path planning, especially in dynamic environments. Through the organic combination of deep learning and visual odometer, it can perceive environmental changes in real time, intelligently remove dynamic interference, and provide a clear and accurate environmental model for the robot, thereby supporting it to perform efficient and reliable navigation and path planning.
[0040] In another exemplary embodiment of the present application, step 102 specifically includes:
[0041] The RGB image is input into the target detection and segmentation model, and the object category detection and instance segmentation are performed on the RGB image at the current moment to obtain the category of the object in the RGB image; the target detection and segmentation model is obtained by iteratively training the YOLOv8 network based on the sample training set; the sample training set includes historical RGB images with the category of the object marked.
[0042] As an optional implementation method, the object detection and segmentation model training process based on the YOLOv8 network can be divided into the following specific steps:
[0043] 1. Data preparation: Prepare a labeled sample training set, usually including images and labels. For object detection, the label includes the category and bounding box of each object. For segmentation tasks, the label also needs to include the segmentation mask of each object, usually a binary image or contour information. The dataset can use standard datasets such as COCO (Common Objects in Context), VOC (Visual Object Classes), or custom datasets. This application uses the COCO dataset.
[0044] 2. Environment configuration: Install the environment that YOLOv8 depends on. First, you need to install Python. In this application, python 3.8 is selected, and then install PyTorch and ultralytics libraries.
[0045] 3. Model selection and initialization: The YOLOv8 open source code provides pre-trained models of different sizes. You can choose the appropriate model based on the computing resources of the task. In this application, YOLOv8s is selected. Use the following code to load the pre-trained model:
[0046] a)fromultralytics importYOLO.
[0047] b) model=YOLO('yolov8s.pt').
[0048] 4. Data enhancement and loading: YOLOv8 automatically handles data loading and enhancement through the ultralytics library. Common data enhancements include random cropping, scaling, flipping, color adjustment, etc.
[0049] 5. Training: Start model training using the following command:
[0050] a)yolo task=detect mode=train model=yolov8n.pt data=your_dataset.yaml epochs=50imgsz=640.
[0051] Among them, task = detect indicates the target detection task, mode = train indicates the training mode, model = yolov8s.pt specifies the use of the pre-trained YOLOv8 model, data = your_dataset.yaml is the dataset configuration file, epochs = 50 indicates 50 training cycles, and imgsz = 640 indicates the size of the input image. During the training process, the model will automatically evaluate on the validation set and output accuracy indicators such as MAP (Mean Average Precision), Recall, Precision, etc.
[0052] 6. Evaluation and verification: During the training process, YOLOv8 will regularly evaluate the performance of the model on the validation set.
[0053] 7. Model saving and reasoning: After training is completed, YOLOv8 will try the code to save the best model and save the model as a .pt format file for subsequent reasoning:
[0054] a)model.save('best_model.pt').
[0055] At inference time, use the following code to load the saved model and make predictions:
[0056] b) results=model.predict('path_to_image_or_video').
[0057] 8. Tuning and optimization: During the training process, you can adjust the model's hyperparameters (such as learning rate, batch size, number of training rounds, etc.) based on the training results and validation set performance. In addition, for segmentation tasks, you can also adjust the resolution of the input image, the enhancement strategy, or choose a larger model to further improve performance.
[0058] Through the above steps, YOLOv8 can complete target detection and instance segmentation tasks, and perform subsequent tuning and optimization according to needs.
[0059] As an optional implementation, the dynamic degree value is specifically expressed as a number ranging from 0 to 100. The larger the value, the higher the dynamic degree. The 80 categories in the COCO dataset cover a wide range of objects, including humans, animals, vehicles, daily objects, etc. According to the dynamic degree of these categories in real life (i.e., whether the objects are often in motion), they can be scored, with scores ranging from 0 (almost motionless) to 100 (frequently moving). As shown in Table 1 below, the scores are based on the dynamic degree of these categories.
[0060] Table 1. COCO dataset category dynamics rating table
[0061] Initial Category Initial dynamic value Initial Category Initial dynamic value Initial Category Initial dynamic value people 100 bike 90 car 90 motorcycle 85 airplane 95 bus 85 train 85 truck 85 Boat 70 Traffic lights 10 fire hydrant 5 Stop Sign 5 Parking meter 5 bench 0 bird 80 cat 70 dog 80 horse 80 sheep 60 dairy cow 50 elephant 40 Bear 60 zebra 60 giraffe 40 Backpack 10 Umbrella 5 handbag 10 tie 0 suitcase 10 Frisbee 70 sled 40 Snowboard 40 Sports Balls 80 Kite 90 bat 20 Baseball Gloves 10 skateboard 80 surfboard 80 Tennis racket 20 bottle 5 Glass 5 Measuring cup 5 fork 0 knife 0 spoon 0 bowl 5 banana 0 apple 0 sandwich 0 orange 0 broccoli 0 carrot 0 hot dog 0 Pizza 0 Donut 0 cake 0 Chair 0 sofa 0 Potted plants 5 bed 0 dining table 0 Flush toilet 0 TV set 0 Laptop 0 mouse 0 Remote Control 0 keyboard 0 cell phone 0 Micro-wave oven 0 oven 0 Toaster 0 Sink 0 refrigerator 0 Book 0 Clocks 0 vase 0 Scissors 0 Teddy bear 0 Hair dryer 0 toothbrush 0
[0062] In another exemplary embodiment of the present application, step 103 specifically includes:
[0063] Use the Depth-Anything depth estimation algorithm to estimate the depth of the original environment image at the current moment and obtain the depth information corresponding to each pixel. The specific steps are as follows:
[0064] 1. Prepare input data (RGB image of the target environment)
[0065] First, you need to prepare an RGB image of the target environment, usually a normal RGB image. This image will be passed as input to the depth estimation algorithm.
[0066] 2. Load and preprocess images
[0067] Before inputting into the Depth-Anything model, the image usually needs to be preprocessed, including: adjusting the image size to 640×480 to meet the input requirements of the Depth-Anything model. Normalizing the image, usually scaling the pixel values to the range [0, 1], or subtracting the mean and dividing by the standard deviation, to make the image consistent with the input format when training the Depth-Anything model.
[0068] 3. Load the pre-trained Depth-Anything model
[0069] The Depth-Anything algorithm usually relies on a pre-trained deep neural network that has been trained on a large number of annotated images. In order to obtain a depth estimate, you need to load the pre-trained model using the following code.
[0070] from depth_anything importDepthAnything.
[0071] model=DepthAnything.from_pretrained("model_checkpoint_or_url").
[0072] 4. Image Reasoning (Depth Estimation)
[0073] The following code implements the inference of the RGB image of the target environment using the loaded pre-trained model. The depth estimation algorithm predicts the depth information of each pixel in the image through the forward propagation of the network. Usually, the depth information is expressed in the form of floating points, representing the distance between the pixel and the camera (usually in meters or centimeters).
[0074] depth_map=model.predict(left_image).
[0075] Among them, left_image is the RGB image of the target environment, and depth_map is the depth image output by the model, in which each pixel value represents the depth of the corresponding pixel in the image (that is, the distance from the camera).
[0076] In another exemplary embodiment of the present application, step 104 specifically includes:
[0077] 2D feature points are extracted from the RGB image using an image processing algorithm, and matched with the corresponding 2D feature points at the previous moment to obtain the camera's motion information; the 2D feature points include intersection points and edge points; the motion information includes translation parameters and rotation parameters;
[0078] Determine the camera pose information at the current moment according to the camera motion information and IMU data; the IMU data includes accelerometer measurement values and gyroscope measurement values;
[0079] Based on the category-dynamic degree labeled image, depth image and camera pose information, each pixel in the depth image is back-projected into the three-dimensional space to obtain the corresponding point cloud data and construct a 3D map.
[0080] The VIO (Visual-IMU Odometry) system combines visual features with inertial measurement unit (IMU) data information to estimate the current camera position. Combined with depth information, 2D image feature points are accurately converted into 3D point cloud data to obtain a 3D map.
[0081] In another exemplary embodiment of the present application, step 105 specifically includes:
[0082] Calculate the modulus of each point cloud data in the 3D map.
[0083] Based on the modulus length of each point cloud data in the 3D map, the mean modulus length of all point cloud data of the same category in the 3D map is calculated, and the difference is made with the mean modulus length of all point cloud data of the corresponding category in the 3D map at the previous moment to obtain the difference of the mean modulus length of each category.
[0084] The dynamic degree value of the same category is corrected according to the difference of the mean value of the module length of each category, and the dynamic degree value of each category at the current moment is obtained.
[0085] The point cloud data whose current dynamic value is greater than the preset dynamic threshold is marked as a dynamic point and removed.
[0086] Based on the point cloud data after dynamic points are eliminated in the 3D map, the ORB-SLAM3 algorithm is used to construct and optimize the map to obtain a global static environment map that reflects the static structure of the target environment.
[0087] As an optional implementation, the depth image at the current moment is obtained using the Depth-Anything depth estimation algorithm described above. Assume D(i t , j t ) is the depth image at time t (i t , j t ) pixel depth value, indicating the straight-line distance from the point to the camera. The camera's intrinsic parameter matrix K (usually the camera's focal length and optical center) is known.
[0088] 1. Calculate the 3D coordinates of the pixel
[0089] First, the 3D spatial coordinates of each pixel need to be calculated from the depth image at the current moment. Assuming that the camera intrinsic parameter matrix K is known, it can usually be expressed as:
[0090]
[0091] Where: f x , f y is the focal length of the camera (in pixels); c x , c y is the principal point coordinate of the camera (unit: pixel)
[0092] For each pixel (i t , j t ), its corresponding three-dimensional coordinates can be calculated by the following formula:
[0093]
[0094] Z=D(i t ,j t )
[0095] Where X, Y, and Z are the three-dimensional spatial coordinates of the pixel point in the image at time t; D(i t , j t ) is the depth value of the pixel in the depth map at the current moment.
[0096] 2. Calculate the module length
[0097] To calculate the modulus of each pixel in the depth map at the current moment (i.e., the distance from the point to the camera origin), the modulus can be calculated using the above three-dimensional coordinates. t , j t )The corresponding three-dimensional coordinate vector is (X, Y, Z), and its modulus can be calculated using the Euclidean distance formula.
[0098] As an optional implementation, the calculation formula for the modulus of each point cloud data in the 3D map is:
[0099]
[0100] Among them, ||v(i t ,j t )|| represents the modulus of the point cloud data in the i-th row and j-th column of the 3D map at time t; fx and f y Respectively represent the horizontal focal length and vertical focal length of the camera, c x and c y Respectively represent the horizontal and vertical coordinates of the camera's principal point; D(i t , j t) represents the depth information of the point cloud data in the i-th row and j-th column in the 3D map at time t.
[0101] As an optional implementation, a global static environment map is generated by the ORB-SLAM3 system, which generates an environment map and performs real-time positioning. This process involves the following main steps:
[0102] 1. Camera initialization and feature extraction.
[0103] Feature extraction (ORB features): ORB-SLAM3 first extracts ORB features (OrientedFAST and RotatedBRIEF) from the image input by the camera. These feature points are highly invariant in the image and can be effectively matched under different viewing angles and lighting conditions.
[0104] Initialization: The system will initialize the camera's position and orientation by matching features in consecutive image frames and combining camera motion information to build a preliminary environment map. If it is a binocular or RGB-D camera, the system can accelerate the initialization process through stereo matching or depth information.
[0105] 2. Keyframes and Map Points.
[0106] Keyframe selection: ORB-SLAM3 automatically selects some representative image frames as keyframes based on the camera's motion trajectory and feature matching. Keyframes are the core of ORB-SLAM3 and contain significant information relative to other frames.
[0107] Map points: Whenever the camera detects enough feature points and they have a stable matching relationship between multiple frames, ORB-SLAM3 will add these feature points to the map. Each map point represents a physical location in space.
[0108] 3. Pose estimation and positioning.
[0109] Pose Estimation: ORB-SLAM3 estimates the relative pose of the camera, that is, the position and direction of the camera in space, by matching features between consecutive frames. This process is optimized by bundle adjustment.
[0110] Positioning: The system reduces the reprojection error by optimizing the camera’s position and posture, so that the feature points in the image match the corresponding points in the map as closely as possible, thereby achieving accurate positioning.
[0111] 4. Map optimization (map point and keyframe optimization).
[0112] Local optimization: ORB-SLAM3 uses a local map optimization method to continuously optimize the map by adding new keyframes and map points in each frame. Local optimization is mainly performed through bundle adjustment, which can reduce reprojection errors and improve positioning accuracy.
[0113] Global Optimization: ORB-SLAM3 also supports global optimization, which is to further optimize the entire map through global loop closure detection (Loop Closing) to eliminate drift caused by accumulated errors. Loop closure detection can compare the system's trajectory and map with the previous trajectory, correct errors and optimize the entire map.
[0114] 5. Loop Closing
[0115] Closed loop detection: In the environment, the camera may return to the area it has passed before. At this time, ORB-SLAM3 will detect this closed loop and align the current map with the previous map through the closed loop detection algorithm. This step is crucial to reduce positioning drift and improve map accuracy.
[0116] Closed-loop optimization: After the closed-loop detection is successful, the system will adjust the map and posture through global optimization to correct the previous accumulated errors.
[0117] 6. Generate environment map
[0118] Finally, through the above steps, ORB-SLAM3 will generate an accurate environment map. This map consists of the following parts:
[0119] 3D point cloud map: Map points form a dense three-dimensional point cloud. Each point represents a feature point in space. The positions of these points are usually obtained through triangulation and other methods.
[0120] Camera trajectory: By estimating the camera's pose, ORB-SLAM3 generates the camera's motion trajectory, that is, the path of the camera in the environment.
[0121] Optimized map: After closed-loop optimization, the map generated by ORB-SLAM3 is globally optimized, with higher accuracy and reduced drift and error.
[0122] 7. Export the map.
[0123] The generated environment map can be exported and visualized in a variety of forms:
[0124] 3D point cloud visualization: Point cloud visualization tools (such as PCL, RViz, etc.) can be used to display the environment map, which is usually composed of 3D map points and camera trajectories.
[0125] Map saving: ORB-SLAM3 can save the generated map in a specific format (such as PLY, PCD, etc.) for subsequent use or further processing.
[0126] Path planning and navigation: The generated map can be directly used for the robot’s path planning and environmental perception, helping the robot understand and navigate its environment.
[0127] 8. Dynamic environment processing
[0128] Dynamic object detection: In a dynamic environment, ORB-SLAM3 may encounter moving objects. In order to maintain the accuracy of the map, ORB-SLAM3 will use dynamic object detection and separation technology to avoid mistaking dynamic objects for part of the environment map.
[0129] Update and Optimize: ORB-SLAM3 will dynamically update the map to handle new objects or structures that appear in the scene and keep the map real-time.
[0130] In summary, the process of ORB-SLAM3 generating an environmental map can be summarized as follows: first, the camera is positioned through ORB feature extraction and key frame selection, and then a preliminary map is constructed through feature matching and pose estimation of consecutive frames. In the continuous optimization process of key frames and map points, the map accuracy is improved through local and global optimization, and the map error is further corrected by closed-loop detection. Finally, ORB-SLAM3 generates an accurate environmental map, usually presented in the form of a three-dimensional point cloud, which can be used for subsequent navigation, positioning, and path planning tasks.
[0131] In another exemplary embodiment of the present application, the path planning algorithm in step 106 includes: a Fast-Planner path planning algorithm.
[0132] As an optional implementation, the optimal path is obtained through the Fast Planner algorithm. FastPlanner is an algorithm for robot path planning, especially suitable for real-time path planning in complex dynamic environments. Fast Planner is an optimization-based method that calculates an optimal path from the starting point to the target by considering factors such as robot kinematic constraints, environmental obstacles, and target position. This path usually takes into account factors such as the shortest path, collision avoidance, and speed / acceleration constraints.
[0133] The basic process of FastPlanner to obtain the optimal path:
[0134] 1. Environment modeling and state space definition
[0135] Environment modeling: First, FastPlanner obtains the environment information of the robot, that is, the global static map obtained by ORB-SLAM3. This map is usually a grid map, indicating whether each spatial position is occupied by an obstacle.
[0136] State Space: The space where the robot is located is discretized into a high-dimensional state space (usually 2D or 3D). Each state includes the robot's position, velocity, acceleration and other kinematic information.
[0137] 2. Determination of the starting state and target state
[0138] Starting point: The current position (position and orientation) of the robot is used as the starting point for path planning.
[0139] Target: The target point can be a specific location.
[0140] 3. Cost function design
[0141] Fast Planner uses a cost function to evaluate the quality of each path. The cost function includes the following aspects:
[0142] Path length: The total length or total distance of a path, usually the goal is to find the shortest path.
[0143] Collision cost: collision detection between the path and obstacles to prevent the robot path from passing through the obstacle area.
[0144] Speed and acceleration constraints: The robot is subject to physical limitations such as speed and acceleration during movement. Path planning needs to take these dynamic constraints into account to avoid overspeeding or sudden acceleration changes.
[0145] Dynamic obstacle avoidance: In a dynamic environment, path planning must not only avoid static obstacles, but also consider the dynamic avoidance of other moving objects.
[0146] These costs constitute a comprehensive objective function, and FastPlanner will try to minimize this cost to obtain the optimal path.
[0147] 4. Path search and planning methods
[0148] Fast Planner uses a sampling-based optimization path search method, among which the common search methods are:
[0149] Sampling method: Use different path search algorithms to sample the state space and evaluate these sampling points based on the cost function.
[0150] Optimization method: After searching for a set of feasible paths, Fast Planner will further reduce the tortuosity of the path through path smoothing and optimization, optimize the smoothness of the path, and ensure the feasibility of the path under the robot's kinematic constraints. Commonly used optimization algorithms include gradient descent method, dynamic programming to minimize path cost, etc.
[0151] During this process, Fast Planner generates paths in an incremental manner, combines environmental information and cost functions, and gradually adjusts the path until an optimal path is found.
[0152] In summary, this application obtains environmental information through sensor data and constructs an obstacle map; converts the robot and its motion constraints into state space; designs a cost function by considering factors such as path length, collision cost, dynamic obstacles, speed and acceleration constraints; searches for paths in the state space through sampling and optimization algorithms, such as RRT* (Rapidly-exploring Random TreesStar), to obtain the optimal planned path.
[0153] The present application also provides an application scenario, which applies the above-mentioned mapping and path planning method based on dynamic removal. Specifically: The mapping and path planning method based on dynamic removal provided in this embodiment can be applied to search and rescue operation scenarios in complex ruins environments. This scenario is divided into three core links: environmental perception and dynamic map construction, intelligent path planning and optimization, and search and rescue mission execution and feedback. In the environmental perception and dynamic map construction link, the micro-quadruped robot makes full use of its visual sensors, inertial measurement units (IMUs) and other multi-sensing devices to comprehensively collect various types of data in the ruins environment. By integrating visual and IMU information, the visual limitations in narrow spaces are effectively overcome. At the same time, with the help of target detection and depth estimation technology, the robot can accurately identify and remove dynamic obstacles, thereby constructing a high-precision static environment map. In the intelligent path planning and optimization link, based on the constructed static map, the robot uses advanced path planning algorithms to quickly plan an optimal or near-optimal search and rescue path. In this process, the robot will monitor environmental changes in real time and dynamically adjust the path to avoid new obstacles. In addition, by combining multi-dimensional environmental information and high-precision positioning technology, the robot can ensure accurate navigation in complex ruins. Finally, in the execution and feedback of the search and rescue mission, the micro-quadruped robot will autonomously navigate to the target area according to the planned path to perform specific tasks such as searching and rescuing trapped people. In the process of performing the mission, the robot will collect mission execution data in real time, such as search and rescue progress, environmental changes, etc., and feed this data back to the control center. These feedback data not only help the control center understand the progress of the search and rescue, but also provide important reference for subsequent path optimization and mission adjustment.
[0154] The mapping and path planning method based on dynamic removal provided in this embodiment belongs to the intelligent path planning and optimization link in the search and rescue operation process in a complex ruins environment. This method not only significantly improves the robot's autonomous navigation ability in a complex environment, but also greatly enhances its ability to cope with dynamic challenges, thereby ensuring the smooth execution of search and rescue missions. Through the application of this method, the micro-quadruped robot has shown great potential and value in high-risk and difficult tasks such as ruins search and rescue.
[0155] This application has the following technical effects:
[0156] 1. Compared with the traditional SLAM method in dynamic environment, this application adopts Depth-Anything algorithm for depth estimation, which requires lower computing power than binocular triangulation method and has light load requirements on the robot's computing platform.
[0157] 2. Use the Faster-Planner algorithm to perform path planning and navigation based on the reconstructed environment map, effectively implement optimal path planning, and generate motion control instructions to achieve autonomous motion control.
[0158] 3. This application fully considers the interference of objects in dynamic environments. Through dynamic object proposal and high-precision map reconstruction technology, it ensures the robustness and stability of the system in unstructured dynamic environments, and improves the adaptability of the robot in practical applications.
[0159] Based on the same inventive concept, the embodiment of the present application also provides a dynamic removal-based mapping and path planning device for implementing the above-mentioned dynamic removal-based mapping and path planning method. The implementation solution provided by the device to solve the problem is similar to the implementation solution recorded in the above-mentioned method, so the specific limitations in one or more embodiments of the dynamic removal-based mapping and path planning device provided below can refer to the limitations of the dynamic removal-based mapping and path planning method above, and will not be repeated here.
[0160] In an exemplary embodiment, Figure 2 As shown, a mapping and path planning device based on dynamic removal is provided, comprising:
[0161] The RGB image acquisition module 201 is used to acquire an RGB image of a target environment; the target environment is the environment where the robot is located at the current moment.
[0162] The object category recognition and dynamic degree marking module 202 is used to determine the category of each object in the RGB image based on the RGB image and adopt a deep learning algorithm; and set the dynamic degree value of the corresponding pixel point in the RGB image according to the category of each object to obtain the category-dynamic degree marking image at the current moment; the categories of the objects include humans, animals, vehicles and daily items.
[0163] The depth image acquisition module 203 is used to obtain a depth image based on the RGB image by using the Depth-Anything algorithm.
[0164] The 3D map construction module 204 is used to convert each pixel into point cloud data in three-dimensional space through a visual odometer method based on the RGB image, the category-dynamic degree labeled image and the depth image to construct a 3D map.
[0165] The dynamic point elimination module 205 is used to determine and eliminate the point cloud data of dynamic points in the 3D map based on the dynamic degree value of each point cloud data in the 3D map, so as to obtain a global static environment map.
[0166] The path planning module 206 is used to obtain the optimal planned path of the robot in the target environment based on the global static environment map using a path planning algorithm.
[0167] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 3 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a mapping and path planning method based on dynamic removal is implemented.
[0168] Those skilled in the art will understand that Figure 3The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components. In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0169] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0170] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0171] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0172] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0173] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.
[0174] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0175] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A mapping and path planning method based on dynamic removal, characterized in that: The mapping and navigation method based on dynamic removal includes: Obtain an RGB image of the target environment; the target environment is the environment where the robot is currently located; Based on the RGB image, a deep learning algorithm is used to determine the category of each object in the RGB image; and the dynamic degree value of the corresponding pixel point in the RGB image is set according to the category of each object to obtain the category-dynamic degree labeled image at the current moment; the categories of the objects include humans, animals, vehicles and daily items; Based on the RGB image, the Depth-Anything algorithm is used to obtain the depth image; Based on RGB images, category-dynamic degree labeled images and depth images, each pixel is converted into point cloud data in three-dimensional space through visual odometry method to build a 3D map; Based on the dynamic degree value of each point cloud data in the 3D map, the point cloud data of the dynamic points in the 3D map is determined and eliminated to obtain a global static environment map; Based on the global static environment map, the path planning algorithm is used to obtain the optimal planning path of the robot in the target environment.
2. The mapping and path planning method based on dynamic removal according to claim 1, characterized in that: Based on RGB images, a deep learning algorithm is used to determine the categories of each object in the RGB image, including: The RGB image is input into the target detection and segmentation model, and the object category detection and instance segmentation are performed on the RGB image at the current moment to obtain the category of the object in the RGB image; the target detection and segmentation model is obtained by iteratively training the YOLOv8 network based on the sample training set; the sample training set includes historical RGB images with the category of the object marked.
3. The mapping and path planning method based on dynamic removal according to claim 1, characterized in that: Based on RGB images, category-dynamic degree labeled images and depth images, each pixel is converted into point cloud data in three-dimensional space through the visual odometry method to build a 3D map, including: 2D feature points are extracted from the RGB image using an image processing algorithm, and matched with the corresponding 2D feature points at the previous moment to obtain the camera's motion information; the 2D feature points include intersection points and edge points; the motion information includes translation parameters and rotation parameters; Determine the camera pose information at the current moment according to the camera motion information and IMU data; the IMU data includes accelerometer measurement values and gyroscope measurement values; Based on the category-dynamic degree labeled image, depth image and camera pose information, each pixel in the depth image is back-projected into the three-dimensional space to obtain the corresponding point cloud data and construct a 3D map.
4. The mapping and path planning method based on dynamic removal according to claim 1, characterized in that: Based on the dynamic degree value of each point cloud data in the 3D map, the point cloud data of dynamic points in the 3D map is determined and eliminated to obtain a global static environment map, specifically including: Calculate the modulus of each point cloud data in the 3D map; Based on the modulus length of each point cloud data in the 3D map, the mean modulus length of all point cloud data of the same category in the 3D map is calculated, and the difference between the modulus length means of all point cloud data of the corresponding category in the 3D map at the previous moment is made, so as to obtain the difference of the mean modulus length of each category; The dynamic degree value of the same category is corrected according to the difference of the mean value of the modulus length of each category, and the dynamic degree value of each category at the current moment is obtained; The point cloud data whose current dynamic value is greater than the preset dynamic threshold is marked as dynamic points and removed; Based on the point cloud data after dynamic points are eliminated in the 3D map, the ORB-SLAM3 algorithm is used to construct and optimize the map to obtain a global static environment map that reflects the static structure of the target environment.
5. The mapping and path planning method based on dynamic removal according to claim 4, characterized in that: The calculation formula for the modulus of each point cloud data in the 3D map is: Among them, ||v(i t ,j t )|| represents the modulus of the point cloud data in the i-th row and j-th column of the 3D map at time t; fx and f y Respectively represent the horizontal focal length and vertical focal length of the camera, c x and c y Respectively represent the horizontal and vertical coordinates of the camera's principal point; D(i t , j t ) represents the depth information of the point cloud data in the i-th row and j-th column in the 3D map at time t.
6. The mapping and path planning method based on dynamic removal according to claim 1, characterized in that: Path planning algorithms include: Fast-Planner path planning algorithm.
7. A mapping and path planning device based on dynamic removal, characterized in that: The mapping and path planning device based on dynamic removal applies the mapping and path planning method based on dynamic removal according to any one of claims 1 to 6, and the mapping and path planning device based on dynamic removal comprises: The RGB image acquisition module is used to acquire the RGB image of the target environment; the target environment is the environment where the robot is located at the current moment; The object category recognition and dynamic degree marking module is used to determine the category of each object in the RGB image using a deep learning algorithm based on the RGB image; and set the dynamic degree value of the corresponding pixel point in the RGB image according to the category of each object to obtain the category-dynamic degree marking image at the current moment; the categories of the objects include humans, animals, vehicles and daily items; The depth image acquisition module is used to obtain the depth image based on the RGB image using the Depth-Anything algorithm; 3D map construction module, which is used to convert each pixel into point cloud data in three-dimensional space through visual odometry method based on RGB images, category-dynamic degree labeled images and depth images to build 3D maps; A dynamic point elimination module is used to determine and eliminate the point cloud data of dynamic points in the 3D map based on the dynamic degree value of each point cloud data in the 3D map, so as to obtain a global static environment map; The path planning module is used to obtain the optimal planned path of the robot in the target environment based on the global static environment map and using the path planning algorithm.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the mapping and path planning method based on dynamic removal as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the mapping and path planning method based on dynamic removal described in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the mapping and path planning method based on dynamic removal described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Semantic SLAM service robot navigation method in indoor dynamic environment
CN112013841A
Visual positioning and static map construction method and system in dynamic environment
CN112991447A
Dynamic environment positioning and mapping method based on binocular vision and related device
CN118053105A
Multi-sensor SLAM (Simultaneous Localization and Mapping) method based on dynamic feature point elimination and loopback detection
CN118225096A
YOLOv8-based point-line fusion visual SLAM (Simultaneous Localization and Mapping) method in indoor dynamic scene
CN119164383A