Positioning system and method for indoor mobile robot in dynamic scene
The dynamic target detection and feature filtering technology combined with the RGB-D camera and YOLOv3 algorithm solves the problem of dynamic non-rigid objects blocking the indoor mobile robot positioning system, improves positioning accuracy and operating efficiency, and meets real-time processing requirements.
Patent Information
- Application Number
- CN202510913886.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-10-17
AI Technical Summary
In indoor mobile robot application scenarios, the presence of dynamic non-rigid objects such as pedestrians and pets leads to the accumulation of feature matching errors in traditional visual SLAM systems, increasing the risk of system tracking loss and affecting positioning accuracy and stability.
An RGB-D camera combined with the YOLOv3 algorithm is used for dynamic target detection. Through the dynamic target detection module, target information screening and processing module, depth information fusion processing module and feature filtering module, a three-dimensional mask area of the dynamic target is generated, dynamic feature points are filtered, and static feature points are retained and input into the SLAM system to improve positioning accuracy and operation efficiency.
It effectively reduces the occlusion of dynamic targets by static scene features, reduces the cumulative amount of feature matching errors, improves the utilization rate of static features, maintains stable tracking and the processing speed of dynamic targets, improves positioning accuracy, and reduces the risk of system tracking loss.
Smart Images

Figure CN120807865A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of machine vision and robot positioning, and more particularly, to an indoor mobile robot positioning system in a dynamic scene and a method thereof. BACKGROUND
[0002] In a complex static scene handled by an indoor mobile robot, there are often high dynamic non-rigid objects such as people and pets. In computer vision, dynamic non-rigid targets appearing at a medium distance will cause large-area unstructured occlusion of the static scene in the image. For the key problem of simultaneous localization and mapping in the field of mobile robot state estimation, the large occlusion and random variation of the surrounding environment features will greatly reduce the estimation capability. For different scenes, a robust robot pose estimation system is essential.
[0003] At present, in the application scene of an indoor mobile robot, the existence of dynamic non-rigid objects (such as pedestrians and pets) will cause significant interference to a traditional visual SLAM system. These dynamic targets will cause large-area occlusion of static scene features, cause feature matching error accumulation, and further increase the risk of system tracking loss.
[0004] Therefore, in view of the above, the existing structure is studied and improved, and an indoor mobile robot positioning system in a dynamic scene and a method thereof are provided, so as to achieve a more practical purpose. SUMMARY
[0005] 1. Technical problem to be solved
[0006] In view of the problems in the prior art, the purpose of the present application is to provide an indoor mobile robot positioning system in a dynamic scene and a method thereof, which can effectively reduce the occlusion of static scene features by dynamic targets, thereby reducing the accumulation of feature matching errors and the risk of system tracking loss, and can improve the utilization rate of static features, maintain stable tracking, and improve the processing speed of dynamic targets.
[0007] 2. Technical solution
[0008] To solve the above problems, the present application adopts the following technical solution.
[0009] An indoor mobile robot positioning system in a dynamic scene, the positioning system comprising a dynamic target detection module, a target information screening processing module, a depth information fusion processing module, and a feature filtering module;
[0010] The dynamic target detection module synchronously collects color images and depth information using an RGB-D camera, and then uses YOLOv3 to detect dynamic targets and output detection frame information.
[0011] a target information screening processing module, which deletes detection boxes with low confidence using non-maximum suppression, retains the target detection box with the highest confidence, and obtains screening data;
[0012] a depth information fusion processing module, which calculates the depth values of the center point coordinates and the surrounding area of the detection box, takes the minimum value of the region depth as the target reference depth, and sets a depth threshold interval;
[0013] a feature filtering module, which generates a three-dimensional mask region of the dynamic target according to the depth information fusion processing result, filters the feature points in the region, retains the static features and inputs them into the SLAM system, and obtains the positioning information of the indoor mobile robot in the dynamic scene.
[0014] Further, the dynamic target detection module receives a picture sequence captured by an RGB-D camera in a dynamic environment, matches the color image and the depth picture through a time label, and then the DD-SLAM uses the target detection algorithm YOLOv3 to identify the dynamic target in the color picture and returns the position coordinates of the identification box.
[0015] Further, in the target information screening processing module, the detection box with the highest confidence is retained by scanning all output results of the neural network, and the class label of the detection box is specified as the class with the highest score in the box, and the coordinates and width and height values of the non-rigid dynamic object detection box are recorded.
[0016] Further, in the depth information fusion processing module, the feature points in the dynamic region are removed, and the valuable static scene feature information is retained to the greatest extent.
[0017] When calculating the depth values of the surrounding area, the surrounding area of 5x5 pixels is taken as the range for calculation.
[0018] When setting the depth threshold interval, the threshold interval range is ±8, and it corresponds to 0-255 quantization values.
[0019] Further, in the feature filtering module, the target detection static information is adjusted to be sufficient by combining the target detection algorithm based on deep learning and the visual SLAM system based on geometric depth information, the mismatch data generated by the visual SLAM system due to the dynamic target is reduced, and the positioning accuracy and running efficiency of the mobile robot are improved.
[0020] A positioning method for an indoor mobile robot in a dynamic scene, the specific implementation steps of the positioning method are as follows:
[0021] Step 1, scan and collect information about dynamic targets: collect image information of dynamic targets through an RGB-D camera, and then use YOLOv3 to detect dynamic targets and output dynamic target information detection boxes;
[0022] Step two, processing of target information: first, the target information is screened to obtain the highest confidence target data, and then based on depth information fusion, the depth threshold interval is obtained, and the unstructured mask of the dynamic non-rigid object is generated;
[0023] Step three, feature filtering and dynamic target positioning: combining the target detection algorithm based on deep learning and the visual SLAM system of geometric depth information, the sufficient static features of the dynamic target are retained, the dynamic form of the object is judged, the best unstructured mask is obtained, and the positioning information of the indoor dynamic target is displayed.
[0024] Further, in step one, the dynamic target detection is performed by the target detection frame center point method, specifically:
[0025] The indoor mobile robot positioning system first receives the picture sequence from the RGB-D camera in the dynamic environment, and matches the color image and the depth picture through the time label;
[0026] Then, the DD-SLAM uses the target detection algorithm YOLOv3 to identify the dynamic target of the color picture, and returns the position coordinates of the identification frame.
[0027] Further, in step two, when the target information is screened, the depth information obtained by the depth camera is used to propose a strategy to remove the feature points in the dynamic area, which maximizes the retention of valuable static scene feature information, making the SLAM system have high accuracy;
[0028] When performing depth fine system fusion processing, the highest confidence target detection frame is retained by scanning all output results of the neural network, and it specifies the class label of the detection frame as the highest scoring class in the frame, while recording the coordinates and width and height values of the non-rigid dynamic object detection frame. After calculating the center point of the target detection frame, the method obtains the geometric depth information of the center point on the corresponding depth image. At this time, the data obtained is the approximate depth information of the dynamic non-rigid object from the camera.
[0029] Further, in step three, when performing feature filtering, the extension of the object in depth is considered, and a fixed depth threshold interval is set for the specific dynamic human object to realize the unstructured mask.
[0030] After reading the depth image in uchar type, the actual depth information is measured by 0-255:
[0031]
[0032] where V D-M represents the maximum measurement value of the depth camera, TH DThe depth threshold value calculated is V O-3D Represent the morphological extension width of the non-rigid object in three-dimensional space.
[0033] 3. Beneficial effects
[0034] Compared with the prior art, the advantages of the present application are:
[0035] ① The scheme, through the dynamic target detection module collects the information of the dynamic target, then, through the target information screening processing module, the data information detection frame, acquires the screening data, then, through the depth information fusion processing module calculates the depth value, and sets the depth threshold interval, finally, through the feature filtering module generates the three-dimensional mask area of the dynamic target, and acquires the positioning information of the indoor mobile robot in the dynamic scene, so that the occlusion of static scene features to dynamic targets can be effectively reduced, and the cumulative amount of feature matching error can be reduced, and the risk of system tracking loss can be reduced.
[0036] ② The scheme, through the dynamic feature filtering, can make the utilization rate of static features increase by about 35%, and in the scene where the dynamic target proportion is less than or equal to 40%, stable tracking can be maintained, and the overall processing speed of the positioning information of the dynamic target reaches 16.7FPS, which can effectively meet the real-time demand. BRIEF DESCRIPTION OF DRAWINGS
[0037] Figure 1 The figure is a system schematic diagram of the positioning system in the present application;
[0038] Figure 2 The figure is a schematic diagram of the influence of the dynamic object on the SLAM system in the present application;
[0039] Figure 3 The figure is a flow schematic diagram of the positioning method in the present application. DETAILED DESCRIPTION
[0040] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application; obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments of the present application, and all other embodiments obtained by those skilled in the art without creative labor on the basis of the embodiments in the present application belong to the scope of protection of the present application.
[0041] Embodiment 1:
[0042] Please refer to Figure 1 - Figure 2 A positioning system of an indoor mobile robot in a dynamic scene, the positioning system comprising a dynamic target detection module, a target information screening processing module, a depth information fusion processing module and a feature filtering module.
[0043] The dynamic target detection module synchronously collects color images and depth information by using an RGB-D camera, and then uses YOLOv3 to perform dynamic target detection and output bounding box information.
[0044] Specifically, the dynamic target detection module receives a picture sequence captured by an RGB-D camera in a dynamic environment, matches color images and depth pictures through a time label, and then uses a target detection algorithm YOLOv3 to perform dynamic target recognition on the color pictures and returns the position coordinates of the bounding box.
[0045] The target information screening processing module uses non-maximum suppression to delete detection boxes with low confidence and retains the detection box with the highest confidence to obtain screening data.
[0046] Specifically, by scanning all output results of the neural network, the detection box with the highest confidence is retained, and the class label of the detection box is specified as the class with the highest score in the box, and the coordinates and width and height values of the non-rigid dynamic object detection box are recorded.
[0047] The depth information fusion processing module calculates the coordinates of the center point of the detection box and the depth values of the surrounding area, takes the minimum value of the depth values of the surrounding area as the target reference depth, and sets a depth threshold interval.
[0048] Specifically, the dynamic region feature points are removed, and valuable static scene feature information is retained to the greatest extent.
[0049] When calculating the depth values of the surrounding area, the surrounding area of 5*5 pixels is taken as the range for calculation.
[0050] When setting the depth threshold interval, the threshold interval range is ±8, and the corresponding quantization value is 0-255.
[0051] The feature filtering module generates a three-dimensional mask region of the dynamic target according to the depth information fusion processing result, filters the feature points in the region, retains the static features and inputs them into the SLAM system, and obtains the positioning information of the indoor mobile robot in the dynamic scene.
[0052] Specifically, in the feature filtering module, the target detection algorithm based on deep learning and the visual SLAM system based on geometric depth information are combined, so that the static information of target detection is adjusted sufficiently, the dynamic target is reduced to reduce the mismatch data generated by the visual SLAM system, and the positioning accuracy and running efficiency of the mobile robot are improved.
[0053] Specifically, the processing speed of the robot positioning system is 16.7FPS, which can meet the real-time demand of indoor movement of the robot.
[0054] Embodiment 2
[0055] Based on the above embodiment 1, further description is given.
[0056] See also Figure 3 , a positioning method for an indoor mobile robot in a dynamic scene, the specific implementation steps of the positioning method are as follows:
[0057] Step 1: Scan and collect information about dynamic targets: Use an RGB-D camera to collect image information of dynamic targets, then use YOLOv3 to detect dynamic targets and output dynamic target information detection frames;
[0058] Specifically, dynamic target detection is performed using the target detection box center point method, specifically:
[0059] The indoor mobile robot positioning system first receives a sequence of images taken by an RGB-D camera in a dynamic environment and matches the color image and depth image using time tags;
[0060] Then, DD-SLAM uses the target detection algorithm YOLOv3 to perform dynamic target recognition on the color image and return the position coordinates of the recognition box.
[0061] Thus, the indoor mobile robot positioning system can roughly know the current position of the dynamic non-rigid object.
[0062] Non-maximum suppression is used to remove low-confidence detection boxes. By scanning all outputs of the neural network, the highest-confidence object detection box is retained. The class label of the detection box is assigned to the class with the highest score within the box, while the coordinates, width, and height of the non-rigid dynamic object detection box are recorded. Therefore, after calculating the center point of the target detection box, this method obtains the geometric depth information of this center point in the corresponding depth image. The resulting data is the approximate depth information of the dynamic non-rigid object from the camera.
[0063] Step 2: Processing target information: First, the target information is screened to obtain the target data with the highest confidence. Then, based on the depth information fusion, the depth threshold interval is obtained and an unstructured mask of the dynamic non-rigid object is generated.
[0064] Specifically, when screening target information, the depth information obtained by the depth camera is used to propose a strategy for eliminating feature points in dynamic areas, which retains valuable static scene feature information to the greatest extent, making the SLAM system more accurate.
[0065] When performing deep and delicate fusion processing, the confidence of the highest target detection frame is retained by scanning all output results of the neural network, and the class label of the detection frame is designated as the highest scoring class in the frame, while recording the coordinates and width and height values of the non-rigid dynamic object detection frame. After calculating the center point of the target detection frame, the method obtains the geometric depth information of the center point on the corresponding depth image. At this time, the data obtained is the approximate depth information of the dynamic non-rigid object from the camera.
[0066] Step three, feature filtering and dynamic target positioning: combining the target detection algorithm based on deep learning and the visual SLAM system of geometric depth information, retaining sufficient static features of dynamic targets, judging the dynamic form of the object, obtaining the best unstructured mask, and displaying the positioning information of the indoor dynamic target.
[0067] Specifically, when performing feature filtering, the extension of the object in depth is considered, and a fixed depth threshold interval is set according to the specific dynamic human object to realize the unstructured mask.
[0068] After reading the depth image in uchar type, the actual depth information is measured by 0-255:
[0069]
[0070] where V D-M represents the maximum measurement value of the depth camera, generally 10m, TH D is the calculated depth threshold, V O-3D represents the morphological extension width of the non-rigid object in three-dimensional space.
[0071] Specifically, due to the existence of non-rigid features, the center point of the detection frame has a great possibility of not directly falling in the actual motion area of the dynamic object. The method selects the pixel area around the frame center point for uniform sampling and calculation, and selects the minimum value in the area as the depth value of the dynamic non-rigid target from the camera.
[0072] The above is only a preferred specific embodiment of the present application; however, the protection scope of the present application is not limited thereto. Any skilled person in the art can make equivalent substitutions or changes to the technical solutions and improved concepts of the present application within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application.
Claims
1. A positioning system for an indoor mobile robot in a dynamic scene, characterized by: The positioning system includes a dynamic target detection module, a target information screening and processing module, a depth information fusion processing module and a feature filtering module; The dynamic target detection module uses an RGB-D camera to synchronously capture color images and depth information. It then uses YOLOv3 to detect dynamic targets and output detection box information. The target information screening processing module uses non-maximum suppression to delete low-confidence detection frames, retain the target detection frames with the highest confidence, and obtain the screening data; The depth information fusion processing module calculates the coordinates of the center point of the detection frame and the depth value of the surrounding area, takes the minimum depth of the area as the target reference depth, and sets the depth threshold interval; The feature filtering module generates a three-dimensional mask area of the dynamic target based on the depth information fusion processing results, filters the feature points in the area, retains the static features and inputs them into the SLAM system to obtain the positioning information of the indoor mobile robot in dynamic scenes.
2. The positioning system for an indoor mobile robot in a dynamic scene according to claim 1, characterized in that: The dynamic target detection module receives a sequence of images taken by an RGB-D camera in a dynamic environment and matches the color image and depth image through time tags. DD-SLAM then uses the target detection algorithm YOLOv3 to identify dynamic targets in the color image and returns the position coordinates of the recognition box.
3. The positioning system for an indoor mobile robot in a dynamic scene according to claim 1, characterized in that: In the target information screening and processing module, the target detection frame with the highest confidence is retained by scanning all the output results of the neural network, and the class label of the detection frame is assigned to the class with the highest score in the frame, while the coordinates and width and height values of the non-rigid dynamic object detection frame are recorded.
4. The positioning system for an indoor mobile robot in a dynamic scene according to claim 1, characterized in that: In the depth information fusion processing module, dynamic area feature points are eliminated to maximize the grid retention of valuable static scene feature information; When calculating the depth value of the surrounding area, the calculation is performed with the surrounding area of 5×5 pixels as the range; When setting the depth threshold interval, the threshold interval range is ±8 and corresponds to a quantization value of 0 to 255.
5. The indoor mobile robot positioning system in dynamic scenarios according to claim 1, characterized in that: In the feature filtering module, the target detection algorithm based on deep learning and the visual SLAM system with geometric depth information are combined to make the static information of target detection sufficiently adjusted, and reduce the mismatching data generated by the visual SLAM system due to dynamic targets, thereby improving the positioning accuracy and operation efficiency of the mobile robot.
6. A method for positioning an indoor mobile robot in a dynamic scene according to any one or more of claims 1 to 5, characterized in that: The specific implementation steps of the positioning method are as follows: Step 1: Scan and collect information about dynamic targets: Use an RGB-D camera to collect image information of dynamic targets, then use YOLOv3 to detect dynamic targets and output dynamic target information detection frames; Step 2: Processing target information: First, the target information is screened to obtain the target data with the highest confidence. Then, based on the depth information fusion, the depth threshold interval is obtained and an unstructured mask of the dynamic non-rigid object is generated. Step 3: Perform feature filtering and form dynamic target positioning: Combine the target detection algorithm based on deep learning and the visual SLAM system with geometric depth information to retain sufficient static features of dynamic targets, determine the dynamic form of the object, obtain the best unstructured mask, and display the positioning information of indoor dynamic targets.
7. The method for positioning an indoor mobile robot in a dynamic scene according to claim 6, characterized in that: In step 1, dynamic target detection is performed using the target detection frame center point method, specifically: The indoor mobile robot positioning system first receives a sequence of images taken by an RGB-D camera in a dynamic environment and matches the color image and depth image using time tags; Then, DD-SLAM uses the target detection algorithm YOLOv3 to perform dynamic target recognition on the color image and return the position coordinates of the recognition box.
8. The method for positioning an indoor mobile robot in a dynamic scene according to claim 6, characterized in that: In the second step, when performing target information screening, the depth information obtained by the depth camera is used to propose a strategy for eliminating feature points in dynamic areas, thereby retaining valuable static scene feature information to the greatest extent, making the SLAM system have higher accuracy; When performing depth and fine-grained fusion processing, the target detection frame with the highest confidence is retained by scanning all the output results of the neural network, and the class label of the detection frame is assigned to the class with the highest score in the frame. At the same time, the coordinates and width and height values of the non-rigid dynamic object detection frame are recorded. After calculating the center point of the target detection frame, this method obtains the geometric depth information of the center point on the corresponding depth image. The data obtained at this time is the approximate depth information of the dynamic non-rigid object from the camera.
9. The method for positioning an indoor mobile robot in a dynamic scene according to claim 6, characterized in that: In step 3, when performing feature filtering, the depth extension of the object is taken into consideration, and a fixed depth threshold interval is set according to a specific dynamic human object to implement an unstructured mask; After reading the depth image with uchar type, the actual depth information is measured in 0-255: Where V D-M Represents the maximum measurement value of the depth camera, TH D is the calculated depth threshold, V O-3D Represents the morphological extension width of a non-rigid object in three-dimensional space.