A Method and System for Mobile Robot Localization and Mapping in a Dynamic Environment

Through the example segmentation network and dense optical flow field fusion, the dynamic area is shielded, and the accuracy and accuracy of robot positioning and map construction in dynamic environments are solved, and higher-precision camera pose estimation and point cloud map construction are achieved.

CN115290072BActive Publication Date: 2025-07-08SEVNCE ROBOTICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210956838.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-10
Publication Date
2025-07-08
Estimated Expiration
2042-08-10

AI Technical Summary

Technical Problem

In the prior art, when robot positioning and mapping is located in dynamic environments, dynamic object detection is inaccurate, resulting in low camera position estimation accuracy and poor mapping accuracy.

Method used

The dynamic object mask is obtained through the example segmentation network, combined with dense optical flow field fusion, the object outline is repaired using deep images, and the dynamic area is blocked, and camera position estimation and mapping are only carried out on the static area.

Benefits of technology

The camera position estimation accuracy and mapping accuracy of the visual SLAM module are improved, errors caused by dynamic objects are reduced, and the robustness of robot positioning is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115290072B_ABST
    Figure CN115290072B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for mobile robot positioning and mapping in a dynamic environment. The method includes: S1, acquiring an RGB image and a depth image; S2, inputting the current frame RGB image into an instance segmentation network to obtain a dynamic object mask; inputting the current frame and the next frame RGB images into an optical flow estimation network to obtain a dense optical flow field; S3, fusing the dense optical flow field and the dynamic object mask to obtain a fused mask; S4, using the current frame depth image to repair the object contour in the fused mask to obtain a repaired mask; S5, using the repaired mask to mask the dynamic region in the current frame RGB image, and the visual SLAM module performs camera pose estimation and constructs a point cloud map. The dense optical flow field information is used to make up for the deficiency that the instance segmentation network can only detect prior dynamic objects, and the depth image is used to repair the object contour in the fused mask, thereby improving the camera pose estimation accuracy, mapping accuracy, and positioning accuracy of the visual SLAM module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robot navigation technology, and in particular to a method and system for positioning and mapping a mobile robot in a dynamic environment. Background Art

[0002] With the development of science and technology and the improvement of living standards, robots have gradually entered people's daily lives. Current home service robots are usually based on simultaneous localization and mapping (SLAM) technology to build sparse landmark maps containing geometric information, which can be used to perform navigation and positioning tasks.

[0003] The traditional ORB-SLAM2 system is a comprehensive system that integrates monocular, binocular, and depth cameras. It has three parallel threads: tracking, local mapping, and closed-loop detection. When a closed loop is detected, a fourth thread is opened for global optimization. It is a representative of the feature point method used in visual SLAM systems. Its operation process assumes that the environment is a static scene and the changes in the main part of the scene are caused by the movement of the camera. However, in the actual environment, there are inevitably moving objects, such as people indoors and vehicles driving outdoors. Although ORB-SLAM2 uses RANSAC to remove outliers in the matching process to improve robustness in dynamic environments, this method is often difficult to calculate in the presence of many dynamic objects.

[0004] In response to the above problems, Bescos et al. proposed the DynaSLAM (B. Bescos, JM Faci l, J. Civera, J. Neira. DynaSLAM: Tracking, Mapping and Inpainting in Dynamic Scenes. IEEE Robotics and Automation Letters, 3 (4), 4076-4083, 2018) system, a visual SLAM system (VSLAM for short). It is based on ORB-SLAM2 and adds dynamic object detection and background inpainting. It can detect dynamic objects through multi-view geometry and deep learning. The obtained static scene map allows the inpainting of frame backgrounds blocked by dynamic objects. Although DynaSLAM performs better than traditional ORB-SLAM2 in high dynamic environments, the DynaSLAM algorithm does not remove feature points on actual moving objects, which leads to large deviations when calculating camera trajectories and poses, resulting in low camera pose estimation accuracy and poor mapping accuracy.

[0005] In addition, some technical solutions for detecting prior dynamic objects based only on deep learning have emerged in the prior art. However, these technical solutions only detect prior dynamic objects based on deep learning, resulting in insufficient robustness in dealing with dynamic environments in many scenarios. For example, in the scenario where a person drags an object to move, it cannot be detected, which also leads to low accuracy in camera pose estimation and poor mapping accuracy. Summary of the Invention

[0006] The present invention aims to at least solve the technical problems existing in the prior art, and provides a method and system for positioning and mapping a mobile robot in a dynamic environment.

[0007] To achieve the above object of the present invention, according to the first aspect of the present invention, a method for positioning and mapping a mobile robot in a dynamic environment is provided, including: Step S1, obtaining a synchronously captured RGB image and a depth image; Step S2, inputting the current frame RGB image into a pre-trained instance segmentation network to obtain a dynamic object mask; inputting the current frame and the next frame RGB images into an optical flow estimation network to obtain a dense optical flow field; Step S3, fusing the dense optical flow field and the dynamic object mask to obtain a fused mask; Step S4, using the current frame depth image to repair the object contour in the fused mask to obtain a repaired mask; Step S5, using the repaired mask to mask the dynamic region in the current frame RGB image, and inputting the RGB image after masking the dynamic region and the current frame depth image into a visual SLAM module for camera pose estimation and building a point cloud map.

[0008] To achieve the above object of the present invention, according to the second aspect of the present invention, a computer storage medium is provided, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the method for positioning and mapping a mobile robot in a dynamic environment as described in the first aspect of the present invention.

[0009] To achieve the above object of the present invention, according to the third aspect of the present invention, there is provided a mobile robot positioning and mapping system in a dynamic environment, including: an image acquisition module for acquiring a synchronously captured RGB image and a depth image; a dynamic object mask acquisition module for inputting the current frame RGB image into a pre-trained instance segmentation network to obtain a dynamic object mask; a dense optical flow field acquisition module for inputting the current frame and the next frame RGB images into an optical flow estimation network to obtain a dense optical flow field; a fusion mask acquisition module for fusing the dense optical flow field and the dynamic object mask to obtain a fusion mask; a repair mask acquisition module for using the current frame depth image to repair the object contour in the fusion mask to obtain a repair mask; a shielding module for using the repair mask to shield the dynamic area in the current frame RGB image; and a visual SLAM module for performing camera pose estimation and building a point cloud map based on the RGB image after shielding the dynamic area and the current frame depth image.

[0010] To achieve the above object of the present invention, according to the fourth aspect of the present invention, there is provided a mobile robot, on which a depth camera and a processor are provided; the processor receives the RGB image and the depth image synchronously captured by the depth camera, and executes the steps of the method for positioning and mapping a mobile robot in a dynamic environment according to the first aspect of the present invention.

[0011] The above technical solution: uses an instance segmentation network to segment the objects in the RGB image to obtain a dynamic object mask, uses an optical flow estimation network to extract the dense optical flow field of two consecutive frames of RGB images, judges the objects moving relative to the camera through the dense optical flow field, and fusing the dense optical flow field and the dynamic object mask can make up for the deficiency that the instance segmentation network can only detect prior dynamic objects, realizes accurate detection of real-time dynamic areas, and avoids the problem of misjudging the object motion information due to pure instance segmentation; uses the depth image to repair the object contour in the fusion mask, further improving the accuracy of the dynamic area in the repair mask; shields the dynamic area in the RGB image based on the repair mask, ensuring that the feature points extracted in the visual SLAM all belong to the static object area, greatly reducing the error caused by the extracted bad feature points, enabling the visual SLAM module to perform camera pose estimation and build a point cloud map only for the area relatively static to the camera, improving the camera pose estimation accuracy, mapping accuracy and positioning accuracy of the visual SLAM module, and solving the problem of the deficiency of the visual SLAM module in a dynamic environment. Description of the Drawings

[0012] Figure 1 is a schematic flow chart of the method for positioning and mapping a mobile robot in a dynamic environment in Embodiment 1 of the present invention;

[0013] Figure 2 is a schematic structural diagram of the optical flow estimation network in Embodiment 1 of the present invention;

[0014] Figure 3 It is a schematic diagram of the instance segmentation network structure in Embodiment 1 of the present invention;

[0015] Figure 4 It is a specific flow block diagram in an application scenario of Embodiment 1 of the present invention. Detailed implementation manners

[0016] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.

[0017] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0018] In the description of the present invention, unless otherwise specified and limited, it should be noted that the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the communication inside two elements. It can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific situations.

[0019] Embodiment 1

[0020] This embodiment discloses a method for mobile robot positioning and mapping in a dynamic environment, as Figure 1 shown, including:

[0021] Step S1, obtaining RGB images and depth images captured synchronously; preferably but not limited to capturing RGB images and depth images synchronously through an RGBD depth camera, and the RGB images and depth images correspond to each other.

[0022] Step S2, inputting the current frame RGB image into a pre-trained instance segmentation network to obtain a dynamic object mask. The pixel values in the dynamic object area of the dynamic object mask are the second numerical value, and the pixel values in the non-dynamic object area are the first numerical value. The first numerical value and the second numerical value are not equal, and the values of the first numerical value and the second numerical value are 1 or 0; inputting the current frame and the next frame of RGB images into an optical flow estimation network to obtain a dense optical flow field.

[0023] The instance segmentation network preferably but not limited to adopt the SOLOV2 network structure. As Figure 3 shown, input each frame of RGB image into the trained instance segmentation network SOLOV2. As Figure 3 shown, SOLOV2 is divided into two branches: the kernel branch and the feature branch. After convolving each pixel of the kernel branch and each feature of the feature branch, a dynamic object mask is obtained.

[0024] Preferably, the training process of the instance segmentation network includes:

[0025] (1) Construct a sample set. The samples are RGB images. Each RGB image sample corresponds to an object mask. The object mask can be obtained by manually prior recognition and segmentation processing of the objects on the RGB image sample. Semantic labels are set for each object (which can be static objects and dynamic objects) on the object mask. The semantic label is used to mark whether the object is dynamic or static. Divide the sample set into a training set, a validation set, and a test set.

[0026] (2) Construct an instance segmentation network, preferably but not limited to the SOLOV2 network. Use the training set to train the instance segmentation network, and use the validation set and the test set to verify and test the trained instance segmentation network respectively to obtain the trained instance segmentation network.

[0027] (3) Input the current frame RGB image into the trained instance segmentation network to obtain the object mask and the semantic label of each object on the object mask. Filter out the dynamic object mask according to the semantic label to obtain a dynamic object mask that only includes dynamic objects.

[0028] Further preferably, when setting the samples, only the dynamic object regions can be manually priorly labeled on the object mask corresponding to each sample. In this way, there is no need to set semantic labels, and the trained instance segmentation network automatically outputs the dynamic object mask, which simplifies the processing process.

[0029] The optical flow estimation network preferably but not limited to adopt the LiteFlowNet network structure, and requires two consecutive frames of RGB images to be input. As Figure 2 shown, select the current frame and the next frame of RGB images. The LiteFlowNet network internally calculates a two-dimensional dense optical flow vector for each pixel point of the current frame RGB image. The dense optical flow vectors of multiple pixel points constitute a dense optical flow field. As Figure 2 shown, the F i is used to describe the image I iCNN features. M:S represents the cascaded flow inference module, including the feature descriptor matching unit M and the sub-pixel refinement unit S. R represents the regularization module. The flow inference and regularization modules correspond to the data fidelity and regularization terms in the traditional energy minimization method, respectively.

[0030] Step S3, fuse the dense optical flow field and the dynamic object mask to obtain a fused mask.

[0031] Since the dynamic object regions in the dynamic object mask are only obtained by the instance segmentation network based on prior knowledge, there is a high probability of missing the judgment of dynamic regions. The dense optical flow field can accurately reflect all dynamic pixel points, and the fusion of the two can effectively expand the dynamic region. To facilitate the fusion of the dense optical flow field and the dynamic object mask, preferably, the process of fusing the dense optical flow field and the dynamic object mask to obtain a fused mask specifically includes:

[0032] Step S31, convert the dense optical flow field into an optical flow RGB image in the Munsell color system, and perform binary processing on the optical flow RGB image to obtain an optical flow mask; preferably, in the optical flow RGB image, the hue represents the motion direction of the pixel point, that is, the optical flow field direction, and the chroma represents the momentum magnitude of the pixel point, that is, the magnitude of the optical vector. The greater the motion of the pixel point, the greater its chroma and the darker the color. During the binary processing of the optical flow RGB image, set the pixel value of the static pixel point to the first value, and set the pixel value of the dynamic pixel point to the second value to obtain the optical flow mask.

[0033] Step S32, perform a fusion operation on the optical flow mask and the dynamic object mask to obtain a fused mask. The fusion method can be selected according to the values of the first value and the second value to ensure that the dynamic region in the fused mask after fusion is larger than the dynamic object region in the dynamic object mask. For example, when the first value is 1 and the second value is 0, perform an AND operation on the optical flow mask and the dynamic object mask to obtain a fused mask; for another example, when the first value is 0 and the second value is 1, perform an OR operation on the optical flow mask and the dynamic object mask to obtain a fused mask.

[0034] Step S4, use the current frame depth image to repair the object contour in the fused mask to obtain a repaired mask. To achieve fast and effective repair of the contour only for the dynamic object region, preferably, step S4 specifically includes:

[0035] Step S41, traverse all pixel points of the current frame depth image based on an edge judgment algorithm to extract edge pixel points. All edge pixel points form multiple closed regions, and an edge information mask including multiple closed regions is established; further preferably, to more accurately identify edge pixel points and reduce the calculation amount, the process of the edge judgment algorithm is: for the pixel point (u, v) in the depth image, judge whether the pixel point (u, v) is a pixel point on the edge Edge according to the following formula:

[0036]

[0037] A block represents a rectangular block (u:u + 1, v:v + 1) with the pixel points (u, v) and (u + 1, v + 1) as the diagonal, and max(d block ) represents the maximum depth value in the block, and min(d block ) represents the minimum depth value in the block. τ represents the depth threshold, preferably but not limited to 0.3 meters to 0.5 meters. f(max(d block ) - min(d block )) represents taking the difference between max(d block ) and min(d block ).

[0038] Further preferably, to reduce the amount of computation and speed up the edge recognition of the dynamic object area in the depth image. In step S41, first filter out a preset area in the current frame depth image. Since the current frame depth image corresponds to the current frame RGB image, the current frame depth image corresponds to the fusion mask. Map the dynamic object area in the fusion mask to the current frame depth image to obtain a mapped area, and select the neighborhood of the surrounding pixel points as a supplementary point set around the mapped area. The neighborhood radius is preferably but not limited to 10 to 100 pixel points. Combine the mapped area and all supplementary point sets to obtain a preset area. Use an edge judgment algorithm to perform edge recognition within the preset area to obtain an edge information mask.

[0039] Step S42, for each dynamic object area in the fusion mask, find the closed area with the most overlapping pixel points with it in the edge information mask, denoted as the first closed area, and take the union of the first closed area and the dynamic object area, and use this union as the repaired dynamic object area corresponding to the dynamic object area.

[0040] Step S5, use the repaired mask to shield the dynamic area in the current frame RGB image, and input the RGB image after shielding the dynamic area and the current frame depth image into the visual SLAM module for camera pose estimation and building a point cloud map. The visual SLAM module includes a visual odometer and a point cloud map generation unit, and the point cloud map generation unit includes a point cloud stitching sub-unit, a noise elimination sub-unit, and a local point cloud acquisition sub-unit. The visual SLAM module preferably but not limited to can also adopt an existing VSLAM process framework, which will not be elaborated here.

[0041] In an application scenario of this embodiment, step S5 includes:

[0042] Use the fusion mask to shield the dynamic object area in the current frame RGB image to obtain a first RGB static image, and input the first RGB static image into the visual odometer of the visual SLAM module for camera pose processing;

[0043] Use the repair mask to mask the dynamic object area in the current frame RGB image to obtain the second RGB static image, and input the second RGB static image and the current frame depth image into the point cloud map generation unit of the visual SLAM module to obtain the global point cloud map. Specifically, the second RGB static image and the current frame depth image are input into the local point cloud acquisition subunit to obtain the local point cloud, and the point cloud stitching subunit generates the global point cloud map based on the local point cloud and the camera pose.

[0044] In another application scenario of this embodiment, as Figure 4 shown, step S5 includes:

[0045] Use the repair mask to mask the dynamic area in the current frame RGB image to obtain the second RGB static image, input the second RGB static image into the visual odometry process of the visual SLAM module to obtain the camera pose, and input the second RGB static image and the current frame depth image into the point cloud map generation unit of the visual SLAM module to obtain the global point cloud map. The working principle of the point cloud map generation unit has been described in detail in the previous application scenario and will not be elaborated here.

[0046] In a preferred implementation manner of this embodiment, for the convenience of fusing with the dynamic object mask and making the generated optical flow mask more accurate, the binary processing process of the optical flow RGB image is as follows:

[0047] Step A, define the area composed of the pixel points with chromaticity less than the chromaticity threshold in the optical flow RGB image as the relatively camera-stationary area, and define the area other than the relatively camera-stationary area in the optical flow RGB image as the initial relatively camera-moving area; count the value range [x min , x max of the chromaticity in the optical flow RGB image, where x min represents the minimum chromaticity value of the optical flow RGB image, x max represents the maximum chromaticity value of the optical flow RGB image, let x τ represent the chromaticity threshold, set the initial value of x τ to be x min , and stepwise increase the value of x τ until P(x≥x τ )≥0.5, where P(x≥x τ ) represents the proportion of the pixel points with chromaticity greater than or equal to x τ in the optical flow RGB image. Since generally the relatively camera-stationary area is larger than the relatively camera-moving area, when its proportion is greater than or equal to 0.5, it can be considered that the pixel points in this chromaticity area belong to the relatively camera-stationary area.

[0048] Step B, obtain the hue mean value of the pixel points in the relatively camera-stationary area.

[0049] Step C: Calculate the difference between the hue of the pixel points in the initial relative camera motion area and the average hue. Calculate the difference between the hue of the pixel points in the initial relative camera motion area and the average hue. If the difference is greater than or equal to the hue difference threshold, add the corresponding pixel points to the actual relative camera motion area. If the difference is less than the hue difference threshold, add the corresponding pixel points to the relative camera static area. The hue difference threshold is preferably but not limited to a value between 5 and 50 sub - division levels.

[0050] Step D: Assign the pixel values of the pixel points in the relative camera static area to the first value, and assign the pixel values of the pixel points in the actual relative camera motion area to the second value to obtain an optical flow mask.

[0051] In this embodiment, aiming at the problems existing in the background technology and combining the current trend of the development of deep learning, a VSLAM algorithm integrating instance segmentation algorithm and optical flow estimation is proposed, which can be used to improve the positioning accuracy of mobile robots in dynamic environments. Among them, the instance segmentation algorithm SOLOv2 can perform instance segmentation on a continuous input RGB image sequence to obtain the masks of each object in the RGB image. The input of the optical flow estimation LiteFlowNet network is two adjacent consecutive RGB images. At each point in the current frame, a two - dimensional optical flow vector is predicted, and an optical flow RGB image with the motion information of each pixel point is output. By using the output optical flow RGB image to judge the relatively moving objects, it makes up for the deficiency that the SOLOv2 algorithm can only detect prior dynamic objects. In the VSLAM system, by removing dynamic feature points from the mask area of dynamic objects, the accuracy and accuracy of the trajectory and pose estimation in the process of mobile robot positioning can be improved.

[0052] This embodiment fully considers the problem of mobile robot positioning in a dynamic environment. By introducing the optical flow estimation network LiteFlowNet to process two consecutive frames of images, a two - dimensional optical flow vector is predicted at each point in the current frame to analyze the relative motion of each object, and the mask of the prior dynamic objects in the current frame of the instance segmentation algorithm SOLOv2 is fused, so as to obtain a more accurate dynamic object area. When extracting feature points at the front end, the dynamic object area is shielded, ensuring that the extracted feature points all belong to the static object area, greatly reducing the errors caused by extracting bad feature points. It solves the problem of insufficient robustness of the traditional ORB - SLAM2 system and the SLAM system based solely on instance segmentation in the positioning of mobile robots in dynamic environments.

[0053] Embodiment 2

[0054] This embodiment discloses a computer storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, at least one program, the code set or the instruction set is loaded and executed by a processor to implement the method for positioning and mapping a mobile robot in a dynamic environment as described in Embodiment 1.

[0055] Embodiment 3

[0056] This embodiment discloses a system for positioning and mapping a mobile robot in a dynamic environment, including: an image acquisition module for acquiring a synchronously captured RGB image and a depth image; a dynamic object mask acquisition module for inputting the current frame RGB image into a pre-trained instance segmentation network to obtain a dynamic object mask; a dense optical flow field acquisition module for inputting the current frame and the next frame RGB images into an optical flow estimation network to obtain a dense optical flow field; a fusion mask acquisition module for fusing the dense optical flow field and the dynamic object mask to obtain a fusion mask;

[0057] a repaired mask acquisition module for repairing the object contour in the fusion mask by using the current frame depth image to obtain a repaired mask; a shielding module for shielding the dynamic area in the current frame RGB image by using the repaired mask; a visual SLAM module for performing camera pose estimation and establishing a point cloud map based on the RGB image after shielding the dynamic area and the current frame depth image.

[0058] In this embodiment, preferably, the repaired mask acquisition module performs:

[0059] Traversing all pixel points of the current frame depth image based on an edge judgment algorithm to extract edge pixel points, all the edge pixel points form a plurality of closed regions, and an edge information mask including the plurality of closed regions is established; for each dynamic object region in the fusion mask, finding the closed region with the most overlapping pixel points with it from the edge information mask, denoted as the first closed region, taking the union of the first closed region and the dynamic object region, and using the union as the repaired dynamic object region corresponding to the dynamic object region.

[0060] Embodiment 4

[0061] This embodiment discloses a mobile robot, on which a depth camera and a processor are provided; the processor receives the RGB image and the depth image synchronously captured by the depth camera, and executes the steps of the method for positioning and mapping a mobile robot in a dynamic environment as described in Embodiment 1.

[0062] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples.

[0063] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

Claims

1. A method for mobile robot positioning and mapping in a dynamic environment, characterized in that Including: Step S1: Obtain the simultaneously captured RGB image and depth image; Step S2: Input the current frame RGB image into a pre-trained instance segmentation network to obtain a dynamic object mask; Input the current frame and the next frame RGB images into an optical flow estimation network to obtain a dense optical flow field; Step S3: Fuse the dense optical flow field and the dynamic object mask to obtain a fused mask; Step S4: Use the current frame depth image to repair the object contour in the fused mask to obtain a repaired mask; Step S5: Use the repaired mask to mask the dynamic area in the current frame RGB image, and input the RGB image after masking the dynamic area and the current frame depth image into a visual SLAM module for camera pose estimation and building a point cloud map; Among them, in step S3, fusing the dense optical flow field and the dynamic object mask to obtain a fused mask specifically includes: Convert the dense optical flow field into an optical flow RGB image in the Munsell color system, and perform binary processing on the optical flow RGB image to obtain an optical flow mask; Perform a fusion operation on the optical flow mask and the dynamic object mask to obtain a fused mask, and ensure that the dynamic area in the fused mask after fusion is larger than the dynamic object area in the dynamic object mask; In the optical flow RGB image, the hue represents the movement direction of the pixel point, and the chroma represents the momentum magnitude of the pixel point; Among them, the binary processing process of the optical flow RGB image is: Define the area composed of pixel points with chroma less than the chroma threshold in the optical flow RGB image as the relatively camera static area, and define the area other than the relatively camera static area in the optical flow RGB image as the initial relatively camera moving area; Obtain the hue mean value of the pixel points in the relatively camera static area; Calculate the difference between the hue of the pixel points in the initial relatively camera moving area and the hue mean value. If the difference is greater than or equal to the hue difference threshold, add the corresponding pixel points to the actual relatively camera moving area. If the difference is less than the hue difference threshold, add the corresponding pixel points to the relatively camera static area; Assign the pixel value of the pixel points in the relatively camera static area to the first value, and assign the pixel value of the pixel points in the actual relatively camera moving area to the second value to obtain an optical flow mask; The pixel value of the dynamic object area in the dynamic object mask is the second value, and the pixel value of the non-dynamic object area in the dynamic object mask is the first value, and the first value and the second value are not equal.

2. The method for positioning and mapping of a mobile robot in a dynamic environment according to claim 1, characterized in that, The values of the first value and the second value are 1 or 0.

3. The method for mobile robot positioning and mapping in a dynamic environment according to claim 1 or 2, characterized in that, In step S4, using the current frame depth image to repair the object contour in the fused mask to obtain a repaired mask specifically includes: Based on an edge judgment algorithm, traverse all pixel points of the current frame depth image to extract edge pixel points. All edge pixel points form multiple closed regions, and establish an edge information mask including the multiple closed regions; For each dynamic object area in the fused mask, find the closed region with the most overlapping pixel points with it in the edge information mask, denoted as the first closed region, and take the union of the first closed region and the dynamic object area, and use the union as the repaired dynamic object area corresponding to the dynamic object area.

4. The method for mobile robot positioning and mapping in a dynamic environment according to claim 3, wherein The process of the edge judgment algorithm is as follows: For the pixel point (u, v) in the depth image, it is judged whether the pixel point (u, v) is a pixel point on the edge Edge according to the following formula: A block represents a rectangular block (u:u + 1, v:v + 1) with the pixel points (u, v) and (u + 1, v + 1) as its diagonal corners. max(d block ) represents the maximum depth value in the block, and min(d block ) represents the minimum depth value in the block. τ represents the depth threshold.

5. The method for positioning and mapping of a mobile robot in a dynamic environment according to claim 1 or 3 or 4, characterized in that The step S5 includes: Using the repair mask to mask the dynamic area in the current frame RGB image to obtain a second RGB static image, inputting the second RGB static image into the visual odometer of the visual SLAM module to process and obtain the camera pose, and inputting the second RGB static image and the current frame depth image into the point cloud map generation unit of the visual SLAM module to obtain the global point cloud map.

6. A computer storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set are loaded and executed by the processor to implement the mobile robot positioning and mapping method according to any one of claims 1-5.

7. A mobile robot positioning and mapping system in a dynamic environment, characterized in that, Including: An image acquisition module for acquiring the synchronously captured RGB image and depth image; A dynamic object mask acquisition module for inputting the current frame RGB image into a pre-trained instance segmentation network to obtain a dynamic object mask; A dense optical flow field acquisition module for inputting the current frame and the next frame RGB images into an optical flow estimation network to obtain a dense optical flow field; A fusion mask acquisition module for fusing the dense optical flow field and the dynamic object mask to obtain a fusion mask; A repair mask acquisition module for using the current frame depth image to repair the object contour in the fusion mask to obtain a repair mask; A masking module for using the repair mask to mask the dynamic area in the current frame RGB image; A visual SLAM module for performing camera pose estimation and building a point cloud map based on the RGB image after masking the dynamic area and the current frame depth image; The fusion mask acquisition module fuses the dense optical flow field and the dynamic object mask to obtain a fusion mask, specifically including: Converting the dense optical flow field into an optical flow RGB image in the Munsell color system, and performing binary processing on the optical flow RGB image to obtain an optical flow mask; Performing a fusion operation on the optical flow mask and the dynamic object mask to obtain a fusion mask, and ensuring that the dynamic area in the fused fusion mask is larger than the dynamic object area in the dynamic object mask; In the optical flow RGB image, the hue represents the motion direction of the pixel point, and the chroma represents the momentum magnitude of the pixel point; Among them, the binary processing process of the optical flow RGB image is: Defining the area composed of pixel points with chroma less than the chroma threshold in the optical flow RGB image as the relatively camera static area, and defining the area other than the relatively camera static area in the optical flow RGB image as the initial relatively camera moving area; Obtaining the hue mean value of the pixel points in the relatively camera static area; Calculating the difference between the hue of the pixel points in the initial relatively camera moving area and the hue mean value, if the difference is greater than or equal to the hue difference threshold, adding the corresponding pixel points to the actual relatively camera moving area, and if the difference is less than the hue difference threshold, adding the corresponding pixel points to the relatively camera static area; Assigning the pixel value of the pixel points in the relatively camera static area to the first numerical value, and assigning the pixel value of the pixel points in the actual relatively camera moving area to the second numerical value to obtain an optical flow mask; The pixel value of the dynamic object area in the dynamic object mask is the second value, and the pixel value of the non-dynamic object area in the dynamic object mask is the first value, and the first value is not equal to the second value.

8. The mobile robot positioning and mapping system in a dynamic environment according to claim 7, wherein The repair mask acquisition module executes: Traverse all pixel points of the current frame depth image based on the edge judgment algorithm to extract edge pixel points, and all edge pixel points form multiple closed areas, and establish an edge information mask including the multiple closed areas; For each dynamic object area in the fusion mask, find the closed area with the most overlapping pixel points with it from the edge information mask, denoted as the first closed area, and take the union of the first closed area and the dynamic object area, and use the union as the repaired dynamic object area corresponding to the dynamic object area.

9. A mobile robot, characterized in that, A depth camera and a processor are provided on the mobile robot; The processor receives the RGB image and the depth image synchronously captured by the depth camera, and executes the steps of the method for positioning and mapping a mobile robot in a dynamic environment according to any one of claims 1-5.

Citation Information

Patent Citations

  • Semantic SLAM robustness improvement method based on instance segmentation

    CN111581313A

  • Visual loopback detection method based on semantic segmentation and image restoration in dynamic scene

    CN111696118A

  • RGB-D visual SLAM method applied to indoor dynamic environment

    CN116758112A

  • Wall-climbing robot visual SLAM positioning method based on dense optical flow and image segmentation in dynamic environment

    CN118628571A

  • Image processing method, image segmentation model training method and related apparatus

    WO2021139625A1