Semantic and Multi-Object Error-Minimized Dynamic SLAM Method and Robot
By extracting static images and calculating reprojection errors in the visual SLAM method, selecting the camera pose estimation matrix with the smallest comprehensive projection errors, the problem of large camera pose estimation errors in dynamic environments is solved, and the accuracy of graph construction is improved.
Patent Information
- Application Number
- CN202210963410.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-11
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-08-11
AI Technical Summary
The existing visual SLAM method is difficult to accurately distinguish dynamic objects in dynamic environments, resulting in large errors in camera position estimation, affecting the accuracy of map construction.
By acquiring continuous RGB image frames, extracting the static images of the reference frame and the current frame, compute the camera pose estimation matrix for each object area, and selecting the camera pose estimation matrix with the smallest comprehensive projection error as the inter-frame optimal camera pose matrix.
The accuracy of camera position estimation and graph construction accuracy are improved, and the camera motion trajectory obtained is more accurate.
Smart Images

Figure CN115326073B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mobile robot vision positioning and mapping, and particularly to a semantic and multi-object error-minimized dynamic SLAM method and a robot. Background Art
[0002] Simultaneous localization and mapping (SLAM) technology is mainly divided into two categories: laser SLAM and visual SLAM. Among them, visual SLAM refers to the use of a camera mounted on a robot to perceive the surrounding environment without prior information, establish an environmental model during movement, and simultaneously estimate its own position. The camera has low cost, simple structure, larger amount of information collected, and wider application range. Currently, many products apply visual SLAM technology for positioning and mapping, such as autonomous vehicles, virtual reality, drones, and household floor-sweeping robots, etc.
[0003] There are many frameworks for visual SLAM, but most of them are constructed based on the premise of assuming a static environment. The traditional ORB-SLAM2 system is an integrated system that combines monocular, binocular, and depth cameras. Its operation process assumes the environment as a static scene and the main part of the scene changes due to camera movement. However, in the actual environment, there are inevitably moving objects, such as people indoors and vehicles driving outdoors, etc. Although ORB-SLAM2 uses the random sample consensus (RANSAC) method to remove outliers in the matching process to improve robustness in a dynamic environment, when dynamic objects dominate the environment, this method often has difficulty distinguishing outliers, resulting in a large error in camera pose estimation.
[0004] In view of the above problems, Chinese Patent (CN112308921A) discloses a method for dynamic SLAM based on semantic and geometric joint optimization. This method uses the instance segmentation network MASK-RCNN to extract semantic information from objects and generate a semantic binary mask, and performs geometric segmentation on the image to generate a geometric binary mask. Based on the dynamic and static information of the two masks, the feature point weights are calculated, and then the dynamic and static attributes of the feature points are judged.
[0005] The SLAM method disclosed in the above patent mainly determines the dynamic and static attributes of feature points by fusing information from two approaches: semantics and geometry. During the fusion process, different weighting coefficients and thresholds are set to calculate the attributes of feature points. However, due to the existence of errors during movement and the large interference of the weighting coefficient setting on the entire system, the robustness of this solution is relatively low in the actual dynamic environment application. This leads to incorrect final fusion results of geometric segmentation and semantic segmentation, resulting in large deviations when calculating the camera trajectory and pose, and having a serious impact on the entire SLAM system. Summary of the Invention
[0006] The present invention aims to at least solve the technical problems existing in the prior art, and provides a semantic and multi-object error-minimized dynamic SLAM method and a robot.
[0007] To achieve the above object of the present invention, according to the first aspect of the present invention, there is provided a semantic and multi-object error-minimized dynamic SLAM method, which acquires continuous RGB image frames, and performs steps S1 to S5 on two adjacent RGB image frames to obtain an optimal inter-frame camera pose matrix, and obtains a camera motion trajectory based on the optimal inter-frame camera pose matrices of all adjacent two RGB image frames, wherein: Step S1, acquiring a reference frame RGB image and a current frame RGB image; Step S2, acquiring a reference frame static image based on the reference frame RGB image, and acquiring a current frame static image based on the current frame RGB image, and the reference frame static image / current frame static image does not include a prior dynamic object area; Step S3, acquiring a camera pose estimation matrix corresponding to each object area in the reference frame static image; Step S4, for each camera pose estimation matrix, obtaining a reprojection error of the object area other than the object area corresponding to the camera pose estimation matrix in the reference frame static image on the current frame RGB image based on the camera pose estimation matrix, and fusing the reprojection errors of all object areas other than the object area corresponding to the camera pose estimation matrix in the reference frame static image to obtain a comprehensive projection error of the camera pose estimation matrix; Step S5, selecting the camera pose estimation matrix with the minimum comprehensive projection error as the optimal inter-frame camera pose matrix.
[0008] To achieve the above object of the present invention, according to the second aspect of the present invention, the present invention provides a semantic and multi-object error-minimized dynamic SLAM system, including: a camera motion trajectory acquisition module that acquires consecutive RGB image frames. For two adjacent RGB image frames, an image acquisition module, a static image acquisition module, a camera pose estimation matrix acquisition module, a comprehensive projection error acquisition module, and an inter-frame optimal camera pose matrix selection module are used to obtain an inter-frame optimal camera pose matrix. Based on the inter-frame optimal camera pose matrices of all adjacent two RGB image frames, a camera motion trajectory is obtained; an image acquisition module that acquires a reference frame RGB image and a current frame RGB image; a static image acquisition module that acquires a reference frame static image based on the reference frame RGB image and a current frame static image based on the current frame RGB image, where the reference frame static image / current frame static image does not include a priori dynamic object regions; a camera pose estimation matrix acquisition module that acquires a camera pose estimation matrix corresponding to each object region in the reference frame static image; a comprehensive projection error acquisition module that, for each camera pose estimation matrix, calculates the reprojection error of the object regions in the reference frame static image other than the object region corresponding to the camera pose estimation matrix on the current frame RGB image based on the camera pose estimation matrix, and fuses the reprojection errors of all object regions in the reference frame static image other than the object region corresponding to the camera pose estimation matrix to obtain the comprehensive projection error of the camera pose estimation matrix; an inter-frame optimal camera pose matrix selection module that selects the camera pose estimation matrix with the minimum comprehensive projection error as the inter-frame optimal camera pose matrix.
[0009] To achieve the above object of the present invention, according to the third aspect of the present invention, the present invention provides a computer-readable storage medium storing at least one instruction, at least one program, a code set, or an instruction set, which is loaded and executed by a processor to implement the semantic and multi-object error-minimized dynamic SLAM method as described in the first aspect of the present invention.
[0010] To achieve the above object of the present invention, according to the fourth aspect of the present invention, the present invention provides a mobile robot, including a mobile robot body, a vision device, and a processor provided on the mobile robot body. The processor receives consecutive frame RGB images captured by the vision device and executes the steps of the semantic and multi-object error-minimized dynamic SLAM method as described in the first aspect of the present invention.
[0011] The present invention is mainly applied to camera positioning and mapping in a dynamic environment. By referring to the static images of the reference frame and the current frame, the influence of absolutely dynamic objects on the accuracy of subsequent calculation of the camera pose matrix (i.e., positioning) is eliminated. The camera pose estimation matrix is calculated separately for each segmented object region, and the reprojection error of each camera pose estimation matrix in other object regions is obtained. The camera pose estimation matrix with the smallest comprehensive projection error is selected as the optimal camera pose matrix, further improving the accuracy of camera pose estimation and the accuracy of mapping, and obtaining a more accurate camera motion trajectory. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 FIG. 6 is a schematic flow chart of the semantic and multi-object error minimum dynamic SLAM method in Embodiment 1 of the present invention;
[0013] Figure 2 FIG. 10 is a schematic diagram for calculating the three-dimensional coordinates of feature points in the reference frame image in Embodiment 1 of the present invention;
[0014] Figure 3 FIG. 14 is a schematic diagram of the global process of an application scenario in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.
[0016] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0017] In the description of the present invention, unless otherwise specified and limited, it should be noted that the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the communication inside two elements. It can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific situations.
[0018] Embodiment 1
[0019] This embodiment discloses a semantic and multi-object error-minimized dynamic SLAM method. Continuous RGB image frames are acquired. For two adjacent RGB image frames, steps S1 to S5 are executed to obtain the optimal inter-frame camera pose matrix. Based on the optimal inter-frame camera pose matrices of all adjacent two-frame RGB images, the camera motion trajectory is obtained, as Figure 1 shown, where:
[0020] Step S1: Obtain a reference-frame RGB image and a current-frame RGB image. The reference-frame RGB image and the current-frame RGB image are two consecutive adjacent RGB image frames. Preferably but not limited to, the reference-frame RGB image is the previous RGB image of the current-frame RGB image.
[0021] Step S2: Obtain a reference-frame static image based on the reference-frame RGB image, and obtain a current-frame static image based on the current-frame RGB image. The reference-frame static image does not include the prior dynamic object regions, and the current-frame static image does not include the prior dynamic object regions. The prior dynamic object regions refer to the object regions that are confirmed as absolutely dynamic according to prior empirical knowledge. The reference-frame static image / current-frame static image includes prior static object regions and the background.
[0022] In this embodiment, specifically, the reference-frame static image is obtained by removing the prior dynamic object regions from the reference-frame RGB image, and the current-frame static image is obtained by removing the prior dynamic object regions from the current-frame RGB image.
[0023] Step S3: Obtain the camera pose estimation matrix corresponding to each object region in the reference-frame static image. Let j be the object region index, and j is a positive integer. Preferably, obtain the camera pose estimation matrix corresponding to the j-th object region of the reference-frame static image, which specifically includes:
[0024] Step S31: Select n feature points from the j-th object region and obtain the n corresponding matching feature points of the n feature points in the current-frame static image; n is a positive integer;
[0025] Step S32: Obtain the three-dimensional spatial coordinates of the n feature points in the j-th object region of the reference-frame static image. Preferably, when the vision device (such as the vision device on a mobile robot) is an RGBD camera or a binocular camera, the depth value of each feature point can be obtained through the depth map or the binocular depth measurement principle, and the three-dimensional spatial coordinates of the feature point can be obtained according to the depth value of the feature point and its two-dimensional coordinates in the reference-frame static image.
[0026] Step S33: Establish an objective function:
[0027] where, x i'Denote the two-dimensional coordinates of the matching feature point i' corresponding to the i-th feature point in the j-th object region of the reference frame static image in the current frame static image; R represents the rotation matrix, and t represents the translation matrix; X i Denote the three-dimensional spatial coordinates of the i-th feature point; π(·) represents the projection function for projecting onto the current frame static image; Denote the calculation of The camera pose estimation matrix T when reaching the minimum value, and use this camera pose estimation matrix as the optimal camera pose estimation matrix T * , R * Denote the optimal rotation matrix, t * Denote the optimal translation matrix; Denote the two-norm.
[0028] Step S34, solve the objective function to obtain the optimal camera pose estimation matrix T * As the camera pose estimation matrix corresponding to the j-th object region of the reference frame static image. Introduce Bundle Adjustment (bundle adjustment algorithm) to solve by minimizing the sum of the reprojection errors of all feature points in the j-th object region of the image frame. Preferably but not limited to using the existing Levenberg-Marquardt method to solve the objective function to obtain the optimal camera pose estimation matrix T * .
[0029] Step S4, for each camera pose estimation matrix, calculate the reprojection error of the object region in the reference frame static image other than the object region corresponding to this camera pose estimation matrix on the current frame RGB image based on this camera pose estimation matrix, and fuse the reprojection errors of all object regions in the reference frame static image other than the object region corresponding to this camera pose estimation matrix to obtain the comprehensive projection error of this camera pose estimation matrix. Calculate the comprehensive projection errors of all camera pose estimation matrices according to Step S4.
[0030] Step S5, select the camera pose estimation matrix with the minimum comprehensive projection error as the optimal inter-frame camera pose matrix.
[0031] In this embodiment, in order to automatically obtain the reference frame static image / current frame static image, in Step S2, the specific process of obtaining the reference frame static image / current frame static image includes:
[0032] Step S21: Input the reference frame RGB image / current frame RGB image into the trained instance segmentation network to obtain the object surface mask. The instance segmentation network preferably but not limited to uses the SOLOV2 network. The object surface mask includes a dynamic object area and a static object area identified by the instance segmentation network according to prior knowledge. The instance segmentation network preferably but not limited to uses a binary mask, such as pixels with a value of 0 in the binary mask are object pixels of semantic dynamics, and pixels with a value of 1 are object pixels of semantic non-dynamics. The object pixels of non-dynamics include static object pixels and the background.
[0033] Step S22: Use the object surface mask to remove the pixel points of the prior dynamic objects in the reference frame RGB image / current frame RGB image to obtain the reference frame static image / current frame static image. Specifically, for example, remove the pixels with a pixel value of 0 in the object surface mask, and number each object area in the area with a pixel value of 1.
[0034] In this embodiment, since the object information detected by the instance segmentation network is related to the training data set used in the early stage, there are objects that cannot be detected (such as due to insufficient prior knowledge) and relevant information of the background in actual applications. Then, all non-object areas identified by the remaining instance segmentation network are defined as background areas. To make up for the deficiencies of the instance segmentation network and avoid missing static object areas, which may affect the later SLAM positioning accuracy, further preferably, in step S3, the background areas in both the reference frame static image and the current frame static image are regarded as an object area and participate in the subsequent camera pose evaluation and positioning.
[0035] In this embodiment, preferably, the training process of the instance segmentation network includes:
[0036] (1) Construct a sample set. The samples are RGB images. Each RGB image sample corresponds to an object surface mask. The object surface mask can be obtained by manually prior identifying and segmenting the objects on the RGB image sample. Each object (which can be a static object and a dynamic object) on the object surface mask is set with a semantic label, and this semantic label is used to mark whether the object is dynamic or static. Divide the sample set into a training set, a validation set, and a test set.
[0037] (2) Construct an instance segmentation network, preferably but not limited to the SOLOV2 network. Use the training set to train the instance segmentation network, and use the validation set and the test set to verify and test the trained instance segmentation network respectively to obtain the trained instance segmentation network.
[0038] (3) Input the current frame RGB image into the trained instance segmentation network to obtain the object mask and the semantic label of each object on the object mask.
[0039] In this embodiment, in step S32, when the vision device for SLAM is a monocular camera, the three-dimensional spatial coordinates of the feature points are obtained through the fundamental matrix between frames. To improve the accuracy of the three-dimensional spatial coordinates of the feature points, the fundamental matrix between frames is initialized to ensure that the fundamental matrix between frames is calculated through the most static object region in the reference frame static image, so as to improve the accuracy of the fundamental matrix between frames and the accuracy of the subsequent camera pose estimation (positioning). Therefore, further preferably, the process of obtaining the three-dimensional spatial coordinates of the feature points in the reference frame static image includes:
[0040] Step S321, traverse all object regions in the reference frame static image to update the reference frame static image. Specifically: determine whether there is a matching object region in the current frame static image that has at least 8 pairs of matching feature point pairs with the object region. If so, retain the object region; if not, discard the object region. Specifically, let the fundamental matrix between frames be F. The F matrix describes the corresponding relationship between the coordinates of points in space in two planes. The size of the F matrix is a 3x3 matrix with 9 unknowns, that is, 9 degrees of freedom. However, due to its scale equivalence, the degree of freedom is 8. At least 8 pairs of matching feature points exist between two consecutive frames of images to calculate the fundamental matrix F between frames. Therefore, traverse each object region to determine whether there are 8 or more feature points in the object region.
[0041] Further preferably, when there are more than 8 pairs of matching feature point pairs, to filter out outliers and improve the robustness of the fundamental matrix between frames, the RANSAC algorithm is used to filter out outliers. Specifically, it includes: first, arbitrarily select 8 pairs of matching points in the object region, and the F matrix can be calculated by the existing eight-point method. Calculate the error of this F matrix through other feature points in the region, and iterate k times to take the F matrix of the model with the smallest error as the calculated value of this step. Similarly, substitute other objects, and finally obtain the initialized F matrix, where k is a positive integer.
[0042] Step S322, obtain the fundamental matrix between frames of each object region based on the matching feature point pairs of each object region in the updated reference frame static image, that is, solve the corresponding fundamental matrix F for each object region in turn.
[0043] Step S323, obtain the sum of the epipolar distances from all the matching feature points of each object region to the epipolar line L in the matching object region; specifically, as shown in, for example, the solution process for the first object region in the updated reference frame static image is: let the feature points x in the first object region M in the reference frame static image, and the corresponding matching object region in the first object region in the current frame static image F and the feature points x 2 of Figure 2 is shown. For example, for the solution process of the first object region in the updated reference frame static image: let the first object region M in the reference frame static image 1 inside the feature points x 1 , the current frame static image F c in the corresponding matching object region of the first object region and the feature points x1 The matching feature point that matches is x 1' , and the normalized coordinates corresponding to the feature point and the matching feature point are as follows:
[0044] x 1 = [u 1 v 1 1] T , x 1' = [u 1' v 1' 1] T ;
[0045] u 1 , v 1 respectively represent the normalized abscissa and ordinate of the feature point x 1 , and u 1' , v 1' respectively represent the normalized abscissa and ordinate of the matching feature point x 1' .
[0046] According to the epipolar constraint, we have: x 1' T Fx 1 = 0;
[0047] The epipolar line L 1 can be expressed as: L 1 = [X Y Z] T = Fx 1 = F[u 1 v 1 1] T ;
[0048] Then the epipolar distance from the matching feature point x 1' to its epipolar line L 2 is:
[0049]
[0050] Among them, X, Y, and Z respectively represent the coordinate values of the x-axis, y-axis, and z-axis of the three-dimensional space of the feature point.
[0051] Step S324, use the inter-frame fundamental matrix of the object region with the minimum sum of epipolar distances as the initialized inter-frame fundamental matrix;
[0052] Step S325, obtain the three-dimensional spatial coordinates of the feature points in the reference frame static image based on the initialized inter-frame fundamental matrix.
[0053] In this embodiment, to accurately calculate the reprojection error, further preferably, in step S4, the steps of obtaining the reprojection error of the object region other than the object region corresponding to the camera pose estimation matrix in the reference frame static image on the current frame RGB image based on the camera pose estimation matrix include:
[0054] Step S41, assume that the reprojection error of the j-th object region on the current frame RGB image is obtained based on the camera pose estimation matrix The camera pose estimation matrix is the camera pose estimation matrix not corresponding to the j-th object region, and the camera pose estimation matrix corresponding to the j-th object region can be expressed as j'≠j, both n and j are positive integers; assume that the j-th object region includes n feature points, and based on the camera pose estimation matrix the estimated matching feature point coordinates of the n feature points in the current frame RGB image are obtained respectively;
[0055] Step S42, project the estimated matching feature point coordinates and the actual matching feature point coordinates onto the normalization plane respectively to obtain the normalized estimated coordinates and the normalized actual coordinates, calculate the distance between the normalized estimated coordinates and the normalized actual coordinates of the matching feature points, and record the distance as the error distance of the matching feature points; assume that the normalized estimated coordinates and the normalized actual coordinates obtained for the i-th feature point are respectively expressed as and (u i , v i ), then the error distance of the i-th feature point is: || || represents taking the absolute value modulus length.
[0056] Step S43, calculate the average value of the error distances of all matching feature points within the j-th object region, and take this average value d_fin as the reprojection error of the j-th object region on the current frame RGB image based on the camera pose estimation matrix .
[0057] In this embodiment, in step S4, the preferred but not limited fusion method for obtaining the comprehensive projection error is to calculate the average value or the difference between the maximum and minimum values of the reprojection errors of all object regions other than the object region corresponding to the camera pose estimation matrix in the fused reference frame static image, etc.
[0058] In this embodiment, to more comprehensively evaluate all reprojection errors, further preferably, in step S4, in the step of obtaining the comprehensive projection error of the camera pose estimation matrix from the reprojection errors of all object regions in the fused reference frame static image except the object region corresponding to the camera pose estimation matrix, the comprehensive projection error is the sum of the reprojection errors of all object regions in the reference frame static image except the object region corresponding to the camera pose estimation matrix. Specifically, assume the camera pose estimation matrix corresponds to the j-th object region, then the sum of the reprojection errors of all object regions in the reference frame static image except the j-th object region corresponding to the camera pose estimation matrix is expressed as:
[0059] s j = d_fin 1 +... + d_fin j-1 + d_fin j+1 +... + d_fin J ;
[0060] J represents the total number of object regions in the reference frame static image, and s j represents the comprehensive projection error of the camera pose estimation matrix . Calculate the comprehensive projection errors of all camera pose estimation matrices according to the above process: s 1 , s 2 , s 3 ,..., s J . Select the camera pose estimation matrix corresponding to the minimum value from s 1 , s 2 , s 3 ,..., s J as the optimal camera pose estimation matrix. The camera pose can be obtained according to the optimal camera pose estimation matrix to achieve camera and robot positioning.
[0061] In an application scenario of this embodiment, as shown in Figure 3 , use the SOLOV2 network to obtain the object surface mask, that is, the original mask. The consecutive frame images are sequentially processed according to the above steps S1 to S5 to obtain a series of inter-frame optimal camera pose matrices, and then a complete camera motion trajectory curve is generated. Select the RGB image frames with better tracking effects as key frames in the consecutive image frames (preferably but not limited to the frames with the minimum comprehensive projection error in several consecutive frames, or the frames with large differences before and after), and input the key frames into the back end of the SLAM system to complete local mapping and loop detection and other links.
[0062] This embodiment provides a SLAM method that combines instance segmentation algorithm and multi-object reprojection error minimization, which can be used to improve the accuracy of system positioning and mapping in dynamic environments. The instance segmentation algorithm can perform instance segmentation on the input image sequence to obtain the regions occupied by each object in the image, and directly eliminate the object regions with prior dynamic semantics. In the actual scenario, based on the premise that each independent rigid body object hardly moves in the same direction relative to the camera orientation, calculations are performed separately for each of the remaining objects segmented by the instance segmentation. First, the pose matrix of an object is calculated through the feature points on one of the objects, and this pose matrix is substituted into other objects to calculate the error between the estimated matching points and the actual matching points. The error values of all objects are accumulated and normalized as the comprehensive projection error value of the pose matrix of this object. Each object is iterated in turn, and finally the camera pose matrix with the minimum comprehensive projection error is taken as the optimal inter-frame pose matrix to calculate the pose of the camera, thereby improving the accuracy of camera pose estimation and the accuracy of mapping.
[0063] Compared with the existing SLAM methods, this technical invention focuses on the problems of camera positioning and mapping in dynamic environments, and comprehensively analyzes the advantages and disadvantages of the SLAM system based on deep learning and the traditional SLAM system. The method of combining an instance segmentation network with multi-object calculation of the pose matrix is used to minimize the overall error finally to improve the accuracy of the SLAM method. While fully utilizing the instance segmentation results and eliminating the absolutely dynamic objects, this embodiment proposes to calculate the inter-frame pose estimation matrix separately using the remaining segmented objects, and substitute the matrix results into other objects and the background, and calculate the distance error between the estimated matching points and the actual matching points as a measurement parameter for the quality of this matrix. Each object is iterated in turn, and finally the optimal inter-frame pose estimation matrix is found.
[0064] Embodiment 2
[0065] This embodiment discloses a semantic and multi-object error-minimized dynamic SLAM system, including: a camera motion trajectory acquisition module that acquires continuous RGB image frames. For two adjacent RGB image frames, the image acquisition module, static image acquisition module, camera pose estimation matrix acquisition module, comprehensive projection error acquisition module, and inter-frame optimal camera pose matrix selection module are used to obtain the inter-frame optimal camera pose matrix, and the camera motion trajectory is obtained based on the inter-frame optimal camera pose matrices of all adjacent two-frame RGB images; an image acquisition module that acquires a reference frame RGB image and a current frame RGB image; a static image acquisition module that acquires a reference frame static image based on the reference frame RGB image and a current frame static image based on the current frame RGB image, where the reference frame static image / current frame static image does not include prior dynamic object regions; a camera pose estimation matrix acquisition module that acquires the camera pose estimation matrix corresponding to each object region in the reference frame static image; a comprehensive projection error acquisition module that, for each camera pose estimation matrix, calculates the reprojection error of the object regions in the reference frame static image other than the object region corresponding to the camera pose estimation matrix on the current frame RGB image based on the camera pose estimation matrix, and fuses the reprojection errors of all object regions in the reference frame static image other than the object region corresponding to the camera pose estimation matrix to obtain the comprehensive projection error of the camera pose estimation matrix; an inter-frame optimal camera pose matrix selection module that selects the camera pose estimation matrix with the minimum comprehensive projection error as the inter-frame optimal camera pose matrix.
[0066] In this embodiment, each module corresponds to the steps in Embodiment 1 and will not be elaborated here.
[0067] Embodiment 3
[0068] This embodiment discloses a computer-readable storage medium that stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the semantic and multi-object error-minimized dynamic SLAM method provided in Embodiment 1.
[0069] Embodiment 4
[0070] This embodiment discloses a mobile robot, including a mobile robot body, a vision device, and a processor provided on the mobile robot body. The processor receives continuous frame RGB images captured by the vision device and executes the steps of the semantic and multi-object error-minimized dynamic SLAM method provided in Embodiment 1.
[0071] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples.
[0072] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
Claims
1. Semantic and multi-object error-minimized dynamic SLAM method, Characterized in that, Continuous RGB image frames are acquired. For two adjacent RGB image frames, steps S1 to S5 are performed to obtain the optimal inter-frame camera pose matrix. Based on the optimal inter-frame camera pose matrices of all adjacent two-frame RGB image frames, the camera motion trajectory is obtained, where: Step S1, obtaining a reference frame RGB image and a current frame RGB image; Step S2, obtaining a reference frame static image based on the reference frame RGB image, and obtaining a current frame static image based on the current frame RGB image. The reference frame static image / current frame static image does not include prior dynamic object regions; Step S3, obtaining the camera pose estimation matrix corresponding to each object region in the reference frame static image; Step S4, for each camera pose estimation matrix, calculating the reprojection error of the object regions in the reference frame static image except the object region corresponding to the camera pose estimation matrix on the current frame RGB image based on the camera pose estimation matrix, and fusing the reprojection errors of all object regions in the reference frame static image except the object region corresponding to the camera pose estimation matrix to obtain the comprehensive projection error of the camera pose estimation matrix; Step S5, selecting the camera pose estimation matrix with the minimum comprehensive projection error as the optimal inter-frame camera pose matrix; In step S4, the step of calculating the reprojection error of the object regions in the reference frame static image except the object region corresponding to the camera pose estimation matrix on the current frame RGB image based on the camera pose estimation matrix includes: Let the reprojection error of the j-th object region based on the camera pose estimation matrix on the current frame RGB image, where j′≠j. Let the j-th object region include n feature points. Based on the camera pose estimation matrix respectively obtain the estimated matching feature point coordinates of the n feature points in the current frame RGB image; Projecting the estimated matching feature point coordinates and the actual matching feature point coordinates onto the normalization plane respectively to obtain the normalized estimated coordinates and the normalized actual coordinates, calculating the distance between the normalized estimated coordinates and the normalized actual coordinates of the matching feature points, and recording the distance as the error distance of the matching feature points; Obtain the average value of the error distances of all matching feature points, and use the average value as the reprojection error of the j-th object region on the current frame RGB image based on the camera pose estimation matrix ; The comprehensive projection error is the sum or average of the reprojection errors of all object regions in the reference frame static image except the object region corresponding to the camera pose estimation matrix.
2. The semantic and multi-object error-minimized dynamic SLAM method according to claim 1, Characterized in that, In step S2, the specific process of obtaining the reference frame static image / current frame static image includes: Inputting the reference frame RGB image / current frame RGB image into a trained instance segmentation network to obtain an object surface mask; Using the object surface mask to remove the pixel points of prior dynamic objects in the reference frame RGB image / current frame RGB image to obtain the reference frame static image / current frame static image.
3. The semantic and multi-object error-minimized dynamic SLAM method according to claim 1 or 2, Characterized in that, In step S3, in both the reference frame static image and the current frame static image, their respective background regions are regarded as one object region.
4. The semantic and multi-object error-minimized dynamic SLAM method according to claim 3, Characterized in that, In step S3, the specific process of obtaining the camera pose estimation matrix corresponding to the j-th object region in the reference frame static image includes: Step S31: Select n feature points from the j-th object region and obtain n corresponding matching feature points of the n feature points in the current-frame static image; where n is a positive integer; j is the object region index and is a positive integer. Step S32: Obtain the three-dimensional spatial coordinates of the n feature points in the j-th object region of the reference-frame static image. Step S33, establish the objective function: where x i′ represents the two-dimensional coordinates of the matching feature point i' corresponding to the i-th feature point in the j-th object region of the reference frame static image in the current frame static image; R represents the rotation matrix, and t represents the translation matrix; X i represents the three-dimensional spatial coordinates of the i-th feature point; π(·) represents the projection function for projecting onto the current frame static image; denotes the calculation of the camera pose estimation matrix T when reaching the minimum value, and taking the camera pose estimation matrix at this time as the optimal camera pose estimation matrix T * , R * represents the optimal rotation matrix, and t * represents the optimal translation matrix; Step S34, solve the objective function to obtain the optimal camera pose estimation matrix T * The camera pose estimation matrix corresponding to the j-th object region of the reference frame static image.
5. The semantic and multi-object error-minimized dynamic SLAM method according to claim 4, characterized in that when the vision device for SLAM is a monocular camera, the process of obtaining the three-dimensional spatial coordinates of the feature points in the reference-frame static image includes: Traverse all object regions in the reference-frame static image to update the reference-frame static image, specifically: Determine whether there is a matching object region in the current-frame static image that has at least 8 pairs of matching feature point pairs with the object region. If so, retain the object region; if not, discard the object region. Obtain the inter-frame fundamental matrix of each object region based on the matching feature point pairs of each object region in the updated reference-frame static image. Based on the fundamental matrix between frames for each object region, obtain the sum of the epipolar distances from all the matching feature points in the matching object region of the object region to the epipolar line L 2 ; Take the inter-frame fundamental matrix of the object region with the minimum sum of epipolar distances as the initial inter-frame fundamental matrix. Obtain the three-dimensional spatial coordinates of the feature points in the reference-frame static image based on the initial inter-frame fundamental matrix.
6. A semantic and multi-object error-minimized dynamic SLAM system, characterized in that it includes: A camera motion trajectory acquisition module that acquires continuous RGB image frames. For adjacent two-frame RGB images, input image acquisition module, static image acquisition module, camera pose estimation matrix acquisition module, comprehensive projection error acquisition module, and inter-frame optimal camera pose matrix selection module obtain the inter-frame optimal camera pose matrix, and obtain the camera motion trajectory based on the inter-frame optimal camera pose matrices of all adjacent two-frame RGB images; An image acquisition module that acquires a reference-frame RGB image and a current-frame RGB image; A static image acquisition module that acquires a reference-frame static image based on the reference-frame RGB image and acquires a current-frame static image based on the current-frame RGB image, where the reference-frame static image / current-frame static image does not include prior dynamic object regions; A camera pose estimation matrix acquisition module that acquires the camera pose estimation matrix corresponding to each object region in the reference-frame static image; A comprehensive projection error acquisition module that, for each camera pose estimation matrix, obtains the reprojection error of the object regions in the reference-frame static image other than the object region corresponding to the camera pose estimation matrix on the current-frame RGB image based on the camera pose estimation matrix, and fuses the reprojection errors of all object regions in the reference-frame static image other than the object region corresponding to the camera pose estimation matrix to obtain the comprehensive projection error of the camera pose estimation matrix; An inter-frame optimal camera pose matrix selection module that selects the camera pose estimation matrix with the minimum comprehensive projection error as the inter-frame optimal camera pose matrix; In the comprehensive projection error acquisition module, the step of obtaining the reprojection error of the object regions in the reference-frame static image other than the object region corresponding to the camera pose estimation matrix on the current-frame RGB image based on the camera pose estimation matrix includes: Let the reprojection error of the j-th object region on the current frame RGB image be obtained based on the camera pose estimation matrix where j′≠j. Let the j-th object region include n feature points. Based on the camera pose estimation matrix respectively obtain the estimated matching feature point coordinates of the n feature points in the current frame RGB image; The estimated matching feature point coordinates and the actual matching feature point coordinates are respectively projected on the normalized plane to obtain the normalized estimated coordinates and the normalized actual coordinates, and the distance between the normalized estimated coordinates and the normalized actual coordinates of the matching feature points is calculated, and the distance is recorded as the error distance of the matching feature points; Calculate the average of the error distances of all matching feature points, and use the average value as the reprojection error of the j-th object region on the current frame RGB image based on the camera pose estimation matrix ; The comprehensive projection error is the sum or average of the reprojection errors of all object regions in the reference frame static image except the object region corresponding to the camera pose estimation matrix.
7. A computer-readable storage medium storing at least one instruction, at least one program, a code set or an instruction set, which is loaded and executed by a processor to implement the semantic and multi-object error minimum dynamic SLAM method according to any one of claims 1-5.
8. A mobile robot, characterized in that, it includes a mobile robot body, a vision device and a processor provided on the mobile robot body, the processor receives consecutive frame RGB images captured by the vision device, and executes the steps of the semantic and multi-object error minimum dynamic SLAM method according to any one of claims 1-5.
Citation Information
Patent Citations
Joint optimization dynamic SLAM method based on semantics and geometry
CN112308921A
Dynamic environment camera pose estimation and semantic map construction method based on semantic SLAM
CN111402336A
Visual odometer implementation method based on semi-direct method
CN113592947A