Simultaneous Localization and Mapping Method and Unmanned Mobile Device
By introducing ground parameters and multiple target three-dimensional points into the SLAM system, and optimizing position and ground parameters is divided into two optimization stages, the problem of lack of ground constraints in positioning and map construction in the existing SLAM system is solved, and a higher accuracy of positioning and map construction is achieved.
Patent Information
- Application Number
- CN202210316869.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-03-28
AI Technical Summary
The existing SLAM system lacks ground constraints in the positioning and map construction process, resulting in large map errors and it is difficult to accurately reflect the high-level information and geometric structure of the environment.
By introducing ground parameters and multiple target three-dimensional points in the SLAM system, it is divided into two optimization stages: the first stage fixes the ground parameters and optimizes the pose parameters of the keyframe; the second stage fixes the pose parameters of the keyframes and optimizes the ground parameters. Optimization is performed using a cost function, including reprojection error and distance error.
By introducing ground constraints and phased optimization, map errors can be significantly reduced, positioning and map construction accuracy can be improved, and the three-dimensional points on the ground are closer to the actual ground.
Smart Images

Figure CN114858156B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of location-based services, and in particular, to a method for instant positioning and map construction and an unmanned mobile device. Background Art
[0002] Positioning technologies have been widely applied in various industries, such as robots, autonomous driving, AR, VR, etc. There are also many sensors used for positioning. Monocular cameras have attracted the attention of a large number of researchers and practitioners due to their low cost, small size, and easy installation on various platforms. Currently, the mainstream monocular SLAM (Simultaneous Localization and Mapping) systems mainly use points as the only map landmarks. For example, ORB-SLAM and DSO technologies. However, points in the map cannot reflect the high-level information and geometric structure in the scene. Extracting ground information from the scanned images of the surrounding environment, and then adding the corresponding ground information during map construction, and combining ground constraints can make the estimation error smaller.
[0003] Therefore, how to add ground constraints during the SLAM mapping process to reduce the error of the constructed map is one of the technical problems that need to be solved currently. Summary of the Invention
[0004] Embodiments of the present disclosure provide a method for instant positioning and map construction and an unmanned mobile device.
[0005] In a first aspect, embodiments of the present disclosure provide a method for instant positioning and map construction, which includes:
[0006] Obtain a plurality of consecutive key frames and a plurality of target three-dimensional points that can be observed in all of the plurality of consecutive key frames;
[0007] Use the consecutive key frames and the target three-dimensional points to perform local map optimization in a first stage; wherein, during the local map optimization in the first stage, after fixing the ground parameters in the distance error to an initial value, optimize the pose parameters of the consecutive key frames.
[0008] Use the ground three-dimensional points in the consecutive key frames and the target three-dimensional points to perform plane optimization in a second stage; during the plane optimization in the second stage, after fixing the pose parameters of the consecutive key frames to the optimization results obtained in the first stage, optimize the ground parameters in the consecutive key frames.
[0009] Further, in both the local map optimization in the first stage and the local map optimization in the second stage, a cost function is used for optimization. The cost function includes the reprojection error of the target 3D points on multiple consecutive key frames and the distance error between the ground 3D points in the target 3D points and the ground. The reprojection error is obtained based on the target 3D points and the pose parameters of the multiple consecutive key frames. The distance error is obtained based on the ground 3D points and the ground parameters.
[0010] Further, obtaining multiple consecutive key frames includes:
[0011] Obtaining the current key frame, the first historical key frame, and multiple consecutive historical key frames between the current key frame and the first historical key frame;
[0012] Wherein, in the local map optimization in the first stage, the pose parameters corresponding to the first historical key frame remain fixed, while the pose parameters of the current key frame and the multiple historical key frames between the current key frame and the first historical key frame are parameters to be optimized.
[0013] Further, the ground parameters include a ground normal vector parameter and a distance parameter from the ground to the origin of the world coordinate system.
[0014] Further, the method further includes:
[0015] Extracting feature points in the current key frame;
[0016] Determining the current 3D points corresponding to the feature points;
[0017] Based on the method of fitting a plane with the current 3D points, determining the ground 3D points among the current 3D points that are located on the ground.
[0018] Further, based on the method of fitting a plane with the current 3D points, determining the ground 3D points among the current 3D points that are located on the ground includes:
[0019] Fitting one or more planes based on the current 3D points;
[0020] Determining the target plane corresponding to the ground among the one or more planes;
[0021] Determining the distance between the current 3D points and the target plane;
[0022] Based on the distance, screening out the ground 3D points from the current 3D points.
[0023] Further, in the local map optimization in the first stage of the current optimization process, the initial value corresponding to the ground parameters is the optimization result obtained in the plane optimization process in the second stage of the previous optimization process.
[0024] In a second aspect, an unmanned mobile device is provided in an embodiment of the present disclosure, including: an image sensor and a processor; wherein,
[0025] The image sensor collects an image of the surrounding environment of the unmanned mobile device and outputs the image to the processor;
[0026] The processor obtains a plurality of consecutive key frames in the image and a plurality of target three-dimensional points that can be observed in all of the plurality of consecutive key frames;
[0027] The processor further performs a first-stage local map optimization using the consecutive key frames and the target three-dimensional points, and performs a second-stage plane optimization using the ground three-dimensional points among the consecutive key frames and the target three-dimensional points; wherein, in the first-stage local map optimization process, after fixing the ground parameters in the distance error to an initial value, the pose parameters of the consecutive key frames are optimized; in the second-stage plane optimization process, after fixing the pose parameters of the consecutive key frames to the optimization result obtained in the first-stage optimization, the ground parameters in the consecutive key frames are optimized.
[0028] In a third aspect, a method for providing a location-based service is provided in an embodiment of the present disclosure, including: performing backend optimization using the method described in the first aspect, constructing a surrounding environment map based on the result of the backend optimization, and providing a location-based service for the service recipient based on the surrounding environment map, where the location-based service includes one or more of navigation, route planning, and map rendering.
[0029] In a fourth aspect, an instant positioning map construction device is provided in an embodiment of the present invention, including:
[0030] An acquisition module configured to acquire a plurality of consecutive key frames and a plurality of target three-dimensional points that can be observed in all of the plurality of consecutive key frames;
[0031] A first optimization module configured to perform a first-stage local map optimization using the consecutive key frames and the target three-dimensional points; wherein, in the first-stage local map optimization process, after fixing the ground parameters in the distance error to an initial value, the pose parameters of the consecutive key frames are optimized;
[0032] A second optimization module configured to perform a second-stage plane optimization using the ground three-dimensional points among the consecutive key frames and the target three-dimensional points; in the second-stage plane optimization process, after fixing the pose parameters of the consecutive key frames to the optimization result obtained in the first-stage optimization, the ground parameters in the consecutive key frames are optimized.
[0033] In a fifth aspect, an embodiment of the present invention provides a location-based service providing device, including: performing backend optimization using the above-mentioned instant positioning and mapping device, constructing a surrounding environment map based on the result of the backend optimization, and providing location-based services for the service recipient based on the surrounding environment map, where the location-based services include one or more of navigation, route planning, and map rendering.
[0034] The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0035] In a possible design, the structure of the above device includes a memory and a processor. The memory is used to store one or more computer instructions that support the above device in executing the corresponding method, and the processor is configured to execute the computer instructions stored in the memory. The above device may further include a communication interface for the device to communicate with other devices or communication networks.
[0036] In a sixth aspect, an embodiment of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the method described in any of the above aspects.
[0037] In a seventh aspect, an embodiment of the present disclosure provides a computer-readable storage medium for storing the computer instructions used by any of the above devices, and when the computer instructions are executed by a processor, they are used to implement the method described in any of the above aspects.
[0038] In an eighth aspect, an embodiment of the present disclosure provides a computer program product that includes computer instructions, and when the computer instructions are executed by a processor, they are used to implement the method described in any of the above aspects.
[0039] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0040] In the SLAM (Simultaneous Localization and Mapping) in the embodiments of the present disclosure, during the back-end optimization process, the ground parameters and multiple target three-dimensional points are used to optimize the current frame and multiple consecutive key frames before the current frame. The optimization process is divided into two stages. In the first stage, after fixing the ground parameters, the pose parameters of multiple consecutive key frames are optimized. In the second stage, based on the optimization in the first stage, after fixing the pose parameters of multiple consecutive key frames, the ground parameters are optimized. Finally, the pose parameters and ground parameters corresponding to each optimized key frame can be obtained. In the SLAM process of the embodiments of the present disclosure, by introducing the constraint of the ground parameters, the scale of the SLAM system is constrained in such a way that the ground three-dimensional points approach the ground during the back-end optimization process. And after introducing the ground constraint, the optimization is carried out in two stages. In the first stage of optimization, a consistent scale can be maintained in the front and back optimizations, while in the second stage of optimization, the ground satisfies this consistent scale. Finally, a local map optimization result with relatively high accuracy can be obtained.
[0041] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Combined with the drawings, through the following detailed description of non-limiting embodiments, other features, objects, and advantages of the present disclosure will become more obvious. In the drawings:
[0043] Figure 1 The flowchart of the simultaneous localization and mapping method according to an embodiment of the present disclosure is shown;
[0044] Figure 2 The schematic diagram of the reprojection error effect according to an embodiment of the present disclosure is shown;
[0045] Figure 3 The schematic diagram of the distance error effect from a point to the ground according to an embodiment of the present disclosure is shown;
[0046] Figure 4 The structural block diagram of an unmanned mobile device according to an embodiment of the present disclosure is shown;
[0047] Figure 5 The schematic diagram of a scene of the simultaneous localization and mapping method according to an embodiment of the present disclosure is shown;
[0048] Figure 6 The structural block diagram of the simultaneous localization and mapping device according to an embodiment of the present disclosure is shown;
[0049] Figure 7 It is a schematic diagram of the structure of an electronic device suitable for implementing the simultaneous localization and mapping method and / or the location-based service providing method according to the embodiments of the present disclosure. Detailed Implementation Modes
[0050] In the following, exemplary implementation modes of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement them. In addition, for clarity, parts unrelated to the description of the exemplary implementation modes are omitted in the drawings.
[0051] In the present disclosure, it should be understood that terms such as "including" or "having" are intended to indicate the presence of features, numbers, steps, actions, components, parts, or combinations thereof disclosed in this specification, and do not exclude the possibility of the presence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0052] In addition, it should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The present disclosure will be described in detail below with reference to the drawings and in combination with the embodiments.
[0053] Details of the embodiments of the present disclosure will be introduced in detail below through specific examples.
[0054] Figure 1 A flowchart showing an instant positioning and mapping method according to an embodiment of the present disclosure is shown. As Figure 1 shown, the instant positioning and mapping method includes the following steps:
[0055] In step S101, a plurality of consecutive key frames and a plurality of target three-dimensional points that can be observed in all of the plurality of consecutive key frames are acquired;
[0056] In step S102, local map optimization in a first stage is performed using the consecutive key frames and the target three-dimensional points; wherein, in the process of local map optimization in the first stage, after fixing the ground parameters in the distance error to an initial value, the pose parameters of the consecutive key frames are optimized;
[0057] In step S103, plane optimization in a second stage is performed using the consecutive key frames and the ground three-dimensional points among the target three-dimensional points; in the process of plane optimization in the second stage, after fixing the pose parameters of the consecutive key frames to the optimization result obtained in the first stage optimization, the ground parameters in the consecutive key frames are optimized.
[0058] In this embodiment, the instant positioning and mapping method can be executed on a terminal. The terminal can be any device, such as a mobile device like a drone, an autonomous vehicle, a robot, etc. An image sensor can be provided on the terminal for collecting images of the surrounding environment during movement. The image sensor provided on the terminal in the embodiments of the present disclosure can be a monocular camera.
[0059] Simultaneous Localization and Mapping (SLAM) technology refers to a technology where a mobile device (such as a robot, drone, mobile phone, car, smart wearable device, etc.) starts from an unknown location in an unknown environment, and during the movement, it observes and locates its own position and attitude through sensors (such as cameras, lidar, IMU, etc.), and then performs incremental map construction based on its own pose, so as to achieve the purpose of simultaneous localization and map construction.
[0060] If the sensors used in the SLAM process are vision-related sensors (such as monocular, binocular, RGB-D, fisheye, panoramic cameras, etc.), it is usually called visual SLAM.
[0061] The process of visual SLAM includes image acquisition, front-end localization, and back-end optimization, etc. The map construction part in SLAM mainly refers to the back-end optimization process, and the embodiments of the present disclosure mainly relate to the back-end optimization process.
[0062] In the embodiments of the present disclosure, after the mobile device enters an unknown environment, it can use the image sensor set on the mobile device to collect images in the surrounding environment in real time. The images are transmitted by the image sensor to the processing device of the mobile device, and the processing device performs front-end processing on them. During the front-end processing, feature points can be extracted from the current image, and then the feature points are matched with the previous image, and the initial pose of the current image relative to the previous image is determined based on the matching result. Information such as the current image and the initial pose can be transmitted to the back-end thread for back-end optimization. The front-end optimization can be a real-time processing process, while the back-end optimization may not be a real-time processing process.
[0063] For the first frame of image collected by the image sensor, the acquisition position of the first frame of image can be directly used as the origin of the map coordinate system to be constructed, and the feature points extracted from the first frame of image, etc. are stored as the information of the initial map. The images collected subsequently can continuously update the initial map.
[0064] In the embodiments of the present disclosure, multiple consecutive key frames can be obtained from the front-end processing process. The multiple consecutive key frames can be multiple consecutive key frames with a fixed number. The back-end optimization process can be an iterative loop optimization process. Each iterative optimization is for a fixed number of key frames. The fixed number of key frames can be regarded as multiple consecutive key frames in a sliding window. After obtaining a new key frame, that is, the current key frame, the sliding window is moved one position towards the current key frame, that is, the current key frame is added to the sliding window, and a historical key frame far from the current key frame in the original sliding window is removed. The multiple consecutive key frames in the sliding window with the current key frame added can jointly complete the current iterative optimization.
[0065] Suppose the current key frame is the nth key frame. If 10 consecutive key frames are used for each local optimization, that is, the length of the sliding window is 10, then the nth key frame and the previous historical key frames n-1, n-2, ……, n-9 can be jointly used for back-end optimization. It can be understood that in the current iterative optimization process, the initial pose of the current key frame is the initial pose output in the front-end processing, while the initial poses of other historical key frames can be the results optimized in the previous iterative optimization process.
[0066] In some embodiments, multiple consecutive key frames can be historical key frames before the current key frame and are temporally continuous with the current key frame, that is, the key frames within the sliding window in each round of iterative optimization are consecutive key frames. The target 3D points can be multiple map points to be tracked, and the multiple target 3D points can be observed in the multiple consecutive key frames within the sliding window, that is, the multiple target 3D points each have corresponding image points in the multiple consecutive key frames. It should be noted that in each iterative optimization process, the target 3D points may not be exactly the same. That is, after the current key frame is added, the set of target 3D points to be tracked can be updated based on the co-visibility relationship between the current key frame and multiple previous historical key frames.
[0067] For example, in the previous optimization process, the key frames used are the n-10th to the n-1st consecutive key frames, and the target 3D points observed in these 10 consecutive key frames include P1 to Px. In the current optimization process, the key frames used are the n-9th to the nth consecutive key frames (where the nth key frame is the current key frame). Then the target 3D points observed in these 10 consecutive key frames may include P2 to Px, Py. That is, the 3D point Px cannot be observed in the current key frame, while the new 3D point Py is observed, and the 3D point Py can also be observed in the n-9th to the nth consecutive key frames.
[0068] In the embodiments of the present disclosure, the back-end optimization is divided into two stages: the first stage is local map optimization, and the second stage is plane optimization. In the prior art, the back-end optimization only performs one-step optimization, that is, all parameters to be optimized are jointly optimized using a cost function. The all parameters to be optimized can include but are not limited to the pose parameters of key frames, ground parameters, etc. In the local map optimization of the first stage in the embodiments of the present disclosure, the pose parameters of multiple consecutive key frames are used as the parameters to be optimized, while the ground parameters are fixed to the initial values and not optimized; in the plane optimization of the second stage, the pose parameters of multiple consecutive key frames are fixed to the optimization results of the first stage, and the ground parameters are optimized. After two-step optimization, the pose parameters and ground parameters corresponding to each key frame can be obtained, and the pose parameters and ground parameters can be used as the initial values for the next iterative optimization process.
[0069] It should be noted that the pose parameters of the key frame include the rotation parameter and the translation parameter of the key frame relative to the world coordinate system; the ground parameters include the representation parameters of the ground in the world coordinate system. There are various ways to represent the ground in the world coordinate system, and the corresponding ground parameters are different for different representation methods. For example, in the nearest point representation method, the ground parameters include the ground normal vector parameter and the distance parameter between the ground and the origin of the world coordinate system.
[0070] In the SLAM (Simultaneous Localization and Mapping) of the embodiments of the present disclosure, during the back-end optimization process, the ground parameters and multiple target three-dimensional points are used to optimize the current frame and multiple consecutive key frames before the current frame. The optimization process is divided into two stages; in the first stage, after fixing the ground parameters, the pose parameters of multiple consecutive key frames are optimized. In the second stage, based on the optimization results of the first stage, after fixing the pose parameters of multiple consecutive key frames, the ground parameters are optimized. Finally, the pose parameters and ground parameters corresponding to each optimized key frame can be obtained. In the SLAM process of the embodiments of the present disclosure, by introducing the constraint of the ground parameters, the scale of the SLAM system is constrained by the way that the ground three-dimensional points approach the ground during the back-end optimization process. And after introducing the ground constraint, the optimization is carried out in two stages. In the first stage of optimization, a consistent scale can be maintained in the front and back optimizations, and in the second stage of optimization, the ground satisfies this consistent scale. Finally, a higher-accuracy optimization result can be obtained, and the accuracy of the SLAM map constructed based on this higher-accuracy optimization result is higher.
[0071] In an optional implementation manner of this embodiment, a cost function is used for optimization in both the local map optimization in the first stage and the local map optimization in the second stage. The cost function includes the reprojection error of the target three-dimensional points on multiple consecutive key frames and the distance error between the ground three-dimensional points in the target three-dimensional points and the ground; the reprojection error is obtained based on the target three-dimensional points and the pose parameters of the multiple consecutive key frames; the distance error is obtained based on the ground three-dimensional points and the ground parameters.
[0072] In this optional implementation manner, the local map optimization in the first stage and the plane optimization in the second stage are both performed using a pre-constructed cost function. The cost function can include two parts: the reprojection error and the distance error; the reprojection error is the error between the observed image point of the target three-dimensional point on the key frame and the projected image point obtained by projecting the target three-dimensional point onto the key frame; the distance error is the distance between the ground three-dimensional point and the ground (theoretically, the distance from this ground three-dimensional point to the ground is 0).
[0073] The construction processes of the cost function, the reprojection error, and the distance error are illustrated by examples below.
[0074] In some embodiments, the cost function can be expressed as follows:
[0075]
[0076] Wherein,
[0077]
[0078]
[0079]
[0080] Wherein, C represents the cost function for local map optimization. During the optimization process, the parameters to be optimized, such as pose parameters R iw , t iw , and ground parameter π w .
[0081] F represents a set of multiple consecutive key frames within the sliding window and the set of target 3D points that can be jointly observed by the multiple consecutive key frames.
[0082] G represents the set of ground 3D points within the sliding window and the set of initial observation frames where these ground 3D points first appear. It should be noted that G is a subset of F, that is, the ground 3D points in G are the target 3D points on the ground in F, and the initial observation frames in G are the key frames in F where a certain or certain ground 3D points first appear.
[0083] ei ,j represents the reprojection error formed by the i-th key frame and the j-th target 3D point in the set F.
[0084] represents the distance error formed by the q-th ground 3D point in the set G and the anchor p -th initial observation frame where the q-th ground 3D point first appears.
[0085] Ω represents the covariance matrix.
[0086] ρ h (·) represents the robust kernel function.
[0087] zi ,j represents the image point observed for the j-th target 3D point in the i-th key frame in the set F.
[0088] P j wRepresents the coordinates of the jth target 3D point in the set F in the world coordinate system. It should be noted that the world coordinate system is usually established with the initial position of the mobile device when it enters an unknown environment and collects images as the initial map as the origin, and the current rotation parameters of the image sensor on the mobile device as the rotation parameters relative to the world coordinate system, and the translation parameters as 0.
[0089] R iw Represents the rotation parameter from the camera coordinate system to the world coordinate system of the i-th key frame in the set F.
[0090] t iw Represents the translation parameter from the camera coordinate system to the world coordinate system of the i-th key frame in the set F.
[0091] c i Represents the projection point of the j-th target 3D point in the set F in the i-th key frame.
[0092] Represents the coordinates of the qth ground 3D point in the set G in the world coordinate system.
[0093] Represents the anchor in set G p The ground parameters in the camera coordinate system of the initial observation frame.
[0094] π w Represents the ground parameters in the world coordinate system.
[0095] Represents the anchor in set G p The rotation parameters of the camera coordinate system and the world coordinate system of the initial observation frame.
[0096] Represents the anchor in set G p The translation parameters of the camera coordinate system and the world coordinate system of the initial observation frame.
[0097] Figure 2 FIG. 2 is a schematic diagram showing the effect of reprojection error according to an embodiment of the present disclosure. Figure 2 As shown, the three-dimensional point in the world coordinate system After the parameter to be optimized R iw and t iw After the transformation, the coordinate c is projected onto the camera imaging plane i and the observed image point z i,j They do not overlap (theoretically they should overlap), which results in a reprojection error.
[0098] Figure 3 FIG. 2 is a schematic diagram showing the effect of the distance error from a point to the ground according to an embodiment of the present disclosure.Figure 3 As shown, the three-dimensional ground points should theoretically be on the ground plane π w However, due to inaccurate estimated variables, there is a certain distance error between them and the ground, thus constructing the distance error from the points to the ground.
[0099] In the embodiment of the present disclosure, the local map optimization in the first stage is actually the optimization of the above cost function. The optimization objective is to obtain the optimized result of the pose parameters corresponding to the key frames when the value of the cost function C is minimized. The pose parameters include the rotation parameter and the translation parameter from the camera coordinate system of the key frame to the world coordinate system.
[0100] During the local map optimization in the first stage, the ground parameter π in the distance error from the points to the ground w can be set as a fixed initial parameter, which can be the ground parameter obtained in the previous optimization. In the local map optimization in the first stage, it is not optimized; while the pose parameters will change continuously during the optimization process until the pose parameters that can minimize the cost function C are finally obtained.
[0101] After the local map optimization in the first stage is completed, the pose parameters of each optimized key frame can be obtained; in the plane optimization in the second stage, the pose parameters obtained during the local map optimization in the first stage can be fixed, and the ground parameter π in the cost function is optimized w . That is to say, in the plane optimization in the second stage, the pose parameters in the cost function are the optimized parameters obtained after the local map optimization in the first stage and are fixed, while the ground parameter in the distance error will change continuously during the optimization process until the ground parameter that can minimize the cost function C is finally obtained.
[0102] After the optimization in the first stage and the second stage, the pose parameters and the ground parameter of multiple consecutive key frames participating in this optimization process can be obtained.
[0103] In an alternative implementation manner of this embodiment, step S101, that is, the step of obtaining multiple consecutive key frames, further includes the following steps:
[0104] Obtain the current key frame, the first historical key frame, and multiple consecutive historical key frames between the current key frame and the first historical key frame;
[0105] Among them, in the local map optimization in the first stage, the pose parameters corresponding to the first historical key frame are fixed, while the pose parameters of the current key frame and multiple historical key frames between the current key frame and the first historical key frame are parameters to be optimized.
[0106] In this alternative implementation, as described above, the current iterative optimization targets multiple consecutive key frames including the current key frame, while the previous iterative optimization targets multiple consecutive key frames excluding the current key frame and a historical key frame between these multiple consecutive key frames. To achieve better optimization results, that is, to ensure that the scale of the optimization results for the key frames targeted by the current iterative optimization is consistent with that of the key frames targeted by the previous iterative optimization (such as the measurement scale for distance), during the current iterative optimization, in addition to the multiple consecutive key frames including the current key frame within the sliding window, one or more historical key frames outside the sliding window can be added. These one or more historical key frames can be referred to as the first historical key frames, and the optimized pose parameters were obtained for these first historical key frames during the previous iterative optimization. Therefore, during the current iterative optimization, these first historical key frames are added, and the pose parameters of these first historical key frames are fixed values, that is, the results of the previous iterative optimization, which are used to constrain the optimization of the pose parameters of other key frames.
[0107] In an alternative implementation of this embodiment, the ground parameters include a ground normal vector parameter and a distance parameter from the ground to the origin of the world coordinate system.
[0108] In this alternative implementation, there are various ways to represent a plane, such as the Hesse form, spherical coordinate form, tangent plane form, and closest point form. In the SLAM map construction of the embodiments of the present disclosure, it is necessary to introduce the ground as a constraint, and the ground, as a plane, can be represented by one of the above existing forms during the optimization process. However, considering that there are various problems when using the Hesse form, spherical coordinate form, and tangent plane form. For example, the Hesse form may over-parameterize the optimization results, the spherical coordinate form may produce singularities in certain situations, and the tangent plane form only fine-tunes the normal vector and requires recalculating the basis vectors for each iterative optimization; while the closest point form has no obvious drawbacks and is relatively intuitive when applied to the back-end optimization process of the embodiments of the present disclosure. Therefore, when constructing the cost function, the ground parameters are represented by the closest point form, and the distance error is also constructed based on the distance between the three-dimensional points on the ground and the ground.
[0109] The ground parameters in the closest point form are represented as follows:
[0110]
[0111] where π represents the ground, is the normal vector parameter of the ground, and d is the distance from the ground to the origin of the world coordinate system. It should be noted that the world coordinate system is usually established with the initial position where the mobile device is located when collecting the image as the initial map after entering an unknown environment as the origin, and with the current rotation parameter of the image sensor on the mobile device as the rotation parameter relative to the world coordinate system and the translation parameter being 0.
[0112] The distance from the ground to the origin of the world coordinate system can be understood as the vertical distance from the origin to the ground.
[0113] In an alternative implementation of this embodiment, the method further includes the following steps:
[0114] Extract the feature points in the current key frame;
[0115] Determine the current 3D points corresponding to the feature points;
[0116] Based on the method of fitting a plane from the current 3D points, determine the ground 3D points among the current 3D points that are located on the ground.
[0117] In this alternative implementation, in order to introduce the ground constraint into the pose optimization process of the key frame, after obtaining the current key frame, the current key frame is processed for feature point extraction, and after triangulation and other processing of the extracted feature points, the 3D points corresponding to the feature points are determined. These 3D points can be understood as the points on the constructed map corresponding to the feature points in the current key frame, and are also the real points in the surrounding environment. The 3D points corresponding to the feature points extracted from the key frame are usually called points to be tracked, which are used for feature matching in each key frame, and then to construct the SLAM map.
[0118] After determining the 3D points corresponding to the feature points in the current key frame, one or more planes can be fitted based on these 3D points. By comparing the relationship between the normal vector of this plane and the normal vector of the ground determined in the previous key frame, it is possible to determine whether there is a plane corresponding to the ground among the one or more planes, and the 3D points on the plane corresponding to the ground and the 3D points closer to the plane corresponding to the ground can be determined as the ground 3D points.
[0119] In an alternative implementation of this embodiment, the step of determining the ground 3D points among the current 3D points based on the method of fitting a plane from the current 3D points further includes the following steps:
[0120] Fit one or more planes based on the current 3D points;
[0121] Determine the target plane corresponding to the ground among the one or more planes;
[0122] Determine the distance between the current three-dimensional point and the target plane;
[0123] Based on the distance, filter out the ground three-dimensional points from the current three-dimensional points.
[0124] In this optional implementation, during the plane optimization process in the second stage, since the optimization only targets the ground parameters, and the constraint of the ground parameters in the cost function is based on the distance error from the ground three-dimensional points theoretically located on the ground to the ground. Therefore, for the newly obtained current key frame, feature points can be extracted from it, and the three-dimensional points corresponding to these feature points can be obtained through matching, and the plane formed by these three-dimensional points can be obtained by fitting a plane based on the three-dimensional points. Then, based on the relationship between the normal vector of the known ground parameters (i.e., the ground parameters optimized using the previous key frames) in the world coordinate system and the normal vector of the fitted plane, it can be determined whether there is a target plane corresponding to the ground in the fitted plane. For example, a plane parallel to the normal vector in the known ground parameters can be determined as the plane corresponding to the ground.
[0125] After determining the plane corresponding to the ground in the current key frame, ground three-dimensional points can also be filtered out from the current three-dimensional points extracted and processed from the current key frame.
[0126] In some embodiments, the distance between the current three-dimensional point and the target plane can be calculated. When the distance is less than or equal to a preset distance threshold, it can be considered that the current three-dimensional point may be a ground three-dimensional point, but there is a certain distance from the ground due to errors, and this distance is small enough. Therefore, such three-dimensional points can be determined as ground three-dimensional points, so that during the plane optimization process in the second stage, the ground three-dimensional points can be used to optimize the ground parameters.
[0127] In an optional implementation of this embodiment, in the local map optimization in the first stage of the current optimization process, the initial value corresponding to the ground parameters is the optimization result obtained in the plane optimization process in the second stage of the previous optimization process.
[0128] In this optional implementation, in the local map optimization in the first stage, for the part of the distance error calculated based on the ground parameters in the cost function, the ground parameters are fixed as the initial value. The back-end optimization in the SLAM system is an iterative loop process. In each iteration, a new key frame, that is, the current key frame, is added, and the key frame with the oldest time is removed. The optimization result of the ground parameters obtained from the plane optimization in the second stage of the previous optimization process can be used as the initial value during the local map optimization in the first stage of the current optimization process.
[0129] Figure 4 Show a structural block diagram of an unmanned mobile device according to an embodiment of the present disclosure. AsFigure 4 As shown in the figure, the unmanned mobile device includes an image sensor 401 and a processing device 402. Among them,
[0130] the image sensor 401 collects images of the surrounding environment of the unmanned mobile device and outputs the images to the processing device 402;
[0131] the processing device 402 obtains a plurality of consecutive key frames in the images and a plurality of target three-dimensional points that can be observed in the plurality of consecutive key frames;
[0132] the processing device 402 also performs local map optimization in the first stage using the consecutive key frames and the target three-dimensional points, and performs plane optimization in the second stage using the ground three-dimensional points among the consecutive key frames and the target three-dimensional points. Among them, in the process of local map optimization in the first stage, after fixing the ground parameters in the distance error to the initial values, the pose parameters of the consecutive key frames are optimized; in the process of plane optimization in the second stage, after fixing the pose parameters of the consecutive key frames to the optimization results obtained in the first stage optimization, the ground parameters in the consecutive key frames are optimized.
[0133] In this embodiment, the image sensor 401 may be a monocular camera or other sensors capable of collecting images of the surrounding environment provided on the unmanned mobile device. The processing device 402 may be a processing unit such as a CPU, GPU, FPGA, NPU, etc. The processing device 402 may execute various processes in the above-mentioned simultaneous localization and mapping method of the present disclosure according to the program stored in the read-only memory inside the unmanned mobile device or the program loaded from an external storage medium into the random access memory.
[0134] In this embodiment, after the unmanned mobile device enters an unknown environment, the image sensor 401 collects images in the unknown environment in real time and transmits them to the processing device 402 through the communication link between the image sensor 401 and the processing device 402. After receiving the images, the processing device 402 performs front-end optimization and back-end optimization on them using the SLAM system, and the results of the back-end optimization are transmitted to the server for the construction of the global map.
[0135] During the back-end optimization process, the processing device 402 adopts the local map optimization in the first stage and the plane optimization in the second stage as described above, so as to obtain the pose parameters and ground parameters of each key frame in the current iterative optimization.
[0136] For the specific details of the simultaneous localization and mapping method executed by the processing device in the embodiments of the present disclosure, reference may be made to the description above, and details are not repeated here.
[0137] A method for providing location-based services according to an embodiment of the present disclosure. The method for providing location-based services includes: performing backend optimization using the above-mentioned simultaneous localization and mapping method, constructing a surrounding environment map based on the result of the backend optimization, and providing location-based services for the object to be served based on the surrounding environment map. The location-based services include one or more of navigation, route planning, and map rendering.
[0138] In this embodiment, the method for providing location-based services can be executed on a server. The object to be served can be an unmanned mobile device, such as an unmanned system like a drone, a driverless vehicle, a robot, etc. An image sensor can be provided on the object to be served. The image sensor collects images of the surrounding environment of the unmanned mobile device. The processing device on the unmanned mobile device can construct a surrounding environment map based on the images. When constructing the surrounding environment map, the processing device can perform backend optimization on the pose parameters and ground parameters corresponding to the key frames in the images. The process of backend optimization can refer to the description of the simultaneous localization and mapping method in the above text and will not be elaborated here.
[0139] Based on the pose parameters, ground parameters, etc. obtained from the backend optimization, a local map can be constructed on the unmanned mobile device, and this local map can be sent to the server for constructing a global map of the surrounding environment. The server can use this global map to locate the unmanned mobile device and provide location services such as navigation, path planning, and map rendering for the unmanned mobile device based on the positioning result.
[0140] Figure 5 A schematic diagram of a scenario showing the simultaneous localization and mapping method according to an embodiment of the present disclosure is shown. As Figure 5 shown, the unmanned mobile device 501 includes an image sensor 5011 and a processing device 5012. After the unmanned mobile device 501 enters an unknown environment, the image sensor 5011 collects images of the surrounding environment, and the processing device 5012 processes the collected images to construct a local map of the surrounding environment. The local map includes pose parameters and ground parameters corresponding to multiple consecutive key frames within a sliding window. This local map can be sent to the server 502, and the server 502 optimizes the global map based on this local map and other local map information previously received from the unmanned mobile device 501.
[0141] The server 502 also provides location services for the unmanned mobile device 501 based on this global map, such as determining the location of the unmanned mobile device 501 in the unknown environment, and then further indicating the movement actions of the unmanned mobile device 501 based on this location.
[0142] The following is an embodiment of the device of the present disclosure, which can be used to execute the embodiment of the method of the present disclosure.
[0143] Figure 6 A structural block diagram of an instant positioning map construction device according to an embodiment of the present disclosure is shown. The device can be implemented as part or all of an electronic device through software, hardware, or a combination of both. As Figure 6 shown, the instant positioning map construction device includes:
[0144] An acquisition module 601, configured to acquire a plurality of consecutive key frames, and a plurality of target three-dimensional points that can be observed in all of the plurality of consecutive key frames;
[0145] A first optimization module 602, configured to perform a first-stage local map optimization using the consecutive key frames and the target three-dimensional points; wherein, during the first-stage local map optimization process, after fixing the ground parameters in the distance error to an initial value, the pose parameters of the consecutive key frames are optimized;
[0146] A second optimization module 603, configured to perform a second-stage plane optimization using the consecutive key frames and the ground three-dimensional points among the target three-dimensional points; during the second-stage plane optimization process, after fixing the pose parameters of the consecutive key frames to the optimization results obtained in the first stage, the ground parameters in the consecutive key frames are optimized.
[0147] In an optional implementation manner of this embodiment, both the first-stage local map optimization and the second-stage local map optimization use a cost function for optimization. The cost function includes the reprojection error of the target three-dimensional points on a plurality of consecutive key frames and the distance error between the ground three-dimensional points in the target three-dimensional points and the ground; the reprojection error is obtained based on the target three-dimensional points and the pose parameters of the plurality of consecutive key frames; the distance error is obtained based on the ground three-dimensional points and the ground parameters.
[0148] In an optional implementation manner of this embodiment, the acquisition module includes:
[0149] A first acquisition sub-module, configured to acquire a current key frame, a first historical key frame, and a plurality of consecutive historical key frames between the current key frame and the first historical key frame;
[0150] Wherein, during the first-stage local map optimization, the pose parameters corresponding to the first historical key frame remain fixed, while the pose parameters of the current key frame and the plurality of historical key frames between the current key frame and the first historical key frame are parameters to be optimized.
[0151] In an optional implementation manner of this embodiment, the ground parameters include a ground normal vector parameter and a distance parameter from the ground to the origin of the world coordinate system.
[0152] In an alternative implementation of this embodiment, the apparatus further includes:
[0153] An extraction module, configured to extract feature points in the current key frame;
[0154] A first determination module, configured to determine a current three-dimensional point corresponding to the feature point;
[0155] A second determination module, configured to determine ground three-dimensional points among the current three-dimensional points that are located on the ground by fitting a plane based on the current three-dimensional points.
[0156] In an alternative implementation of this embodiment, the second determination module includes:
[0157] A fitting sub-module, configured to fit one or more planes based on the current three-dimensional points;
[0158] A third determination sub-module, configured to determine a target plane corresponding to the ground among the one or more planes;
[0159] A fourth determination sub-module, configured to determine the distance between the current three-dimensional points and the target plane;
[0160] A screening sub-module, configured to screen out ground three-dimensional points from the current three-dimensional points based on the distance.
[0161] In an alternative implementation of this embodiment, in the local map optimization in the first stage of the current optimization process, the initial value corresponding to the ground parameter is the optimization result obtained in the plane optimization process in the second stage of the previous optimization process.
[0162] The simultaneous localization and mapping apparatus in this embodiment corresponds to the simultaneous localization and mapping method described above. For specific details, reference can be made to the description of the simultaneous localization and mapping method above, and details will not be elaborated here.
[0163] According to a location-based service providing apparatus according to an embodiment of the present disclosure, the location-based service providing apparatus includes: performing backend optimization using the above-mentioned simultaneous localization and mapping apparatus, constructing a surrounding environment map based on the result of the backend optimization, and providing location-based services for the service object based on the surrounding environment map, where the location-based services include one or more of navigation, route planning, and map rendering.
[0164] In this embodiment, the location-based service providing device can be implemented on a server, and the object to be served can be an unmanned mobile device, such as an unmanned system like a drone, a driverless vehicle, a robot, etc. An image sensor can be provided on the object to be served. The image sensor collects images of the surrounding environment of the unmanned mobile device, and the processing device on the unmanned mobile device can construct a map of the surrounding environment based on the images. When constructing the map of the surrounding environment, the processing device can perform backend optimization on the pose parameters and ground parameters corresponding to the key frames in the images. The backend optimization process can refer to the description of the simultaneous localization and mapping device in the above text and will not be elaborated here.
[0165] Based on the pose parameters, ground parameters, etc. obtained from the backend optimization, a local map can be constructed on the unmanned mobile device, and this local map can be sent to the server for constructing a global map of the surrounding environment. The server can use this global map to locate the unmanned mobile device and provide location services such as navigation, path planning, and map rendering for the unmanned mobile device based on the positioning result.
[0166] Figure 7 It is a schematic structural diagram of an electronic device suitable for implementing the simultaneous localization and mapping method and / or the location-based service providing method according to the embodiments of the present disclosure.
[0167] As Figure 7 shown, the electronic device 700 includes a processing unit 701, which can be implemented as a processing unit such as a CPU, GPU, FPGA, NPU, etc. The processing unit 701 can execute various processes in the embodiments of any of the above methods of the present disclosure according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage section 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.
[0168] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that the computer program read from it can be installed into the storage section 708 as needed.
[0169] In particular, according to an embodiment of the present disclosure, any of the methods described above with reference to the embodiments of the present disclosure may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing any of the methods in the embodiments of the present disclosure. In such an embodiment, the computer program may be downloaded and installed from a network via the communication section 709, and / or installed from the removable medium 711.
[0170] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0171] The units or modules described in the embodiments of the present disclosure may be implemented in software or in hardware. The described units or modules may also be provided in a processor, and the names of these units or modules do not in some cases constitute a limitation on the units or modules themselves.
[0172] As another aspect, the present disclosure also provides a computer-readable storage medium, which may be the computer-readable storage medium included in the device described in the above embodiments; or it may exist separately and be a computer-readable storage medium not assembled into the device. The computer-readable storage medium stores one or more programs, and the one or more programs are used by one or more processors to execute the methods described in the present disclosure.
[0173] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the present disclosure that have similar functions.
Claims
1. A method for simultaneous localization and mapping, wherein, Including: Obtaining a plurality of consecutive key frames, and a plurality of target three-dimensional points that can be observed in all of the plurality of consecutive key frames; Performing local map optimization in a first stage using the consecutive key frames and the target three-dimensional points; wherein, during the local map optimization in the first stage, after fixing the ground parameter in the distance error between the ground three-dimensional points in the target three-dimensional points to an initial value, optimizing the pose parameters of the consecutive key frames; Performing plane optimization in a second stage using the consecutive key frames and the ground three-dimensional points in the target three-dimensional points; during the plane optimization in the second stage, after fixing the pose parameters of the consecutive key frames to the optimization result obtained in the first stage optimization, optimizing the ground parameters in the consecutive key frames.
2. The method according to claim 1, wherein In both the local map optimization in the first stage and the local map optimization in the second stage, a cost function is used for optimization, and the cost function includes the reprojection error of the target three-dimensional points on a plurality of consecutive key frames and the distance error between the ground three-dimensional points in the target three-dimensional points and the ground; the reprojection error is obtained based on the target three-dimensional points and the pose parameters of the plurality of consecutive key frames; The distance error is obtained based on the ground three-dimensional points and the ground parameters.
3. The method according to claim 1 or 2, wherein, Obtaining a plurality of consecutive key frames includes: Obtaining a current key frame, a first historical key frame, and a plurality of consecutive historical key frames between the current key frame and the first historical key frame; Wherein, in the local map optimization in the first stage, the pose parameters corresponding to the first historical key frame are fixed, while the pose parameters of the current key frame and the plurality of historical key frames between the current key frame and the first historical key frame are parameters to be optimized.
4. The method according to claim 1 or 2, wherein The ground parameters include a ground normal vector parameter and a distance parameter from the ground to the origin of the world coordinate system.
5. The method according to claim 1 or 2, wherein The method further includes: Extracting feature points in the current key frame; Determining the current three-dimensional points corresponding to the feature points; Determining the ground three-dimensional points located on the ground among the current three-dimensional points based on the method of fitting a plane based on the current three-dimensional points.
6. The method according to claim 5, wherein Determining the ground three-dimensional points located on the ground among the current three-dimensional points based on the method of fitting a plane based on the current three-dimensional points, including: Fitting one or more planes based on the current three-dimensional points; Determining the target plane corresponding to the ground among the one or more planes; Determining the distance between the current three-dimensional points and the target plane; Screening out the ground three-dimensional points from the current three-dimensional points based on the distance.
7. The method according to claim 1 or 2, wherein In the local map optimization in the first stage of the current optimization process, the initial value corresponding to the ground parameters is the optimization result obtained in the plane optimization process in the second stage of the previous optimization process.
8. An unmanned mobile device, comprising: An image sensor and a processor; wherein, The image sensor collects images of the surrounding environment of the unmanned mobile device and outputs the images to the processor; The processor obtains a plurality of consecutive key frames in the images, and a plurality of target three-dimensional points that can be observed in all of the plurality of consecutive key frames; The processor also performs local map optimization in the first stage using the consecutive key frames and the target 3D points, and performs plane optimization in the second stage using the consecutive key frames and the ground 3D points among the target 3D points; wherein, in the process of local map optimization in the first stage, after fixing the ground parameter in the distance error of the ground 3D points among the target 3D points from the ground to the initial value, the pose parameters of the consecutive key frames are optimized; in the process of plane optimization in the second stage, after fixing the pose parameters of the consecutive key frames to the optimization result obtained in the first stage optimization, the ground parameters in the consecutive key frames are optimized.
9. A location-based service providing method, wherein, including: Performing back-end optimization using the method according to any one of claims 1-7, constructing a surrounding environment map based on the result of the back-end optimization, and providing location-based services for the service object based on the surrounding environment map, the location-based services including one or more of navigation, route planning, and map rendering.
10. A computer program product comprising computer instructions, wherein, When the computer instruction is executed by a processor, it implements the method according to any one of claims 1-7.
Citation Information
Patent Citations
Visual SLAM rear-end optimization method based on layered architecture
CN107300917A
Method and device for real-time mapping and localization
EP3078935A1