Systems and methods for estimating the pose of a location device using reflective landmarks and other features - Patents.com
The integration of pre-placed landmarks with SLAM and injective mapping estimates improves localization accuracy and scalability for indoor navigation systems, addressing inaccuracies and scaling issues in existing methods.
Patent Information
- Application Number
- JP2023535398
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-30
- Filing Date
- 2021-12-13
- Publication Date
- 2025-12-22
- Estimated Expiration
- 2041-12-13
AI Technical Summary
Existing methods for determining the state of a localization device, such as robots in indoor environments, suffer from inaccuracies due to signal attenuation and multipath effects, especially in complex environments, and scaling issues with outside-in optical localization systems, making them inadequate for precise navigation.
A method incorporating simultaneous localization and mapping (SLAM) with pre-placed landmarks having known positions, using a combination of injective mapping estimates and extended Kalman filters to correct and initialize the localization device's state, reducing error accumulation by integrating pre-placed landmarks into the SLAM algorithm.
Enhances the accuracy and scalability of localization by reducing error propagation and providing precise positioning of the localization device, even in large indoor spaces.
Smart Images

Figure 0007789782000030 
Figure 0007789782000031 
Figure 0007789782000001
Abstract
Description
[Technical Field]
[0001] The present invention is directed to k The state of the locating device at x k The present invention relates to a method for determining a parameter, and to a computer program product and an assembly. [Background technology]
[0002] Indoor navigation of robots (e.g., drones) is a significant problem, for example, in the field of automated warehouses. Such robots are localization agents. To facilitate indoor navigation, the robot (e.g., drone) needs to know its current position relative to its environment. In contrast to outdoor environments where GNSS (Global Navigation Satellite System) can be employed, which provides high localization accuracy, GNSS in indoor environments is often unreliable due to signal attenuation and multipath effects. Existing RF localization technologies for indoor and outdoor spaces also suffer from signal attenuation and multipath effects that limit their usefulness in complex environments (e.g., in the presence of significant amounts of metal).
[0003] Optical localization systems for indoor localization are known in the prior art. Such optical localization systems extract information from images captured by a camera. The location of an object whose pose is to be determined can then be calculated using triangulation techniques after relating the coordinates of features in a two-dimensional camera image to the three-dimensional light rays corresponding to those features. The association between image coordinates and three-dimensional light rays is typically captured using a first-principles camera model (such as a pinhole or fisheye camera model) in combination with a calibrated distortion model (which typically captures lens characteristics, mounting tolerances, and other deviations from the first-principles model).
[0004] In optical localization systems for determining the location of a localization device known in the prior art, a camera is either mounted firmly outside the localization device and can observe the movement of the localization device ("outside-in tracking"), or a camera is mounted on the localization device itself and can observe the apparent movement of the environment ("inside-out tracking"). While outside-in tracking localization systems typically determine the location of the localization device relative to the known location of the camera, inside-out tracking systems, similar to SLAM (simultaneous localization and mapping), typically generate a map of the environment through which the localization device moves. The map is represented in a coordinate system that can be related to the external coordinate system if the location of at least a portion of the environment is already known relative to the external coordinate system, or if the initial pose of the camera is known relative to the external coordinate system. In both cases, some error will accumulate as the map expands away from the initial field of view of the camera or from the portion of the environment with the known location. The potential for error propagation is an issue for applications where the location information must reference external information (e.g., to display the location of the locating device on a given map; to relate it to the location of another such locating device; or when the location is used to navigate the locating device to a known location in an external coordinate system).
[0005] Outside-in optical localization systems typically scale very poorly to larger localization systems because the localization device must be seen by several cameras at every point in order to triangulate the 3D position of the localization device, which is not economically feasible, especially for larger volumes where only a few localization devices are tracked.
[0006] For example, the position and orientation of a camera mounted on a drone as an example of a location device can be used to determine the state of the location device, which can be tracked over time. However, existing methods for determining the state of a location device do not provide an adequate level of accuracy, thus making them insufficient for use in many applications.
[0007] The object of the present invention is to provide a method for determining the state x of a localization device known from the state of the art. k The object of the present invention is to alleviate at least some of the disadvantages associated with methods for determining Summary of the Invention [Means for solving the problem]
[0008] According to a first aspect of the present invention, the state x of the location determination device is k A method for determining is provided, comprising the steps as defined in claim 1. Further optional features and embodiments of the method of the invention are set out in the dependent patent claims.
[0009] The present invention is directed to k The state of the locating device at x k Regarding how to determine the state x k is the state random variable X k The method includes the steps of: a) receiving a first image of a scene of interest in an indoor environment, the indoor environment comprising N pre-placed landmarks having known positions in a world coordinate system, where N is a natural number; b) receiving a second image of the scene of interest in the indoor environment; and c) receiving a second image of the scene of interest at time t k State estimation of the location device in
number
number
[0010] Simultaneous localization and mapping (SLAM) involves three fundamental operations: A suitable combination of these three fundamental operations can provide a SLAM solution, as discussed below.
[0011] The first fundamental operation of SLAM is to model the movement of a localization device through an indoor environment. A mathematical model of such movement through an indoor environment may be referred to as a motion model. Alternatively or in addition to such a motion model, an inertial measurement unit may also be used. Because movement of a localization device through an indoor environment is noisy and prone to error, the motion model may account for such uncertainty due to noise and error. Each movement of the localization device through the indoor environment may therefore increase the uncertainty of the location of the localization device in the indoor environment. The motion model may be embodied as a function that takes as input the current state of the localization device, control signals, and perturbations, and provides as output a new state of the localization device. The new state of the localization device may, for example, be an estimate of the new position to which the localization device will move.
[0012] The second fundamental operation of SLAM involves determining the location of SLAM landmarks in an indoor environment from an image of a scene of interest in the indoor environment. For example, a SLAM feature detector embodied as an edge detector, corner detector, or SIFT feature detector can be applied to the image of the indoor environment, and SLAM features can be detected in the image of the indoor environment. Such SLAM features detect the projection of SLAM landmarks into the image of the indoor environment. To determine the location of SLAM landmarks from the image of the indoor environment, an inverse observation model can be required. Such an inverse observation model can be used for initialization purposes, as described later, where initialization refers to the detection of new SLAM landmarks not detected in previous iterations of the SLAM algorithm. The inverse observation model can be embodied as a function that takes as input the current state of the localization device and the (locations of) measured SLAM features and provides as output an estimate of the location of the SLAM landmark corresponding to the measured SLAM features. Because 2D images of SLAM landmarks with 3D positions in an indoor environment typically do not provide sufficient information to determine the 3D position of the SLAM landmark (however, a depth camera embodied as, for example, a time-of-flight camera may provide sufficient information to determine the 3D position of the SLAM landmark), additional knowledge (e.g., in the form of prior values provided as additional input to the inverse observation model) is required or needs to be employed in at least two different images of the SLAM landmark; in the latter case, the 3D position of the SLAM landmark is obtained using triangulation, and the inverse observation model may employ two states of the localization device and two measured SLAM features as inputs. To enable triangulation, the SLAM features may need to be tracked across at least two different images. Methods for initializing newly discovered SLAM landmarks are well known in the art.
[0013] The third basic operation of SLAM involves addressing determined SLAM features that correspond to SLAM landmarks already detected in a previous iteration of the SLAM algorithm. The third basic operation is primarily concerned with updating the state of the localization device based on newly acquired images of the scene of interest and updating the positions of those SLAM landmarks detected in a previous iteration of the SLAM algorithm. To perform the third basic operation of SLAM, a direct observation model (also referred to as an observation model) may be used, which predicts the positions of SLAM features in an image based on the current state of the localization device and current estimates of the positions of previously detected SLAM landmarks. The observation model may be embodied as a function that takes as input the current state of the localization device and current estimates of the positions of SLAM landmarks and provides as output the positions of the SLAM features. The direct observation model and the indirect observation model may be inverses of each other.
[0014] Three basic operations, along with an estimation algorithm such as an extended Kalman filter (EKF), can be used to build a SLAM algorithm. A SLAM algorithm based on an EKF can operate on a map state comprising the state of the localization device and the locations of previously detected SLAM landmarks. These previously detected SLAM landmarks may be referred to as currently mapped SLAM landmarks, and an iteration of the SLAM algorithm may be given n such currently mapped SLAM landmarks determined in a previous iteration. The map state used by the SLAM algorithm can therefore grow over time as new SLAM landmarks are detected. The map state can also shrink, and the SLAM algorithm can be implemented in such a way that a SLAM landmark can be removed from the map state if it has not been reobserved within a preset (or dynamically adapted) amount of time or if only limited memory is available to store the SLAM landmark. In the first iteration of the SLAM algorithm, the map state may comprise only the assignment of initial values to the state of the localization agent, and the locations of previously detected SLAM landmarks may not be part of the map state.
[0015] In each iteration of the EKF-SLAM algorithm, the following operations may therefore be performed: a prediction step, a correction step, and an initialization step. In the prediction step, the mean and covariance matrix of the map state may be modified based on a motion model. In the correction step, SLAM features corresponding to previously determined SLAM landmarks (each detected SLAM feature may have a unique signature, and the signature of the detected SLAM feature may therefore be compared with the signatures of previously detected SLAM features to determine whether the detected SLAM feature corresponds to a SLAM landmark that has already been observed) are used in update equations provided by the EKF to update the state of the localization agent and to update the positions of the n currently mapped SLAM landmarks (in the considered iteration of EKF-SLAM, the n SLAM landmarks are assumed to have been previously observed). The update equations in EKF-SLAM, also referred to as correction equations, are based on a (direct) observation model. The update equations as part of the correction step may be applied to the output of the prediction step. In the initialization step, SLAM features corresponding to previously unobserved SLAM landmarks may be used to grow a map state. To determine the locations of the SLAM landmarks from the detected SLAM features, an indirect observation model may be used, which also takes as input the current state (e.g., as may be obtained after a prediction step) of the localization agent (which may be assumed to have a fixed geometric relationship with respect to the camera acquiring the images in which the SLAM features for growing the map state are detected).
[0016] In the method according to the invention, the map state s kalso comprises the positions of N pre-placed landmarks in the indoor environment. The positions of the N pre-placed landmarks in the indoor environment are known. The map state therefore comprises both state information needed for SLAM, such as the state of the localization agent and the positions of currently mapped SLAM landmarks, and the positions of pre-placed landmarks, which do not need to be estimated and tracked over time. A first camera imaging a scene of interest in the indoor environment may capture a first image in which at least some of the N pre-placed landmarks (some of which are not empty) are visible, and a second camera imaging the scene of interest in the indoor environment may capture a second image. The first image and the second image may comprise update information for determining / updating the state of the localization device. The first camera capturing the first image and the second camera capturing the second image may be embodied as a depth camera, an event camera, or a conventional camera. Applying a SLAM feature detector, such as a scale-invariant feature transform (SIFT) detector, or a speeded-up robust feature (SURF) detector, or any other known feature detector suitable for SLAM and known in the art, to the second image can provide SLAM features in the second image. Alternatively, both the first image and the second image can be captured by the same camera, such as the first camera. The first image and the second image can also be coincident, i.e., one image can be both the first image and the second image.
[0017] Through the determined injective mapping estimation, the M determined features in the first image can be mapped into a set of N pre-positioned landmarks. Analog SLAM injective mapping estimation maps m SLAM features into a set of n currently mapped SLAM landmarks (these m SLAM features correspond to previously observed SLAM landmarks). The joint observation model describes the mapping of a map state into / onto a set of observations, where the observations comprise the positions of the M determined features in the first image and the m determined SLAM features in the second image. The joint observation model may comprise a first observation model and a SLAM observation model, which implement the respective mappings. If the first image and the second image are captured by different cameras, the first observation model and the SLAM observation model may respectively model the mapping of each scene of interest onto different image planes of different cameras. The joint observation model is used in the method according to the present invention to update / correct the state of the localization agent and, for example, after the prediction step, to update / correct the positions of the n currently mapped SLAM landmarks. Locating device status x k and the positions of the n currently mapped SLAM landmarks are updated (e.g., using an update formula (correction formula) provided by the EKF) based on the determined positions of at least one of the M features and at least one of the determined m SLAM features.
[0018] The Lm SLAM features may correspond to SLAM landmarks that were not determined in previous iterations. The locations of such newly discovered SLAM landmarks, obtained using the inverse observation model, may be used to grow the map state.
[0019] In the method according to the invention, (images of) the N pre-placed landmarks in the indoor environment are therefore also used to update the positions of the n currently mapped SLAM landmarks in the scene of interest. Since the positions of the N pre-placed landmarks may be known with great accuracy, such a flow of information from the N pre-placed landmarks to the n currently mapped SLAM landmarks may help to reduce the accumulation of errors in both the state of the localization device and the (estimated) positions of the n currently mapped SLAM landmarks.
[0020] The method according to the present invention can therefore be considered as part of a SLAM algorithm operating in an environment with pre-placed landmarks with known positions. As part of the SLAM algorithm, the pre-placed landmarks are used to correct possible errors accumulated by the SLAM algorithm during its execution. In addition to fusing measurements of pre-placed landmarks with SLAM measurements during the execution of the iterative SLAM algorithm, measurements of pre-placed landmarks can also be used to initialize the SLAM algorithm before it begins operation; in this case, a first map state can be initialized with a first state of the localization device (the first state is determined using images of the pre-placed landmarks) and with the positions of the pre-placed landmarks. If the first state is not determined using images of the pre-placed landmarks, the first state can be initialized with an arbitrary value, for example, set to zero. Initializing the first state with information obtained from the pre-placed landmarks can provide a subsequently operating SLAM algorithm with absolute position information (with respect to a world coordinate system in which the positions of the pre-placed landmarks are known) so that the SLAM algorithm can build a map of the environment indirectly referenced to the world coordinate system. Drift of the SLAM algorithm is then corrected using the method according to the present invention. If the pre-placed landmarks are not visible, for example, if the localization device is too close to the ceiling in the scene of interest where the pre-placed landmarks are located, the method according to the present invention can rely on conventional SLAM known from the prior art.
[0021] Alternatively, the estimation of the localization device state may proceed independently, i.e., two estimates of the localization device state may be determined in parallel, where a first estimate of the localization device state may be obtained using only images of pre-placed landmarks, and a second estimate of the localization device state (and the growing set of SLAM landmarks in the indoor environment) may be obtained through a conventional SLAM algorithm. If the two estimates of the localization device state are not within a certain threshold of each other, the currently constructed SLAM landmark map comprising the currently mapped SLAM landmarks may be discarded and reinitialized. The first estimate of the localization device state may be used to provide loop termination for a conventional SLAM algorithm, i.e., when the localization device reaches a state it was in previously in a previous iteration, this information may be used to inform the SLAM algorithm that it should expect to see a simultaneous SLAM landmark that it already saw in the previous iteration. Such informing the SLAM algorithm may be achieved, for example, by lowering the threshold for signature comparison of two features so that they are considered to match (i.e., are the same feature).
[0022] In an embodiment of the method according to the invention, the first image and the second image are taken at a time t k is captured at
[0023] The first and second images are taken at the same time t kThis may be beneficial because, in this case, processing of the first and second images may occur in parallel. If the two images are captured at different times, updating the map state may proceed in two steps: in a first update step, for example, based on the first captured image, intermediate updates to the map state may be determined based on information in the first captured image; and this intermediate map step may be updated based on a motion model and information determined from a subsequently captured image (the motion model may be used to propagate the intermediate updates to the map state in time until the capture of the subsequently captured image). If the first and second images are not captured at the same time, one or both images may be algorithmically shifted in time so that the two images are aligned in time; such algorithmic time shifting may be based, for example, on interpolation or extrapolation techniques applied to a sequence of first images (captured over time by a first camera) and / or a sequence of second images (captured over time by a second camera).
[0024] In a further embodiment of the method according to the present invention, the joint observation model comprises a first observation model, which describes a mapping of at least one of the N pre-placed landmarks corresponding to at least one of the M features via injective mapping estimation onto a position of at least one of the M features, and the joint observation model comprises a SLAM observation model, which describes a mapping of at least one of the n currently mapped SLAM landmarks corresponding to at least one of the m SLAM features using SLAM injective mapping estimation onto a position of at least one of the m SLAM features.
[0025] In a further embodiment of the method according to the invention, an injective mapping estimation (hereinafter referred to as IME) links M features with M pre-assigned landmarks, a SLAM injective mapping estimation (hereinafter referred to as SLAM-IME) links m SLAM features with m SLAM landmarks, where for feature i among the M features, the corresponding landmark is landmark IME(i), and for feature j among the m SLAM features, the corresponding SLAM landmark is SLAM landmark SLAM-IME(j), and for feature-landmark pair (i, IME(i)) and for SLAM feature / SLAM landmark pair (j, SLAM-IME(j)), the joint observation model is the observation random variable Z k,i , Z k,j (the linked observation random variable Z k are the observed random variables Z k,i , Z k,j (equipped with) the state random variable X k and link to landmark POS(IME(i)), SLAM landmark POS(SLAM-IME(j),k)): Z k,i =h1(X k ,POS(IME(i)),i=1,···,M,and, Z k,j =h SLAM (X k ,POS(SLAM-IME(i),k)),j=1,...,m where h1(·) is the first observation model and h SLAM (·) is the SLAM observation model.
[0026] The position of the SLAM landmark POS(SLAM-IME(j),k) generally changes over time. The position of the SLAM landmark POS(SLAM-IME(j),k) therefore similarly depends on the index k.
[0027] In a further embodiment of the method according to the present invention, a priori known positions in a world coordinate system of S SLAM landmarks are provided, the S SLAM landmarks are associated with a pre-configured SLAM feature detector, the application of which to a second image provides L SLAM features, and associations are provided between the S SLAM landmarks and their respective feature signatures, and between the feature signatures themselves, the feature signatures being related to outputs provided by application of the pre-configured SLAM feature detector to the image of the S SLAM landmarks, and the method additionally includes: 1) determining t of the L SLAM features associated with the S SLAM landmarks with the a priori known positions by comparing the feature signatures of the L SLAM features with the feature signatures in the associations; and 2) as part of determining a SLAM injective mapping estimate, establishing an injective mapping from the t SLAM features to the corresponding SLAM landmarks based on a comparison of the feature signatures of the L SLAM features to the feature signatures in the associations. In principle, t can be equal to any natural number between "0" (SLAM features determined using corresponding SLAM landmarks do not have a priori known positions in the world coordinate system) and L, the latter case corresponding to a situation where all L determined SLAM features have corresponding SLAM landmarks with a priori known positions in the world coordinate system.
[0028] In a further embodiment of the method according to the invention, the n currently mapped SLAM landmarks comprise only SLAM landmarks with a priori unknown positions in the world coordinate system. The determined SLAM features with corresponding SLAM landmarks with a priori known positions in the world coordinate system are used to determine the state x of the localization device. k and to determine the positions of SLAM landmarks with a priori unknown positions in the world coordinate system.
[0029] In a further embodiment of the method according to the invention, M features associated with the N pre-placed landmarks are substantially indistinguishable from one another, the indoor environment comprises K additional landmarks with a priori unknown positions in a world coordinate system, H (≦K) additional landmarks of the K additional landmarks are within a scene of interest, and currently estimated positions of P of the K additional landmarks are determined by a map state s k and H additional landmarks are captured in the first image as additional features substantially indistinguishable from the M features, and the method additionally includes: 1) determining the locations of the H additional features together with determining the locations of the M features, thereby providing M+H locations of the features and the additional features in the first image; 2) separating the locations of the M+H features and the additional features into a first set having M+R locations (where R≦P) and a second set having Q locations (where H=R+Q); 3) determining an injective mapping estimate from the first set to a set of indistinguishable landmarks comprising the locations of the N pre-placed landmarks and the P currently estimated locations; and 4) based on the second set, generating a map state s k This embodiment of the method according to the invention may also be combined with an embodiment of the method according to the invention comprising S SLAM landmarks with a priori known positions in the world coordinate system.
[0030] The N pre-placed landmarks may be indistinguishable from one another. For example, the N pre-placed landmarks may be embodied and constructed as the same type of retroreflector. The projections of the pre-placed landmarks into the first image, of which the projections are characteristic, may therefore be substantially indistinguishable from one another; i.e., the pre-placed landmark to which an inspected feature corresponds may not be unambiguously determined by inspecting only one feature at a time in the first image and ignoring other features in the first image. In addition to the N pre-placed landmarks, K additional landmarks may similarly be present in the indoor environment, and the K additional landmarks may be indistinguishable from the N pre-placed landmarks. Alternatively, only the additional features corresponding to the additional landmarks may be indistinguishable from the features corresponding to the pre-placed landmarks. The number assigned to the variable K may be known or unknown. Thus, the number of additional landmarks present in the indoor environment may not be known. First, the positions of the K additional landmarks may not be known in the world coordinate system, contrary to the a priori known positions of the N pre-placed landmarks in the world coordinate system. In the scene of interest of the indoor environment captured by the first image, H (≦K) additional landmarks may be captured as the H additional features, and without loss of generality, it may be assumed that the H additional features may be discernible from the first image in a robust and stable manner. In total, therefore, there may be M+H features and additional features in the first image, which may correspond to (some of) the N+K pre-placed landmarks and the additional landmarks.
[0031] Map states kmay additionally comprise currently estimated positions of P of the K additional landmarks. The currently estimated positions of the P additional landmarks may have been determined in a previous iteration of the algorithm into which the method according to the present invention may be embedded as previously described. Thus, although the a priori positions of the additional landmarks may be unknown, these positions may be estimated over time, whereby in a particular iteration of the overall algorithm, e.g., iteration k, the positions of the P additional landmarks may have been previously estimated, implying that, given K is known, from the N+K landmarks and additional landmarks, the positions of the N+P landmarks and additional landmarks in the world coordinate system may be known, and the positions of the KP additional landmarks may still be unknown.
[0032] The observed M+H positions of the feature and additional features may be separated into a first set comprising those features and additional features corresponding to landmarks and additional landmarks with known (or currently estimated) positions in the world coordinate system, and a second set comprising those additional features corresponding to additional landmarks with currently unknown positions in the world coordinate system. An injective mapping estimate may then be determined from the first set into a set of indistinguishable landmarks comprising the (a priori known) positions of the N pre-placed landmarks and the P currently estimated positions of the P additional landmarks.
[0033] The second set may comprise additional features corresponding to additional landmarks whose positions have not been previously estimated. The second set is therefore used for initialization, i.e., to store the estimated positions of the additional features corresponding to the second set in the map state s k Based on the second set, the map state s k can thus be extended: the information required to estimate the position of additional features in the world coordinate system can be obtained by using additional sensors, for example, depth cameras, or by suitably linking additional features across multiple images acquired by the first camera and triangulating based on such established relationships.
[0034] The features corresponding to pre-placed landmarks and the additional features corresponding to previously initialized additional landmarks are of a priori unknown locations, but with a posteriori estimated locations, and can therefore be processed in the same way, i.e., the same algorithm for estimation of the injective mapping estimate can be applied to the elements of the first set, and the additional features corresponding to the previously uninitialized additional landmarks can be initialized, i.e., the map state s k But it can be grown.
[0035] The first set may comprise M+R elements, where M of these elements may be associated with features corresponding to pre-placed landmarks, and R≦P of these elements may be associated with features corresponding to additional landmarks that have been previously initialized. The second set may have Q elements, where H=R+Q.
[0036] In a further embodiment of the method according to the invention, the first image is captured by a first camera with a first camera center and a first image sensor, and the position of the first camera center in the world coordinate system is determined by the state estimation.
number
number
[0037] Separating the M+H features and additional features into the first and second sets may proceed by geometric means, given a current estimate of the pose of the first camera capturing the first image in a world coordinate system (such a current estimate of the pose of the first camera may be used, for example, to estimate the state of the localization device).
number
number
[0038] In a further embodiment of the method according to the invention, the observation additionally comprises the location of at least one of the R additional features in the first image, and together with updating the locations of the n currently mapped SLAM landmarks, the currently estimated locations of the P additional landmarks are updated. However, the currently estimated locations of the P additional landmarks may also be updated based on observed features corresponding to pre-located landmarks with a priori known locations and / or based on observed SLAM features.
[0039] Since the currently estimated positions of the P additional landmarks are only estimates, these positions may also be updated along with the positions of the n currently mapped SLAM landmarks.
[0040] In a further embodiment of the method according to the invention, the first observation model additionally describes a mapping of at least one of the R additional landmarks corresponding to at least one of the R additional features using an injective mapping estimate onto the position of at least one of the R additional features.
[0041] In a further embodiment of the method according to the invention, (i) state estimation
number
number
number
[0042] In a further embodiment of the method according to the invention, the update equations corresponding to all M features and all m SLAM features are called consecutively and independently.
[0043] In a further embodiment of the method according to the invention, the first image is captured by a first camera when a light source is operated to emit light that illuminates the scene of interest. The second image can also be captured by a second camera when the light source is operated to emit light that illuminates the scene of interest. The light source used during capture of the first image and the light source used during capture of the second image can be the same light source.
[0044] According to a further aspect of the invention there is provided a computer program product comprising instructions which, when executed by a computer, cause the computer to perform a method according to the invention.
[0045] According to a further aspect of the present invention, an assembly is provided, the assembly comprising: (a) a localization device; (b) a first camera mounted on the localization device; (c) a plurality of pre-placed landmarks in an indoor environment; and (d) a controller, the controller configured to perform a method according to the present invention.
[0046] The first camera may be configured to provide a first image. The first camera may also provide a second image, and the first and second images may be coincident. The first camera may be embodied as an RGB-IR (red-green-blue-infrared) camera, where the RGB channel may provide the first or second image and the IR channel may provide the other image. The first camera may also be embodied as providing a high dynamic range (HDR) image. The method according to the present invention may determine the state of the localization device through a known geometric relationship between the local coordinate system of the localization agent and the first camera coordinate system of the first camera, and a known coordinate transformation between the local coordinate system and the first camera coordinate system may be used to map the state of the first camera onto the state of the localization device.
[0047] In an embodiment of the assembly according to the present invention, a second camera is disposed on the positioning device, the first camera is disposed on the positioning device in such a way that the ceiling of the indoor environment is within the field of view of the first camera, the camera axis of the first camera is within the field of view of the first camera, and the second camera is disposed on the positioning device in such a way that the camera axis of the second camera is substantially orthogonal to the camera axis of the first camera. For example, the camera axis of the first camera may be defined by the center of projection of the first camera and the center of the image sensor of the first camera, and the camera axis of the second camera may be defined by the center of projection of the second camera and the center of the image sensor of the second camera. Preferably, the first camera is thus disposed on the positioning device such that the ceiling of the indoor environment is within the field of view of the first camera, and the second camera is disposed on the positioning device in such a way that the camera axis of the second camera is substantially orthogonal to the camera axis of the first camera. More preferably, the first camera is positioned on the location device in such a way that the ceiling is within the field of view of the first camera when the first image is captured by the first camera, and the second camera is positioned on the location device in such a way that the camera axis of the second camera is substantially orthogonal to the camera axis of the first camera. The first camera and the second camera may thus be positioned on the location device in such a way that their respective fields of view are substantially different from each other.
[0048] A drone or a mobile land-based robot is an exemplary embodiment of a localization device. A normal movement condition for such a localization device is horizontal movement relative to the ground surface of the scene of interest, either in the air (such as in the case of a drone) or on the ground surface itself (such as in the case of a land-based robot). The N pre-placed landmarks may be located on the ceiling in the scene of interest. In such a normal movement condition, the cameras may be positioned on the localization device in such a way that the camera axis of the first camera is pointing toward the ceiling of the indoor environment. The first camera may be positioned on the localization device in such a way that, during capture of the first image by the first camera, a dot (inner) product between the camera axis of the first camera and the vector of gravity provides a negative result, and the camera axis of the first camera may be within a 45-degree cone around the vector of gravity at the time of capture of the first image by the first camera (the axis of the cone may be aligned with the vector of gravity and may be pointed in the opposite direction compared to the vector of gravity (the axis of the cone points from the apex of the cone to its base), i.e., the axis of the cone may point away from the ground in an indoor environment). On the other hand, SLAM landmarks may typically be found within the scene of interest itself, for example, because typical SLAM features may be determined based on strong local contrast. On the other hand, the ceiling of a room (e.g., the ceiling of a room is the scene of interest) is typically fairly homogeneous, and therefore only a few SLAM features will typically be detected on the ceiling of an indoor room. The first camera may therefore be positioned on the localization agent in such a way that during normal movement conditions, the first camera faces upward toward the ceiling, i.e., it captures images of the ceiling most of the time. The second camera may be positioned on the localization agent in such a way that it primarily captures images of the interior of the indoor environment. The camera axis of the second camera may therefore be oriented in a manner that is substantially orthogonal to the camera axis of the first camera.
[0049] In a further embodiment of the assembly according to the invention, the assembly comprises a light source, which is arranged on the localization device and / or on the at least one additional landmark. The at least one additional landmark may also be arranged in the indoor environment with an a priori unknown position, for example, during the arrangement of the at least one additional landmark in the indoor environment, the position of the at least one additional landmark may not be recorded. However, the additional landmark may also naturally be present in the indoor environment. [Brief explanation of the drawings]
[0050] Exemplary embodiments of the invention are disclosed in this description and illustrated by the following drawings.
[0051] [Figure 1] FIG. 1 shows a schematic representation of an embodiment of a method according to the invention for determining a state xk of a locating device at a time tk. [Figure 2] FIG. 2 shows a schematic depiction of a drone equipped with a light source and a camera, the drone configured to fly in an indoor environment, and markers positioned at multiple locations in the indoor environment. DETAILED DESCRIPTION OF THE INVENTION
[0052] Figure 1 shows the time t k The state of the locating device at x k 1 shows a schematic representation of an embodiment of a method according to the invention for determining a state x k is the time t k The 3D position and 3D orientation may be expressed relative to a world coordinate system. k In addition, time t k The location information may comprise 3D velocity information of the location device at a given location, which may also be expressed relative to a world coordinate system, for example. As the location device may move through an indoor environment over time, its state may need to be tracked to determine the current position and orientation of the location device.
[0053] Time t k In the indoor environment, a first image 1 and a second image 2 of a scene of interest are received, and a state estimation 3 of a localization device in the indoor environment is performed.
number
[0054] The first image 1 is an image of a scene of interest in an indoor environment equipped with N pre-placed landmarks. The positions (and possibly orientations) of the N pre-placed landmarks in the indoor environment are known in a given world coordinate system. At time t k In (1), the first camera capturing the first image 1 has a particular position and orientation, so that all of the N pre-placed landmarks may not be in the field of view (i.e., not visible) to the first camera. For example, J(≦N) pre-placed landmarks may be captured at time t kIn this case, it may be visible to the first camera, and these J pre - arranged landmarks are projected by the first camera onto the first image 1 of the target scene in the indoor environment. The projection of the pre - arranged landmarks into the image can be referred to as "features". From the J pre - arranged landmarks projected onto the first image 1, M (≤ J) features can be identified, and the 2D positions of these features in the image are determined (5). The 2D position of a feature can be, for example, the 2D position of the centroid of the feature. Some of the J pre - arranged landmarks are such that their projection into the image is too small / unclear / poorly detectable at time t k In this case, it can be positioned and oriented relative to the first camera. In this case, M can be strictly less than J, i.e., M < J, and the remaining J - M pre - arranged landmarks projected onto the first image 1 by the first camera can be ignored. It can also be assumed that the M features correspond to the pre - arranged landmarks, i.e., outliers have been removed. Generally, it is possible that more than N features are determined, for example, due to outliers, and among these determined features, there are M features that actually correspond to the pre - arranged landmarks. These M features can be determined prior to the determination of the injective mapping estimation, or these M features can be determined during the determination of the injective mapping estimation. Among the determined features, there can be more than M features that actually correspond to the pre - arranged landmarks, and the M features can thus correspond to an appropriate part of those features among the determined features that actually correspond to the pre - arranged landmarks.
[0055] The SLAM features in the second image 2 are determined, and more specifically, the positions of the SLAM features in the second image 2 are determined. To determine (6) the positions of the SLAM features in the second image 2, any SLAM feature detector known from the prior art can be used. Applying the SLAM feature detector to the second image 2 provides a list of L determined SLAM features with the determined positions of the SLAM features. Each determined SLAM feature corresponds to a SLAM landmark in the scene of interest in the indoor environment depicted in the second image. Since different SLAM feature detectors typically detect different SLAM features in the second image 2, the SLAM landmarks depend on the selection of the SLAM feature detector. Generally, only m of the L determined SLAM features correspond to some of the n SLAM landmarks observed in previous iterations.
[0056] An injective mapping estimation from M features to N pre - placed landmarks is also determined (5). Typically, since it maintains M < N, the injective mapping estimation is typically not injective enough and not surjective. The injective mapping estimation explains which of the N pre - placed landmarks induced which of the M features in the first image 1. To determine (5) such an injective mapping estimation, the current state of the first camera at time t k may need to be known, and the current state of the first camera can be derived from the current state of the positioning device to which the first camera can be attached. The injective mapping estimation can also be determined starting from all determined features (i.e., including outliers). During the determination of the injective mapping estimation, outliers are identified, whereby only the M features corresponding to the pre - placed landmarks are mapped into the set of N pre - placed landmarks. The SLAM injective mapping estimation from m SLAM features to n currently mapped SLAM landmarks can be determined (6), for example, using a feature signature that identifies individual SLAM features.
[0057] Using the determined injective mapping estimate and the determined SLAM injective mapping estimate, in the next step a joint observation model is established (7). The joint observation model can be used by the extended Kalman filter to perform the task of updating / correcting the output of the prediction step of the extended Kalman filter (8). The joint observation model is used to map the state random variables S k is constructed to map onto the concatenated observation random variables, and maps the state random variables S k comprises (i) the current state of the localization device, (ii) the positions of n currently mapped SLAM landmarks, and (iii) the known positions of N pre-placed landmarks, and the concatenated observation random variable statistically describes the positions of features in a first image 1 and the positions of SLAM features in a second image 2. The observations corresponding to the concatenated observation random variable are obtained through an actual measurement process or through calculations performed on data provided by the actual measurement process.
[0058] The coupled observation model is part of a state-space model used to track the movement of the localization device through a scene of interest and to build a SLAM landmark map of the scene. In addition to the coupled observation model, the state-space model typically includes a state-transition model (motion model) used in the prediction step. As an alternative to or in addition to the motion model, an inertial measurement unit may be used. In this case, only the localization device may be assumed to move through the scene of interest, while both the SLAM landmarks and pre-placed landmarks may be assumed to be static over time (although the estimated positions of the SLAM landmarks are typically not static). The state-transition model describes how the state itself evolves over time. If the localization device is embodied as an aircraft, for example, the state-transition model may include equations that model the aircraft's flight, possibly including control inputs and perturbation inputs used to control the aircraft's flight. The coupled observation model and / or the state-transition model may be linear or nonlinear in their inputs.
[0059] If both the coupled observation model and the state-transition model are linear, the Kalman filter can be used to estimate at least the state
number
number
number
number
[0060] The update equations for the Kalman filter or extended Kalman filter may be invoked for all M features and m SLAM features simultaneously, or for each feature of the M features and each SLAM feature of the m SLAM features separately, or for all M features simultaneously and m SLAM features separately. In some embodiments of the method according to the invention, the update equations may be invoked only for those features of the M features whose Mahalanobis distance to predicted feature positions is less than a certain threshold and the predicted feature positions are determined, e.g., by the state estimate
number
[0061] FIG. 2 shows a schematic depiction of an aircraft in the form of a drone equipped with a light source 10 and a first camera 11, the drone flying in an indoor environment 15, and a scene of interest residing within the indoor environment. Each of the pre-placed landmarks 16 is preferably embodied as a retroreflector and is positioned at multiple locations within the indoor environment 15. The pre-placed landmarks 16 may be mounted on a ceiling in the scene of interest 15. The pre-placed landmarks may be embodied in a substantially identical manner, i.e., the pre-placed landmarks may be completely interchangeable with one another. At any given pose of the drone (with its position and orientation), some landmarks 16 may be visible to the first camera 11, as indicated by the lines between the pre-placed landmarks 16 and the first camera 11 in FIG. 2, while other pre-placed landmarks 16 may not be visible to the first camera 11. In FIG. 2, the first camera 11 is mounted below the drone. The first camera 11 may alternatively be mounted on top of the drone in such a way that the camera is preferably positioned facing the ceiling so that landmarks 16 located on the ceiling are visible to the first camera, and the first camera may be located adjacent to the light source 10. Preferably, the drone further comprises a second camera (not shown in FIG. 2). Preferably, the second camera is oriented orthogonally to the first camera. Images captured by the second camera may be used as input to a SLAM algorithm.
[0062] The positions of pre-placed landmarks 16 may be known in world coordinate system 12, the drone's current location may be expressed as drone coordinate system 13, and a coordinate transformation 14 may be known between world coordinate system 12 and drone coordinate system 13. If first camera 11 and light source 10 are fixed to the drone and their pose relative to the drone is known, the pose of first camera 11 and light source 10 may be related to world coordinate system 12 using drone coordinate system 13. The drone's current position may be determined using a first image of a scene of interest 15 in an indoor environment 15, specifically pre-placed landmarks 16 having known positions, and using a second image of the scene of interest 15 in the indoor environment 15, which may be used by a SLAM algorithm. Alternatively, or in addition, the drone may be equipped with an inertial measurement unit, which may also be used to determine the drone's pose. The light source 10 may be an isotropically emitting light source, or it may be a directional light source that emits in an anisotropic manner. The light source 10 and the first camera 11 are ideally close to each other, especially if the pre-placed landmarks 16 are embodied as retroreflectors. (Item 1) Time t k The state of the locating device at x k (9) A method for determining the state x k is the state random variable X k and the method is a) receiving a first image (1) of a scene of interest (15) in an indoor environment (15), the indoor environment (15) comprising N pre-placed landmarks (16) having known positions in a world coordinate system (12), where N is a natural number; b) receiving a second image (2) of a scene of interest (15) in said indoor environment (15); c) the time t k State estimation of the position determination device in
number
number
number
number
number
number
number
Claims
1. Time t k The state of the locating device at x k (9), wherein the state x k comprises the 3D position and 3D orientation of the localization device relative to the world coordinate system (12), and the state random variable X k is a mathematical representation of the state x k , and the method comprises: a) receiving a first image (1) of a scene of interest (15) in an indoor environment (15), the indoor environment (15) comprising N pre-placed landmarks (16) having known positions in the world coordinate system (12), where N is a natural number; b) receiving a second image (2) of a scene of interest (15) in said indoor environment (15); c) said time t k State estimation of the position determination device in [Equation 1] (3) receiving the d) receiving the locations of n currently mapped Simultaneous Localization and Mapping (SLAM) landmarks (4) in the scene of interest (15), wherein the map state s k However, at least (i) the state x of the location device k (ii) the locations of the n currently mapped SLAM landmarks (4); and (iii) the locations of the N pre-placed landmarks (16). e) determining (5) the locations of M features in the first image (1), where M is a natural number less than or equal to N, and determining (5) an injective mapping estimate from the M features to the set of N pre-placed landmarks (16); f) determining (6) the locations of L SLAM features in the second image (2); determining m SLAM features among the L SLAM features, the m SLAM features being related to the n currently mapped SLAM landmarks (4); and determining a SLAM injective mapping estimate (6) from the m SLAM features to the set of the n currently mapped SLAM landmarks (4); g) using the determined injective mapping estimate and the determined SLAM injective mapping estimate to establish a joint observation model as part of a state-space model (7), the state-space model being a mathematical model used to track movement of the localization device through the scene of interest and to construct a SLAM landmark map of the scene of interest, the joint observation ... k On top of the map state s k The map state random variable S k a mathematical model describing the mapping of the time t k In the above, the linked observed random variable Z k is the connected observation z k is a mathematical expression of the connected observation z k comprises the location of at least one of the M features in the first image (1) and the location of at least one of the m SLAM features in the second image (2); h) (i) the state estimation [Equation 2] (3), (ii) the linked observation model, and (iii) the linked observation z k (8) By using (8), the time t k The state x of the location device at k (9) and updating the positions of the n currently mapped SLAM landmarks. A method comprising:
2. The first image (1) and the second image (2) are taken at the time t k The method of claim 1 , wherein the soluble solid is captured in a
3. 3. The method of claim 1, wherein the joint observation model comprises a first observation model that describes the mapping of at least one of the N pre-placed landmarks (16) corresponding to at least one of the M features using the injective mapping estimate onto the location of at least one of the M features, and the joint observation model comprises a SLAM observation model that describes the mapping of at least one of the n currently mapped SLAM landmarks (4) corresponding to at least one of the m SLAM features using the SLAM injective mapping estimate onto the location of at least one of the m SLAM features.
4. the injective mapping estimation, hereafter referred to as IME, links the M features with the M pre-assigned landmarks; the SLAM injective mapping estimation, hereafter referred to as SLAM-IME, links the m SLAM features with the m SLAM landmarks, where for feature i of the M features, the corresponding landmark is landmark IME(i), and for feature j of the m SLAM features, the corresponding SLAM landmark is SLAM landmark SLAM-IME(j); For a feature-landmark pair (i, IME(i)), and for a SLAM feature / SLAM landmark pair (j, SLAM-IME(j)), The joint observation model is k and the landmark position POS(IME(i)), X k and SLAM landmark POS (SLAM-IME(j), k) respectively, and the observed random variable Z k,i , Z k,j and linking the linked observation random variable Z k is the observed random variable Z k,i , Z k,j Equipped with Z k,i =h 1 (X k , POS(IME(i))), i=1,...,M, and Z k,j =h SLAM (X) k ,POS(SLAM-IME(j),k)),j=1,・・・,m where h 1 (·) is the first observation model, and h SLAM The method of claim 3 , wherein (·) is the SLAM observation model.
5. The M features associated with the N pre-placed landmarks (16) are substantially indistinguishable from one another, the indoor environment (15) comprises K additional landmarks with a priori unknown positions in the world coordinate system (12), H (≦K) additional landmarks of the K additional landmarks are within the scene of interest (15), and currently estimated positions of P of the K additional landmarks are determined by the map state s k wherein the H additional landmarks are captured as additional features in the first image (1) that are substantially indistinguishable from the M features; The method comprises: 1) determining (5) positions of the H additional features together with the determination (5) of the positions of the M features, the determination (5) thereby providing M+H positions of the features and the additional features in the first image (1); 2) separating the locations of the M+H features and additional features into a first set comprising M+R locations and a second set comprising Q locations (5), where R≦P and H=R+Q; 3) determining (5) the injective mapping estimate from the first set to a set of indistinguishable landmarks comprising the positions of the N pre-placed landmarks and the P currently estimated positions; 4) based on the second set, k Update and 5. The method of claim 3 or 4, further comprising:
6. The first image (1) is captured by a first camera (11) with a first camera center and a first image sensor, and the position of the first camera center in the world coordinate system (12) is determined by the state estimation [Equation 3] (5) is determined based on (3), and the separating (5) is M+H positions, the M+H positions being 2D positions measured in the coordinate system of the first image sensor, and for each element in the set of M+H positions: (i) determining a ray having the first camera center as an initial point and having as a further point a 3D position corresponding to the 2D position of each of the elements, the 3D position being determined by the state estimate [Equation 4] (3) and a known spatial relationship between the localization device and the first image sensor. (ii) if any landmark in the set of indistinguishable landmarks is closer to the respective ray than a preset distance threshold as measured by orthogonal projection onto the respective ray, assigning the respective element to the first set, otherwise assigning the respective element to the second set; 6. The method of claim 5, which proceeds by performing
7. The connected observation z k further comprising the location of at least one of the R additional features in the first image (1), and together with the updating of the locations of the n currently mapped SLAM landmarks, the currently estimated locations of the P additional landmarks are updated.
8. 8. The method of claim 7 , wherein the first observation model further describes the mapping of at least one of the R additional landmarks corresponding to at least one of the R additional features using the injective mapping estimate onto the location of at least one of the R additional features.
9. (i) the state estimation [Equation 5] (3), (ii) the linked observation model, and (iii) the linked observation z k and the state x k The determining (8) of (8) is performed using an update equation provided by applying an extended Kalman filter to the state-space model, the update equation including a Jacobian matrix of the joint observation model, the Jacobian matrix of the first observation model being [Equation 6] (3), and the Jacobian matrix of the SLAM observation model is evaluated by at least the state estimate [Equation 7] The method of any of claims 3-8, wherein the method is evaluated in (3) and at the locations of the n currently mapped SLAM landmarks (4).
10. 10. The method of claim 9, wherein the update equations corresponding to all M features and all m SLAM features are called consecutively and independently.
11. 7. The method of claim 6, wherein the first image (1) is captured by the first camera (11) when a light source (10) is operated to emit light that illuminates the scene of interest (15).
12. A program comprising instructions which, when executed by a computer, cause the computer to perform the method of any one of claims 1-11.
13. 12. A system comprising: (a) a localization device; (b) a first camera (11) disposed on the localization device; (c) a plurality of pre-placed landmarks (16) in an indoor environment (15); and (d) a controller, the controller configured to perform the method of any one of claims 1-11.
14. 14. The system of claim 13, further comprising a second camera arranged on the localization device, wherein the first camera (11) is arranged on the localization device in such a way that a ceiling of the indoor environment is within a field of view of the first camera, a camera axis of the first camera is within a field of view of the first camera, and the second camera is arranged on the localization device in such a way that a camera axis of the second camera is substantially orthogonal to the camera axis of the first camera.
15. 15. The system according to claim 13 or 14, further comprising a light source (10), said light source (10) being arranged on said localization device and / or at least one additional landmark.
Citation Information
Patent Citations
Position recognition method, device thereof, program thereof, recording medium thereof, and robot device provided with position recognition device
JP2003266349A
Self-location detection device and system for traveling object
JP2007257226A
Information processor, information processing method, and program
JP2011043419A
Information processor, method, program, and movable body control system
JP2020056757A