Method and system for determining a state of a camera
By identifying feature points and establishing injective mappings with landmarks in the indoor environment, and using a TOF camera and Kalman filter to update the camera state, the problem of insufficient camera state determination accuracy in existing technologies is solved, achieving high-precision indoor positioning and navigation.
Patent Information
- Application Number
- CN202180088100.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-30
- Filing Date
- 2021-12-13
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2041-12-13
AI Technical Summary
Existing methods cannot effectively determine the camera's status in indoor environments, especially in large spaces, resulting in insufficient positioning accuracy and failing to meet the needs of many applications.
By receiving indoor environment images captured by a camera, feature points are identified and an injective mapping is established with landmarks at known locations. The distance between the features and landmarks is measured using a TOF camera, and the camera state is updated by combining a Kalman filter or an extended Kalman filter, thus achieving an accurate estimate of the camera state.
It improves the positioning accuracy of the camera in indoor environments, effectively tracking the camera's position and orientation, and is suitable for indoor navigation and positioning tasks in large spaces.
Smart Images

Figure CN116724336B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to a method for determining a state x k of a camera at a time t k , as well as a computer program product and an assembly. BACKGROUND
[0002] Indoor navigation of robots, e.g. drones, is an important issue, e.g. in the field of automated warehousing. To facilitate indoor navigation, a robot, e.g. a drone, needs to know its current position relative to its environment. In contrast to outdoor environments, where GNSS (Global Navigation Satellite System) can be employed, thereby providing high positioning accuracy, GNSS is often unreliable in indoor environments due to signal attenuation and multipath effects. Existing RF positioning technologies for indoor and outdoor spaces also struggle with signal attenuation and multipath effects that limit availability in complex environments, e.g. in the presence of a lot of metal.
[0003] In prior art, optical positioning systems for indoor positioning are known. Such optical positioning systems extract information from images captured by a camera. After relating the coordinates of a feature in a two-dimensional camera image and a three-dimensional ray corresponding to said feature, a position of an object whose pose is to be determined can then be computed using triangulation techniques. The relation between image coordinates and three-dimensional rays is typically captured in a combination of a first-principle camera model, such as a pinhole or fish-eye camera model, and a calibrated distortion model, which typically captures lens characteristics, mounting tolerances, and other deviations from the first-principle model.
[0004] In optical positioning systems for determining a position of an object known in prior art, a camera can be rigidly mounted outside the object, thereby observing the motion of the object ("outside-in tracking"), or the camera can be mounted on the object itself, thereby observing the apparent motion of the environment ("inside-out tracking"). Outside-in tracking positioning systems typically determine the position of the object relative to a known position of the camera, while inside-out tracking systems, like SLAM (Simultaneous Localization and Mapping), typically generate a map of the environment in which the object moves. The map is expressed in an unknown coordinate system, but can be related to a known coordinate system if the position of at least part of the environment is already known or if the initial pose of the camera is known. In both cases, some error will accumulate as the map expands away from the initial field of view of the camera or the part of the environment with a known position. For applications where the position information must be referenced to external information, e.g. displaying the position of the object in a pre-defined map, relating it to the position of another such object, or when the position is used to guide the object to a position known in an external coordinate system, the possibility of propagated error is a problem.
[0005] Outward-inward optical positioning systems scale very poorly to larger positioning systems, because at each point the object has to be seen by several cameras in order to triangulate the 3D position of the object. This is not economically feasible, especially for large spaces where only a few objects are tracked.
[0006] The position and orientation of a camera, for example mounted on a drone, can be outlined in a state and tracked over time. However, existing methods for determining the state of a camera do not provide a sufficient level of precision, thus rendering them insufficient for use in many applications.
[0007] It is an object of the present invention to alleviate at least some of the drawbacks associated with methods for determining the state x k of a camera known from the prior art. SUMMARY
[0008] According to a first aspect of the present invention, a method for determining the state x k of a camera comprises the steps recited in claim 1. Further optional features and embodiments of the method of the present invention are described in the dependent patent claims.
[0009] The present invention relates to a method for determining a state x k of a camera at a time t k , the state x k being an implementation of a state random variable X k , wherein the state is related to a state space model of movements of the camera. The method comprises the steps of: a) receiving an image of a scene of interest in an indoor environment captured by the camera at a time t k , wherein the indoor environment comprises N landmarks having known positions in a world coordinate system, N being a natural number; b) receiving a state estimate k of the camera at a time t k ; c) determining positions of M features in the image, M being a natural number smaller than or equal to N, wherein there is a one-to-one mapping between the M features and the N landmarks; d) receiving distance data respectively indicating distances between the M features and corresponding M landmarks; e) determining a one-to-one mapping estimate from the M features to a set of the N landmarks using at least (i) the positions of the M features in the image and (ii) the state estimate; f) using the determined one-to-one mapping estimate to establish an observation model in the state space model, wherein the observation model is configured for mapping the state random variable X k to a joint observation random variable Z k , wherein at a time t k , an observation z k is a joint observation random variable Zk and wherein the observation z k includes (i) a position of at least one of the M features in the image and (ii) the distance data indicative of a distance; andg) using (i) the state estimate, (ii) the observation model, and (iii) an observation z k to determine a state x k of the camera at time t k .
[0010] The orientations of the N landmarks in the world coordinate system can also be known. Alternatively, the distance data indicative of a distance can also relate to distances between a camera center of the camera and the M landmarks corresponding to the M features.
[0011] In principle, the number of features can also be greater than N in case outliers are detected as features. In this case, M will be greater than N. Such outliers can be removed during different processing steps: for example, they can be removed during determination of the injective map estimate; outliers can also be removed prior to determining the injective map estimate based on (i) the received distance data indicative of a distance, (ii) the state estimate, and (iii) the known positions of the N landmarks in the world coordinate system, for example, by excluding features that cannot be identified to a seemingly reasonable landmark with respect to the respective distance data indicative of a distance. It can thus be assumed - in case outliers are present - that such outliers are removed: the M features are features corresponding to actual landmarks.
[0012] In an embodiment of the method according to the invention, the joint observation random variable Z k comprises M observation random variables Z k,i , i = 1,..., M, wherein each of the M observation random variables Z k,i comprises a distance data random variable D k,i , and wherein the observation z k comprises the observations z k,i , i = 1,..., M.
[0013] In a further embodiment of the method according to the invention, the observation model is configured to model a 3D to 2D projection of each of the M landmarks corresponding to the M features to the corresponding feature, respectively. The injective map estimate, subsequently referred to as IME, links the M features to the M landmarks, wherein for a feature i of the M features, the corresponding landmark is landmark IME(i), and wherein for a feature-landmark pair (i, IME(i)), the observation model links an observation random variable Z k,i to a state random variable X k : Z k,i = h IME(i) (Xk ), where the observation model function h IME(i) (·) depends on the landmark IME(i), and the observation model mentioned therein includes the observation model function h. IME(i) (·), i=1…,M.
[0014] In another embodiment of the method according to the invention, the observation model function h IME(i) (·) is configured to set the state random variable X k Mapping to distance data random variable D k,i Above, where the distance data d indicates the distance. k,i Involves the distance between feature i and landmark IME(i), d k,i It is a random variable D that is a distance from the data. k,i The implementation of.
[0015] Alternatively, the distance data d indicates the distance. k,i and distance data random variable D k,i This can involve the distance between the camera center and the landmark IME(i).
[0016] In another embodiment of the method according to the invention, (i) state estimation is used. (ii) the observation model, and (iii) the observation z k Determine state x k This is accomplished by using update equations, which are provided by applying an extended Kalman filter to the state-space model, wherein the update equations include a Jacobian array of the observation model, wherein the Jacobian array is used for state estimation. The site was evaluated.
[0017] In another embodiment of the method according to the invention, for each observation model function h of i = 1, ..., M IME(i) (·) Use a separate Jacobian array, and wherein the update equation using the corresponding separate Jacobian array is called consecutively and independently for all M features.
[0018] In another embodiment of the method according to the invention, the distance data indicating distance is implemented as distances provided by a time-of-flight (TOF) camera as the distance between the TOF camera and the M landmarks corresponding to the M features.
[0019] A TOF camera can have a camera center and determine the distance between its camera center and M landmarks corresponding to M features.
[0020] The TOF camera functionality can be provided as part of the camera. Alternatively, the TOF camera can be a separate device. In case the TOF camera is a separate device, it can be assumed that the coordinate transformation between the TOF camera and the camera is known. The measurements performed by the TOF camera can then be transferred into the local coordinate system of the camera for comparison with the image captured by the camera.
[0021] In a further embodiment of the method according to the application, the image is captured by the camera when the light source is operated to emit light illuminating the scene of interest.
[0022] In a further embodiment of the method according to the application, the distance data indicative of the distance is implemented as intensity information of each of the M features.
[0023] The term intensity information can for example refer to an average intensity of a feature or a maximum intensity of a feature. The average intensity and the maximum intensity of a feature can be determined from pixels that capture the feature, which is part of the image captured by the image sensor.
[0024] In a further embodiment of the method according to the application, each of the M observation model functions h IME(i) (·) comprises a respective illumination model configured to map the state random variable X k to a respective distance data random variable D k,i above, wherein the illumination model i for i = 1,..., M uses at least (i) a power of light emitted by the light source and an estimated light source position, (ii) a directionality of the light emission of the light source, (iii) a reflectivity of the landmark IME(i), and (iv) a known position of the landmark IME(i) in a world coordinate system, for mapping the state random variable X k to the distance data random variable D k,i for i = 1,..., M.
[0025] In case the geometric relationship of the light source and the camera is known, the state estimate can be estimated from the estimated light source position.
[0026] In a further embodiment of the method according to the application, the M observation model functions each comprise a camera model of the camera.
[0027] In a further embodiment of the method according to the application, the camera model is implemented as a pinhole camera model.
[0028] According to a further aspect of the application, there is provided a computer program product comprising instructions which, when executed by a computer, cause the computer to perform a method according to the application.
[0029] According to a further aspect of the application, there is provided an assembly comprising (a) a camera, (b) a plurality of landmarks, and (c) a controller, wherein the controller is configured to perform a method according to the application.
[0030] In embodiments of the assembly according to the application, the assembly further comprises a time-of-flight (TOF) camera and / or a light source.
[0031] The assembly can comprise a camera and a separate TOF camera. It can be assumed that a known coordinate transformation between the camera and the separate TOF camera is known, which means that measurements obtained by either camera can be converted between the respective local coordinate systems of the two cameras. BRIEF DESCRIPTION OF DRAWINGS
[0032] Exemplary embodiments of the application are disclosed in the specification and illustrated by the accompanying drawings, wherein:
[0033] Figure 1 A schematic depiction of a method according to the application for determining a state x k of a camera at a time t k is shown; and
[0034] Figure 2 A schematic depiction of a drone comprising a light source and a camera is shown, wherein the drone is configured to fly in an indoor environment, wherein landmarks are arranged at a plurality of positions in said indoor environment. DETAILED DESCRIPTION
[0035] Figure 1 A schematic depiction of a method according to the application for determining a state x k of a camera at a time t k is shown. The state x k may comprise a 3D position and a 3D orientation of the camera at a time t k . The 3D position and the 3D orientation can be expressed relative to a world coordinate system as a predefined frame of reference. The state x k may additionally comprise 3D velocity information of the camera at a time t k , wherein said 3D velocity information may, for example, also be expressed relative to the world coordinate system. Since the camera can move over time in the indoor environment, it can be necessary to track its state to determine a current position and orientation of the camera.
[0036] A schematic depiction of a method according to the application for determining a state x k, the camera can capture an image 1 of a scene of interest in an indoor environment comprising N landmarks. The positions (possibly also orientations) of the N landmarks in the indoor environment are known in a world coordinate system. Since at time t k The camera has a certain position and orientation, so it can not be that all N landmarks are visible to the camera. For example, at time t k , it can be that J < N landmarks are visible to the camera, which J landmarks are projected by the camera onto the image 1 of the scene of interest. The projection of a landmark into an image is referred to as a "feature". From the J landmarks projected onto the image 1, M < J features can be identified, and their 2D positions in the image are determined 3. The 2D position of a feature can relate to, for example, the 2D position of the centroid of said feature. Some of the J landmarks can be positioned and oriented at time t k in such a way that their projection onto the image is too small / dark / difficult to detect for the camera. In this case, M can be strictly smaller than J, i.e. M < J, and the remaining J - M landmarks projected by the camera onto the image 1 can be ignored / not detected. The features can be determined using, for example, a scale-invariant feature transform, or using a speeded up robust features detector, or using a histogram of oriented gradients detector, or using any other feature detector known from prior art, or using a custom feature detector tailored to the likely shape of a landmark in the indoor environment. It is also assumed that the M features are features corresponding to the projection of a landmark onto the image, i.e. outliers are removed from the image, which outliers are the projection of other objects than landmarks onto the image.
[0037] The image 1 is captured by an image sensor of the camera. The image sensor has a position and an orientation in the world coordinate system, wherein said position and orientation of the image sensor at time t k can be implicitly encoded in the state x k of the camera. A feature having a certain 2D position in the image thus also has a 3D position in space, wherein the 3D position corresponds to the 3D position of the point on the image sensor corresponding to the certain 2D position of the feature.
[0038] In a next step, a surjective mapping estimate from the M features to the N landmarks is determined 5. Since it is generally assumed that M < N, the surjective mapping estimate is generally only injective but not surjective. The surjective mapping estimate describes which of the N landmarks induces which of the M features in the image. To determine 5 such a surjective mapping estimate, the position / orientation of the camera at time t k may be needed. However, not the current state x k , but only the state estimate 2 is available. Starting from the state estimate 2, the surjective mapping estimate can be determined, wherein during the surjective mapping estimate, a state xk An injective mapping estimate is a function from one set to another set. It can be denoted in function notation as IME(·), where the injective mapping estimate is configured to operate on a domain of a set of M features and a range of a set of N landmarks: feature i is linked to landmark IME(i) through the injective mapping estimate.
[0039] Using the determined feature-to-landmark assignment IME(·), in a next step, an observation model is established 6. The observation model is configured to map the state random variable X k to a joint observation random variable Z k wherein the state x k is an implementation of the state random variable, and the joint observation random variable is called joint because it probabilistically describes observation values related to the M features. The observation values z k are implementations of the joint observation random variable Z k , wherein the observation values are obtained through an actual measurement process, or by performing calculations on data provided by the actual measurement process. The joint observation random variable Z k may comprise M observation random variables Z k,i , i = 1,..., M, wherein each of the M observation random variables can statistically describe observation values related to the respective feature. The M observation random variables can be statistically independent from each other, or the joint observation random variable can comprise a probability distribution that is not accounted for in the product of the probability distributions of the M observation random variables.
[0040] The observation model can comprise M observation model functions h IME(i) (·), i = 1,..., M, wherein each of the M observation model functions can be configured to map the state random variable X k to a respective observation random variable Z k,i , i = 1,..., M. Each observation random variable Z k,i , i = 1,..., M, can comprise a distance data random variable D k,i , i = 1,..., M, and a random variable related to the 2D position of feature i, i = 1,..., M, respectively. The distance data d k,iD, i = 1,..., M, can be a realization of a distance data random variable. The presence of a distance data random variable in the observation random variable means that a quantity related to the distance between feature i and its corresponding landmark is measured. The corresponding landmark can be the landmark that actually caused the feature i (by the projection of the camera) and the observation value associated with the feature i. In case the injective mapping estimate of 5 assigns the features to the landmarks in the correct way, the corresponding landmark can be equal to the landmark IME(i). The term distance between feature i and its corresponding landmark can relate to the distance between the known 3D position of said corresponding landmark and the 3D position of said feature i in the world coordinate system. Instead of the distance between the 3D position of a feature and its corresponding landmark, the distance between the camera center of the camera and the corresponding landmark can be used.
[0041] Distance data random variable D k,i , i = 1,..., M, can statistically model the intensity of a feature, and / or the actual distance between a feature and its corresponding landmark. The intensity of a feature includes, for example, information about the distance between the feature and its corresponding landmark, because the intensity of a feature typically decreases with increasing distance between the feature and its corresponding landmark. The method according to the present invention receives 4 distance data indicating distances as part of the observations z k .
[0042] Thus, the observation model models the mapping of the M landmarks, which correspond to the M features via the determined injective mapping estimate, onto the image plane on which the image sensor is located— according to the state random variable X k . The observation model can also include processing steps, for example, for extracting the 2D position of the projected landmark (i.e., feature), e.g., the centroid of the feature in the image. The observation model can include a mathematical camera model, for example, implemented as a pinhole camera model, which mathematically describes the projection of a point in three-dimensional space onto the image plane on which the image sensor of the camera is located. In order to map the state random variable X k to the distance random variable D k,i , i = 1,..., M, in case the distance random variable relates to the intensity of a feature, the observation model can include an illumination model, or it can include a distance estimation model for determining the distance between feature i and landmark IME(i).
[0043] The illumination model can model the power loss of light emitted by a light source between emission by the light source and reception by the camera. A landmark can be implemented as a retroreflector with a reflectivity specific to the retroreflector, and a light source can be used to illuminate the retroreflector. Light reflected by the retroreflector can then appear brightly in an image 1 captured by the camera. The illumination model can include the reflectivity of the landmark IME(i) that reflects the emitted light. The illumination model can also include the distance (possibly with statistical uncertainty) between feature i and its landmark IME(i), where the distance can be based on the known position of the landmark IME(i) in the world coordinate system and the state estimate . The illumination model can also include the power of the light emitted by the light source, the directionality of the light source, and the estimated light source position. The estimated light source position can be determined using the state estimate given the known relative position and orientation of the camera with respect to the light source. The illumination model can be part of the observation model.
[0044] In the case that the distance random variables D k,i , i = 1,..., M statistically model the actual distance between feature i, i = 1,..., M and landmark IME(i), i = 1,..., M, respectively, the distance can be measured using a time-of-flight (TOF) camera. The TOF camera can be a phase-based TOF camera, or a pulse-based TOF camera. The TOF camera can provide the distance between a feature and its corresponding landmark. The TOF camera functionality can be part of the camera, or the TOF camera can be a separate device. In the case that the TOF camera and the camera are separate devices, the geometric transformation between the TOF camera and the camera can be known, meaning that the measurements performed using the TOF camera can be related to the measurements performed by the camera.
[0045] The observation model is part of a state space model used to track the movement of the camera in space. In addition to the observation model, the state space model can typically include a state transition model. The state transition model describes how the state itself evolves over time. For example, in the case that the camera is mounted on a drone, the state transition model can include equations that model the flight of the drone, which can include control inputs used to control the flight of the drone. The state transition model typically also includes additional terms that model the statistical uncertainty in the state propagation. The observation model and / or the state transition model can be linear or nonlinear in their input, which is the state of the camera. The state transition model can also take control inputs as input.
[0046] In the case that both the observation model and the state transition model are linear, a Kalman filter can be used to determine 7 the state estimate at time t using the observation model and the observation value z k .k the state x k 8, which is a joint observation random variable Z k The observations comprise (i) 2D positions of the M features in the image and (ii) distance data indicative of distances, e.g. implemented as intensities of the features or measured distances between the features and their respective corresponding landmarks. During the determination 7 of the state x k the state estimate is used as input to the observation model (alternatively, an approximation to the state determined during the determination of the injective map estimate can be used as input to the observation model). In case the observation model and / or the state transition model are non-linear, an extended Kalman filter can be used, wherein the extended Kalman filter linearizes the non-linear equations. Both the Kalman filter and the extended Kalman filter provide an update equation for updating the state estimate using the observation model and the measured observations. Once the state x k 8 has been determined 7, it can be propagated in time, e.g. from time t k to time t k+1 , using the state transition model, providing a state estimate of the state of the camera at time t k+1 instead of the Kalman filter, a particle filter can be used, or a state observer such as a Luenberger observer can be used, or any other known filtering technique known from the prior art. The state estimate can be used as input for determining a new state estimate of the state x k+1 at time t k+1 .
[0047] The update equation of the Kalman filter or the extended Kalman filter can be invoked once for all M features or separately for each of the M features. In case the extended Kalman filter is used, the Jacobian determinant of the observation model needs to be computed with respect to the state and evaluated at the state estimate (or alternatively at an approximation to the state determined during the determination of the injective map estimate). In case the extended Kalman filter is invoked separately for each of the M features, a separate Jacobian determinant can be determined for each of the M observation model functions h IME(i) (·), i = 1..., M. In case the time t k+1 - t k between the capturing of consecutive images by the camera is not long enough to process all M features, not all M features can be taken into account during the update of the state.
[0048] Figure 2 A schematic depiction of a drone, including a light source 10 and a camera 11, is shown, where the drone is flying in an indoor environment 15. Landmarks 9, implemented as retroreflectors in this particular example, are arranged at multiple locations within the indoor environment 15. The landmarks 9 may be mounted on the ceiling of the indoor environment 15. In any given posture of the drone (including position and orientation), some of the landmarks 9 may be visible to the camera 11—in Figure 2 The location of the drone is indicated by a line between landmark 9 and camera 11—while other landmarks 9 may not be visible to camera 11. The position of landmark 9 may be known in world coordinate system 12, which serves as a predefined reference frame 12, and the current position of the drone may be expressed in drone coordinate system 13, which serves as a second reference frame 13, where the coordinate transformation 14 between world coordinate system 12 and drone coordinate system 13 may be known. With camera 11 and light source 10 rigidly mounted to the drone and their poses relative to the drone known, the poses of camera 11 and light source 10 can be correlated with world coordinate system 12 using drone coordinate system 13. The current position of the drone can be determined using an image of scene 15 of interest (specifically landmark 9 with a known position) in the indoor environment 15. Alternatively or additionally, the drone may be equipped with an inertial measurement unit (IMU) which can also be used for pose determination. Light source 10 may be an isotropically emitted light source, or it may be a directional light source emitted in an isotropic manner. Light source 10 and camera 11 are ideally close to each other, especially when landmark 9 is implemented as a retroreflector. Camera 11 can also be mounted on top of the drone, i.e., adjacent to light source 10, during the drone's normal movement. The term "normal movement" can refer to the drone's usual movement relative to the ground of the scene of interest. The drone may additionally include a time-of-flight (TOF) camera for directly measuring distances to landmark 9. TOF camera functionality can be provided by a separate TOF camera, or it can be included within camera 11.
Claims
1. A method for determining a state x k of a camera at a time t k , the state x k being an implementation of a state random variable X k , wherein the state relates to a state space model of movement of the camera, the method comprising: a) receiving at time t k an image of a scene of interest in an indoor environment captured by the camera, wherein the indoor environment comprises N landmarks having known positions in a world coordinate system, N being a natural number; b) receiving a state estimate of the camera at time t k c) determining positions of M features in the image, M being a natural number smaller than or equal to N, wherein there is a one-to-one mapping between the M features and the N landmarks; d) receiving distance data indicative of distances between 3D positions of the M features and M landmarks, respectively, projections of the M landmarks being the M features; e) determining a one-to-one mapping estimate from the M features to the set of N landmarks using at least (i) the positions of the M features in the image and (ii) the state estimate; and f) using a determined injective mapping estimate to establish an observation model in the state space model, wherein the observation model is configured for mapping a state random variable X k of the camera to a joint observation random variable Z k , wherein at time t k , an observation z k is an implementation of the joint observation random variable Z k , and wherein the observation z k comprises (i) a position of at least one of the M features in the image and (ii) the distance data indicative of a distance.
7. The method of claim 1, wherein the distance data indicative of distances are implemented as distances, the distances being provided by a time-of-flight (TOF) camera as distances between the TOF camera and the M landmarks corresponding to the M features, respectively. g) using (i) the state estimate, (ii) the observation model, and (iii) an observation z k to determine the state x k of the camera at time t k .
2. The method of claim 1, wherein the joint observation random variable Z k includes M observation random variables Z k,i , i = 1,..., M, where each of the M observation random variables Z k,i includes a distance data random variable D k,i , and where the observation z k includes the observations z k,i , i = 1,..., M.
3. The method of claim 2, wherein the observation model is configured to model a 3D to 2D projection of each of the M landmarks corresponding to the M features, respectively, to a corresponding feature, and wherein the subsequent single-shot map estimation, referred to as IME, links the M features with the M landmarks, wherein, For feature i in the M features, the corresponding landmark is landmark IME(i), and wherein for the feature-landmark pair (i, IME(i)), the observation model relates the observation random variable Z k,i to the state random variable X k : Z k,i = h IME(i) (X k ), wherein the observation model function h IME(i) (·) depends on the landmark IME(i), and wherein the observation model comprises observation model functions h IME(i) (·), i = 1...M.
4. The method of claim 3, wherein the observation model function h IME(i) (·) is configured to map a state random variable X k to a distance data random variable D k,i wherein the distance data d k,i indicates a distance between a feature i and a landmark IME(i), d k,i is an implementation of the distance data random variable D k,i .
5. The method of claim 1, wherein the state estimate (ii) the observation model, and (iii) the observation z k to determine the state x k is done by using an update equation, which is provided by applying an extended Kalman filter to the state space model, wherein the update equation includes a Jacobian array of the observation model, wherein the Jacobian array is evaluated at the state estimate .
6. The method of claim 5, wherein for each observation model function h IME(i) (·) uses separate Jacobian arrays, and wherein the update equations using respective separate Jacobian arrays are invoked sequentially and independently for all M features.
8. The method of claim 1, wherein the image is captured by the camera when a light source is operated to emit light illuminating the scene of interest.
9. The method of claim 8, wherein the distance data indicative of distances are implemented as intensity information of each of the M features.
11. The method of claim 1, wherein the M observation model functions each comprise a camera model of the camera.
10. The method of claim 9, wherein each of the M observation model functions h IME(i) (·) i=1,...,M includes a respective illumination model configured to map the state random variable X k to a respective distance data random variable D k,i above, which statistically models the intensity information, wherein the illumination model i for i = 1,...,M uses at least (i) a power of light emitted by the light source and an estimated light source position, (ii) a directionality of the light emission of the light source, (iii) a reflectivity of the landmark IME(i), and (iv) a known position of the landmark IME(i) in a world coordinate system, for mapping the state random variable X k to the distance data random variable D k,i for i = 1,...,M.
12. The method of claim 11, wherein the camera model is implemented as a pinhole camera model.
13. A computer program product comprising instructions which, when executed by a computer, cause the computer to perform the method of claim 1.
14. An assembly comprising (a) a camera, (b) a plurality of landmarks, and (c) a controller, wherein the controller is configured to perform the method of claim 1.
15. The assembly of claim 14, further comprising a time-of-flight (TOF) camera and / or a light source.