Method and device for determining the absolute pose of a camera on board an aerial or spacecraft
The method addresses the challenge of accurately determining the absolute pose of a camera on board an aerial or spacecraft by using a sequence of images to estimate a local 3D model and comparing it with a reference 3D model, achieving robustness against viewing condition and seasonal changes.
Patent Information
- Application Number
- FR2023005008
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-05-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-05-24
AI Technical Summary
Existing vision-based navigation systems face challenges in accurately determining the absolute pose of a camera on board an aerial or spacecraft due to variations in scene viewing conditions and seasonal changes, which can lead to significant errors in position estimation.
A method is proposed that involves obtaining a sequence of images from the camera, estimating a local 3D model of the scene, and determining the absolute pose of the camera by comparing this local 3D model with a predetermined reference 3D model, thereby being robust to changes in viewing conditions and seasonal variations.
This approach enables accurate and robust estimation of the absolute pose of the camera, improving navigation accuracy regardless of changes in scene illumination or seasonal variations, and is applicable to various types of cameras used for navigation.
Smart Images

Figure 00000025_0000 
Figure 00000025_0001 
Figure 00000026_0000
Abstract
Description
Title of the invention: Method and device for determining the absolute pose of a camera on board an aerial or spacecraft Technical field
[0001] The present invention belongs to the general field of vision-based navigation (VBN) and relates more particularly to a method for determining an absolute pose, in a reference frame, of a camera on board an aerial or spacecraft. State of the art
[0002] In vision-based navigation systems, it is notably known to acquire an image of a scene, by means of a camera mounted on a mobile vehicle relative to this scene, and to compare this image to a reference image of said scene. This comparison aims to register the image acquired by the camera with the reference image, that is to say to find and reposition the acquired image within the reference image. The reference image being georeferenced, it is then possible to estimate the position of the camera, and thus that of the vehicle carrying it, relative to the scene.
[0003] However, many factors can affect the accuracy of the vehicle position estimation.
[0004] For example, if the reference image is generally an image acquired in good observation conditions, for example in the absence of clouds and / or with a scene generally illuminated by the Sun (in particular in the case of a reference image acquired in visible wavelengths), this is not necessarily the case for the image acquired by the camera on board the craft whose position is to be estimated. If the image is acquired by the camera on board the air or space craft while the scene is partly masked by clouds and / or is poorly illuminated by the Sun, then the comparison with the reference image may prove complicated, which may lead to significant errors in the estimation of the position of the craft relative to the scene.
[0005] Furthermore, the scene itself may be subject to seasonal variations. Thus, if the reference image represents the scene in summer and if the craft flies over this same scene in winter, the comparison of the reference image with the image acquired by the camera on board this craft may prove complicated. Statement of the invention
[0006] The present invention aims to remedy all or part of the drawbacks of the prior art, in particular those set out above, by proposing a solution of vision-based navigation that is robust to variations in scene viewing conditions and seasonal variations.
[0007] For this purpose, and according to a first aspect, a method is proposed for determining an absolute pose of a camera, in a reference frame, said camera being monocular and passive, and being on board an aerial or spacecraft mobile relative to a scene, said method comprising: - obtaining a sequence of images of said scene acquired by the camera at different respective times, each image comprising a plurality of pixels, each pixel of an image representing a portion of the scene seen by the camera at the time of acquisition of said image, - an estimation, from the sequence of images, of a local 3D model in a reference frame centered and oriented relative to the camera, said local 3D model representing a part of the scene in three dimensions corresponding to the acquisition of a so-called target image among the sequence of images, - a determination of the absolute pose of the camera at the time of acquisition of the target image, by comparison of the local 3D model in the camera frame of reference with a predetermined reference 3D model corresponding to said scene represented in three dimensions in the reference frame.
[0008] The proposed method therefore aims to estimate the absolute pose of the camera in a reference frame. By "absolute pose" is meant the pose (position and attitude) of the camera in this reference frame. The absolute pose is to be distinguished from the "relative pose" of the camera, which corresponds to the pose of the camera in an arbitrary frame which corresponds for example to a frame of the camera at the instant of acquisition of a previous image of the sequence (in which case the relative pose corresponds to the variation in pose of the camera between two instants of image acquisition).
[0009] Several images acquired by the camera are used to determine a local 3D model of the scene. It should be noted that we are considering here a passive type camera (i.e. one which simply passively measures electromagnetic radiation coming from the scene, unlike, for example, lidars or radars which actively measure by emitting electromagnetic radiation in the direction of the scene and measuring the electromagnetic radiation reflected by said scene). In addition, since the camera is monocular, it is therefore a single camera which is used to acquire the sequence of images, which are therefore acquired at different times by the same camera (monocular vision system) which has in principle moved between two acquisition times and therefore observes the scene from in principle different viewpoints.
[0010] Furthermore, it should be noted that the term "3D model" here means a representation of the 3D geometry of the scene, and that any type of representation of the 3D geometry can be considered in the present disclosure. Such a 3D model can therefore be presented, for example, in the form of a simple Digital Elevation Model, DEM (Digital Elevation Model, sometimes referred to as a "pseudo 3D" or "2.5D" model), or in a more complex form making it possible to represent in space the external envelope of the volume formed by the scene (for example a mesh consisting of vertices, edges and polygonal surfaces), etc.
[0011] Then, the absolute pose of the camera (and of the aerial or spacecraft which carries it), in the reference frame, is estimated by comparing the local 3D model (estimated by means of the sequence of images) of a part of the scene with a reference 3D model of said scene in the reference frame.
[0012] Given that the reference 3D model is established in the reference frame, the fact of realigning (in position and in attitude) the local 3D model with respect to the reference 3D model makes it possible to estimate the absolute pose of the camera in the reference frame, that is to say to estimate both the position and the attitude of said camera in said reference frame.
[0013] Furthermore, when comparing 3D models, one is essentially interested in the 3D geometry of the scene, which is independent of the conditions of observation of the scene and which in principle varies little with seasonal variations. Similarly, the 3D geometry of the scene does not depend on the type of physical characteristics measured of the scene. For example, a reference 3D model established from measurements (active or passive) carried out in visible wavelengths can quite easily be used to determine the absolute pose of a passive type camera measuring electromagnetic radiation in infrared wavelengths.
[0014] The present disclosure therefore proposes a vision-based navigation solution which is advantageously applicable regardless of the type of camera used for navigation. Indeed, since the images acquired by this camera are used to estimate a local 3D model which represents the 3D geometry of the scene (and not the electromagnetic radiation coming from the scene in a particular band of wavelengths), the same reference 3D model can be used to determine the absolute pose of any type of camera used for navigation.
[0015] In particular embodiments, the determination method may further comprise, optionally, one or more of the following characteristics, taken in isolation or in all technically possible combinations.
[0016] In particular modes of implementation, the estimation of the local 3D model comprises: - an estimation of a relative pose of the camera for several images of the image sequence, - a matching of the pixels of the target image with pixels of other images among the sequence of images, so that the matched pixels are pixels of different images representing the same portion of the scene, - for each pixel of the target image matched with pixels of other images: an estimate of a 3D position, in the camera frame of reference, of the portion of the scene represented by said pixel of the target image, as a function of the estimated relative poses, of the position of said pixel in the target image and of the positions in the other images of the pixels matched with said pixel of the target image, in which the local 3D model is formed from the estimated 3D positions of the matched pixels of the target image.
[0017] In particular embodiments, the matching of the pixels of the target image with the pixels of another image among the sequence of images comprises: - a determination of a registration transformation between the target image and said other image among the sequence of images, from the estimated relative poses, - a registration, by means of the registration transformation, of the target image and of said other image among the sequence of images, - a comparison of the target image and said other image among the sequence of images, obtained after registration.
[0018] In particular embodiments, the registration transformation is a homography.
[0019] In particular modes of implementation, the comparison of the target image and said other image among the sequence of images, obtained after registration, implements a dense optical flow algorithm.
[0020] In particular embodiments, the estimation of the relative pose of the camera for several images of the image sequence comprises: - a determination of the relative poses of the camera for several images of the image sequence by visual odometry, - an update of the relative poses determined by beam adjustment.
[0021] In particular modes of implementation, the determination of the absolute pose of the camera at the time of acquisition of the target image comprises: - obtaining an approximate absolute pose of the camera at the time of acquisition of the target image, - a projection of the local 3D model of a part of the scene at the time of acquisition of the target image towards the reference frame, according to the approximate absolute pose, - a comparison, in the reference frame, of the projected local 3D model with the reference 3D model.
[0022] In particular embodiments, the approximate absolute pose is determined from an estimated absolute pose for the camera during an acquisition of a previous target image.
[0023] In particular embodiments, the local 3D model corresponds to a digital elevation model and / or the reference 3D model corresponds to a digital elevation model.
[0024] In particular embodiments, the local 3D model and the reference 3D model are filtered using a high-pass filter before being compared to determine the absolute 3D pose of the camera.
[0025] In particular embodiments, the comparison between the local 3D model and the reference 3D model comprises a correlation of said local 3D model with said reference 3D model.
[0026] In particular implementation modes, the images of the image sequence obtained correspond to so-called key images selected from among the so-called initial images successively acquired by the camera.
[0027] In particular implementation modes, an initial image is selected as a key image when a predetermined criterion of movement of the aerial or space vehicle since the acquisition of the previous key image is verified.
[0028] In particular implementation modes, the camera is sensitive in visible and / or infrared wavelengths.
[0029] According to a second aspect, there is provided a computer program product comprising instructions which, when executed by at least one processor, configure said at least one processor to implement a determination method according to any one of the implementation modes of the present disclosure.
[0030] According to a third aspect, a calculation device is proposed comprising at least one processor and at least one memory, said at least one processor being configured to implement a determination method according to any one of the implementation modes of the present disclosure.
[0031] According to a fourth aspect, there is provided an aerial or spacecraft, comprising a camera and a computing device according to any one of the embodiments of the this disclosure. It should be noted that spacecraft means any craft operating outside the Earth's atmosphere, including a craft operating on the ground on a celestial body other than Earth (satellite, space shuttle, rover, etc.) or flying over such a celestial body. Presentation of figures
[0032] The invention will be better understood on reading the following description, given by way of non-limiting example, and made with reference to the figures which represent: • [Fig.l] [Fig.l]: a schematic representation of an example of a machine carrying a camera for acquiring images of a scene relative to which the machine is likely to move, • [Fig.2] [Fig.2]: a diagram illustrating the main steps of an example of implementation of a method for determining an absolute pose of the camera on board the machine, • [Fig.3] [Fig.3]: a diagram illustrating the main steps of an example of implementation of a step of estimation of a local 3D model of the scene of the absolute pose determination process, • [Fig.4] [Fig.4]: a diagram illustrating the main steps of an example of implementation of a step of estimation of relative poses of the camera for different images acquired by the camera, • [Fig.5] [Fig.5]: a diagram illustrating the main steps of an example of implementation of a pixel matching step of the local 3D model estimation step of the scene, • [Fig.6] [Fig.6]: a schematic representation of pixels of two images representing the same elements of a scene, • [Fig.7] [Fig.7]: a diagram illustrating the main steps of an example of implementation of an absolute pose determination step of the absolute pose determination method, • [Fig.8] [Fig.8]: a schematic representation of an example of a local 3D model before and after high-pass filtering.
[0033] In these figures, identical references from one figure to another designate identical or similar elements. For reasons of clarity, the elements represented are not to scale, unless otherwise indicated.
[0034] Furthermore, the order of steps shown in these figures is given solely as a non-limiting example of the present disclosure which may be applied with the same steps performed in a different order. Description of the embodiments
[0035] As indicated above, the present invention relates to a method 20 for determining an absolute pose of a camera 11 in a reference frame, the camera 11 being on board an aerial or spacecraft 10 mobile relative to a scene. It should be noted that the term “spacecraft” means any craft moving outside the Earth’s atmosphere, including a craft moving on the ground on a celestial body other than the Earth (satellite, space shuttle, rover, etc.). An aerial craft can be any craft flying in the Earth’s atmosphere (airplane, helicopter, drone, etc.).
[0036] By “absolute pose” is meant the pose of the camera 11 in the reference frame considered, that is to say the position (3D) and the attitude (orientation) of said camera 11 in this reference frame.
[0037] The reference frame may be of any type suitable for locating the position and attitude of the camera 11 in a three-dimensional space and is typically defined by an origin and by three non-coplanar axes, for example orthogonal. In certain cases, the reference frame considered depends on the scene flown over. For example, if the scene flown over is a scene on the surface of the Earth, the reference frame may be a geocentric frame or a frame having as its origin a particular point on the surface of the Earth, etc. If the craft 10 flies over another type of celestial body, for example another planet or an asteroid, the reference frame is for example centered on this celestial body or on a particular point on the surface of this celestial body, etc.
[0038] [Fig.l] schematically represents an exemplary embodiment of a space or air vehicle 10. As illustrated by [Fig.l], the vehicle 10 comprises the camera 11 and a computing device 12.
[0039] The camera 11 may be of any type suitable for acquiring two-dimensional (2D) images of the scene, passively (i.e., simply measuring electromagnetic radiation from the scene, without first emitting electromagnetic radiation towards the scene). The camera 11 is configured to measure electromagnetic radiation in one or more bands of determined lengths. For example, the camera 11 is configured to measure electromagnetic radiation in visible wavelengths, i.e., wavelengths between 380 nanometers (nm) and 780 nm. Alternatively or additionally, the camera 11 may be configured to measure electromagnetic radiation in infrared wavelengths, i.e., wavelengths between 780 nm and 5 millimeters (mm).For example, camera 11 is sensitive to near infrared (« Near Infrared », NIR, in the English literature, that is, in wavelengths between 780 nm and 3 micrometers (μm)) and / or to mid infrared (« Mid Infrared », MIR, in the English literature, that is, in wavelengths between 3 μm and 50 μm).
[0040] The images acquired by camera 11 are presented, for example, in the form of a pixel matrix providing physical information on the area of the scene located in the field of view of camera 11. The images are, for example, Nx x Ny pixel matrices, Nx and Ny being, for example, each of the order of a few hundred to a few tens of thousands of pixels. The images are typically acquired recurrently by camera 11, for example periodically with a frequency that is, for example, of the order of a few Hertz (Hz) to a few hundred Hz.
[0041] It should be noted that a monocular vision system (as opposed in particular to stereoscopic), comprising a single camera 11, is sufficient for implementing the method 20 for determining absolute pose. The camera 11 is implemented to acquire a sequence of images of the scene, which are therefore acquired at different times by the same camera (monocular vision system) which has in principle moved between two acquisition times and therefore observes the scene from points of view which are in principle different.
[0042] It should be noted that the machine 10 may include vision sensors other than the camera 11, which are outside the scope of the present disclosure, and the present disclosure concerns the case where the absolute pose in the reference frame is estimated from images acquired by the camera 11 alone.
[0043] The calculation device 12 is configured to implement all or part of the steps of the method 20 for determining the absolute pose of the camera 11. In the non-limiting example illustrated by [Fig.l], the calculation device 12 is on board the space or air vehicle 10. It is however possible, in other examples, to have a calculation device 12 which is not on board the vehicle 10 and which is remote from said vehicle 10. Where appropriate, the images acquired by the camera 11 are transmitted to the calculation device 12 by any type of suitable communication means.
[0044] The computing device 12 comprises for example one or more processors (CPU, DSP, GPU, FPGA, ASIC, etc.). In the case of several processors, these may be integrated into the same equipment and / or integrated into hardware separate equipment. The computing device 12 also comprises one or more memories (magnetic hard disk, electronic memory, optical disk, etc.) in which is for example stored a computer program product, in the form of a set of program code instructions to be executed by the processor(s) to implement all or part of the steps of the method 20 for determining the absolute pose of the camera 11.
[0045] [Fig. 2] schematically represents the main steps of a method 20 for determining the absolute pose of the camera 11 in the reference frame. It should be noted that the position and orientation of the camera 11 relative to the machine 10 which the onboard being known or determinable at any time, the absolute pose of the camera 11 in the reference frame can for example be used to determine the absolute pose of the space or air vehicle 10 in the reference frame, for example for vision-based navigation purposes, possibly in combination with navigation measurements provided by navigation sensors (GPS receiver, accelerometer, odometer, gyroscope, etc.) possibly on board the vehicle 10.
[0046] As illustrated in FIG. 2, the method 20 for determining absolute pose comprises a step S20 of obtaining a sequence of images of said scene. The sequence of images obtained during step S20 are images acquired by the camera 11 at different respective times and are therefore likely to represent the scene from different points of view. For example, the sequence of images obtained during step S20 comprises a predetermined number Nc of images, with N c - -• Preferably, Nc > 3 or Nc > 5. For example, 6 < Nc <10.
[0047] According to a first example, it is possible to use all the images acquired successively (for example periodically) by the camera 11, so that the sequence of images obtained corresponds to images acquired successively by the camera 11.
[0048] According to another example, it is possible not to use, to determine the absolute pose of the camera 11, all the images acquired by said camera. If we designate by “initial images” all the images acquired successively by the camera, then the Nc images obtained during step S20, designated hereinafter by “key images”, are images selected from the initial images acquired successively by the camera 11.
[0049] Generally speaking, such a selection of key images, which amounts to discarding certain initial images, aims to try to increase the probability of obtaining a sequence of key images representing the scene from points of view which are actually different. Indeed, since the time difference between the acquisition times of two successive initial images may be small, the variation in point of view may sometimes be negligible between two successive initial images, and this may further depend on the movement of the machine 10 relative to the scene. For example, it is possible to select, as a key image, an initial image every nc initial images. In this case, if the initial images are acquired at a period then the key images are acquired at a period nc ■ Th nc being for example equal to or greater than 10, or for example equal to or greater than 100.
[0050] According to another example, the selection can implement a predetermined criterion of movement of the machine 10 since the acquisition of the previous key image, and an initial image is selected as a key image when this movement criterion is verified. The evaluation of the movement of the machine 10, even approximate, can implement any method known to the person skilled in the art. For example, it is possible to use an estimate of the displacement of the machine 10 provided by a navigation filter, or measurements provided by an inertial unit of the machine 10, to evaluate whether the displacement criterion is verified (i.e. whether the estimated displacement is greater than a minimum required displacement). In preferred modes of implementation, the evaluation of the displacement is carried out from the content of the initial images. For example, from a given key image (the first key image being able to be chosen arbitrarily), it is possible to compare the content of an initial image with the content of this key image. If the contents of the initial image and the key image considered are substantially different, this means that the machine 10 has moved and the initial image considered can be selected as another key image.For example, it is possible to identify characteristic vignettes (a characteristic vignette corresponding to a set of pixels representing a characteristic element of the scene) in the key image considered and to pursue these characteristic vignettes in the following initial images ("tracking" in the Anglo-Saxon literature). Where appropriate, the displacement criterion is for example considered to be verified if a predetermined percentage of reference vignettes could not be pursued in the initial image considered.
[0051] In the remainder of the description, the case is considered in a non-limiting manner where the sequence of images obtained during step S20 are key images selected from the initial images acquired successively by the camera 11. One of these key images, from the sequence of key images obtained during step S20, is designated by “target image”, and corresponds to the key image with respect to which the absolute pose of the camera 11 must be determined. In other words, the absolute pose to be determined corresponds to the absolute pose at the time of acquisition of the target image. The target image may be any one of the key images from the sequence of key images. Preferably, the target image corresponds to the key image acquired last in the sequence, that is to say the one whose acquisition time is the oldest.
[0052] In order to determine successive absolute poses of the camera 11, it is for example possible to consider a sliding sequence of key images. For example, the sequence of key images (i.e. the Nc key images forming the sequence of key images) previously used to determine the absolute pose of the camera 11 can be updated during step S20 by deleting the oldest key image from this sequence and adding to this sequence a new key image which has just been selected from among the last initial images acquired by the camera 11.
[0053] As illustrated by [Fig.2], the method 20 for determining absolute pose comprises a step S21 of estimating, from the sequence of key images, a local 3D model of the scene. The local 3D model represents the 3D geometry of a part of the scene as seen by the camera 11 at the time of acquisition of the target image. The local 3D model represents the 3D geometry of a part of the scene in a frame of reference of the camera 11, that is to say in a frame of reference whose origin and orientation are defined relative to the camera 11 (and which is different from the reference frame considered).
[0054] Thus, the sequence of key images, which represent a part of the scene as seen from different viewpoints, is used to reconstruct the 3D geometry of this part of said scene. Assuming that most of the constituent elements of the scene are immobile, the different key images, acquired at different respective times and representing the scene as seen from different respective viewpoints, can in fact be used to perform a reconstruction of the 3D geometry of a part of the scene, visible at the time of acquisition of the target image, by implementing multi-view 3D reconstruction methods (based on the principle of structure acquired from a movement, or "Structure From Motion", SFM, in the English literature).Generally speaking, any multi-view 3D reconstruction method known to those skilled in the art can be implemented during step S21 of estimating the local 3D model, and the choice of a particular method constitutes only a non-limiting variant of implementation of the determination method 20.
[0055] [Fig.3] schematically represents the main steps of an example of implementation of step S21 of estimating the local 3D model of the scene in the camera frame of reference.
[0056] As illustrated by [Fig.3], step S21 of estimating the local 3D model comprises in this example a step S210 of estimating a relative pose of the camera 11 for all or part of the key images of the sequence of key images.
[0057] As indicated previously, the relative pose of a key image corresponds to the pose of the camera 11 in an arbitrary reference frame, the position and / or orientation of which in the reference frame are not known a priori (or at least not with sufficient precision). For example, the arbitrary reference frame may be the reference frame of the camera 11 at the time of acquisition of one of the key images of the sequence, in which case the relative pose of the camera 11 at the time of acquisition of another of the key images of the sequence corresponds to the (absolute) pose variation between the respective times of acquisition of the two key images considered. The arbitrary reference frame may also vary from one key image to another, for example if it is the reference frame of the camera 11 at the time of acquisition of the previous key image in the sequence of key images.
[0058] For example, the estimation step S210 can implement a visual odometry method, which makes it possible, for example, to estimate the variation in pose from one key image to another, by analyzing the content of the key images.
[0059] [Fig.4] schematically represents the main steps of an example of implementation of step S210 of estimating the relative poses of the camera 11. As illustrated by [Fig.4], step S210 comprises: - a step S2100 of determining the relative poses of the camera 11 for all or part of the key images of the sequence of key images by visual odometry, - a step S2101 of updating the relative poses determined by beam adjustment.
[0060] Thus, in this example, the relative poses are estimated in at least two stages. A first estimation of the relative poses is first obtained during step S2100 by visual odometry, for example by continuing from one key image to another pixels (or pixel thumbnails) supposed to represent the same portions of the scene. It should be noted that, in certain implementation examples, this step S2100 of determining the relative poses of the camera 11 can be implemented to determine the relative pose of the camera 11 for each initial image acquired by said camera 11 (i.e. not only for the key images of the sequence of key images).
[0061] However, in some cases, the accuracy of the relative poses determined by visual odometry may be limited due to a possible accumulation of errors (the error made on the estimated relative pose for an image may propagate from one image to another). The step S2101 of updating by beam adjustment aims to improve the accuracy of the relative poses determined during the step S2100.
[0062] The beam adjustment is based on the principle that, for two key images acquired by the camera 11 at different acquisition times, the lines of sight, coming from the camera 11 and associated with pixels representing the same portion of the scene in the two key images, must intersect at the level of said same portion. Pixels of different key images representing the same portion of the scene are said to be "matched". Several pixels of the same key image are respectively matched with several pixels of one or more other key images during the visual odometry step S2100, to determine the relative poses.During step S2101, the relative poses are updated to ensure that the lines of sight associated with pixels of different images, mapped to each other, substantially intersect in the same portion (point), and this for several portions of the scene which are represented in several key images. In practice, it is generally not possible to have lines of sight which strictly intersect, and “substantially intersect” means that said lines of sight are at least close to intersecting.
[0063] For example, it is possible to use a Levenberg-Marquardt algorithm to update the relative poses of the first and last keyframes in the sequence using an epipolar error (in pixels) as a metric. This allows one to determine approximate 3D positions, in an arbitrary coordinate system, of the portions of the scene identified as visible in both the first and last keyframes in the sequence (the approximate 3D positions corresponding to the coordinates at which the corresponding lines of sight substantially intersect). Then, the relative poses of the other keyframes in the keyframe sequence can be updated from these approximate 3D positions using a Perspective-n-Point (PNP) problem-solving algorithm, for example the SolvePnP function in the OpenCV library.
[0064] As illustrated by [Fig.3], step S21 of estimating the local 3D model comprises in this example a step S211 of matching pixels of the target image with pixels of other key images among the sequence of key images. As indicated previously, pixels matched together correspond to pixels of different images which are considered to represent the same portion (or point) of the scene, which is therefore in principle seen from different points of view. It should be noted that the matching carried out during step S211 is advantageously a so-called “dense” matching, that is to say a matching of all the pixels of the target image with pixels of other key images among the sequence of key images.Obviously, it is not always possible, for a given pixel of the target image, to identify a pixel of another key image representing the same portion of the scene (since, the machine 10 being mobile, this portion was not necessarily in the field of view of the camera 11 during the acquisition of this other key image). However, this search for corresponding pixels in the other key images is preferably carried out for each pixel of the target image, in order to identify a large number of portions of the scene visible in several key images including the target image. In particular, if a pixel matching is carried out during the step S210 of estimating relative poses, this concerns a much smaller number of pixels than the number of pixels matched during the step S211 of dense matching.Such arrangements make it possible to improve the accuracy and resolution of the local 3D model and, ultimately, to improve the accuracy of determining the absolute pose of the camera 11.
[0065] The matching step S211 may implement any dense matching method known to those skilled in the art, and the choice of a particular method constitutes only a non-limiting variant of the implementation of the local 3D model estimation step S21. For example, the matching step S211 may implement a dense optical flow algorithm.
[0066] [Fig.5] schematically represents the main features of an example of implementation of the matching step S211. In this example, the matching step S211 uses the relative poses estimated during the step S210. For example, for matching the pixels of the target image with the pixels of another key image among the sequence of key images, it is possible to determine, during a step S2110, a registration transformation between the target image and this other key image, from the relative poses estimated for these two key images. The registration transformation aims to bring the target image and this other key image into close acquisition geometries, for example by bringing this other key image into an acquisition geometry close to that of the target image (or vice versa). For example, the registration transformation makes it possible to predict the pixel of the other key image theoretically representing the same portion of the scene.In fact, the positions of the pixels representing the same portion of the scene may vary from one key image to another, especially due to the change in the pose of the camera 11 relative to the scene.
[0067] Figure 6 schematically represents two successive key images and, in these key images, pixels representing the same portions of the scene. More particularly, pixel P] and pixel p^ represent the same portion of the scene, pixel P2 and pixel p'2 represent the same portion of the scene, and pixel P 3 and pixel p'^ represent the same portion of the scene. However, the position of pixel Pi in the left key image is different from the position of pixel p^ in the right key image (similarly for pixels P2 and p\, and for pixels P3 and p'-^). The registration transformation aims to try to predict, from the position of a pixel in the target image, the position of the pixel representing the same portion of the scene in the other key image (or vice versa).
[0068] In some cases, the registration transformation can be determined by making simplifying assumptions, in particular on the geometry of the scene. For example, in some cases, the registration transformation corresponds to a homography. In a manner known per se, a homography models the movement by assuming that the portions of the scene represented by the pixels are located in the same plane, that is to say by assuming that the scene is generally a flat surface. Such a homography can be modeled in a simple manner and also proves to be sufficient in many cases.
[0069] However, nothing excludes, in other examples, modeling the approximate geometry of the scene differently for determining the registration transformation. It should be noted that this modeling of the approximate geometry of the scene is only intended to establish the registration transformation that is used to facilitate pixel matching, and is therefore distinct from the local 3D model estimated during step S21.
[0070] The registration transformation is used, during a step S2111, to register the target image and the other key image with each other, that is to say to bring the target image and the other key image as close as possible to the same acquisition geometry. The appearance of an element of the scene in an image varies according to the acquisition point of view of this image. This registration therefore makes it possible to obtain a target image which more closely resembles the other key image. For example, the target image is brought as close as possible to the acquisition geometry of the other key image, that is to say that the target image is used to predict (by applying the registration transformation to the target image), a predicted image representing the scene from the point of view of the other key image, which therefore more closely resembles this other key image than the target image, and which can therefore be more easily compared to this other target image.
[0071] In the example illustrated by [Fig.5], the matching step S211 then comprises a step S2112 of comparing the target image and said other image among the sequence of images, obtained after registration. During the comparison step S2112, the images obtained after registration are compared to enable identification of all the pixels that can be matched with each other. Such a comparison can be carried out according to any method known to the person skilled in the art. For example, this comparison can implement a correlation between pixel thumbnails of the target image and the key image obtained after registration, or a dense optical flow algorithm, etc.
[0072] Generally speaking, the use of the registration transformation during matching, to register the images considered, makes it possible to improve the robustness and the precision of the matching. Different types of registration transformation can be used, which can in particular model the approximate geometry of the scene differently. The use of a homography, however, has the advantage of being simple to determine and to apply, while making it possible to obtain good results in many cases for the dense matching of pixels.
[0073] As illustrated by [Fig.3], step S21 of estimating the local 3D model comprises, for each pixel of the target image matched with at least one pixel of at least one other key image among the sequence of key images, a step S212 of estimating a 3D position, at the time of acquisition of the target image in the frame of reference of the camera 11, of the portion of the scene represented by said pixel of the target image. The estimation of the 3D position, relative to the camera 11 at the time of acquisition of the target image, of each portion of the scene represented by pixels mapped to each other in different keyframes, takes into account the estimated relative poses and positions in the keyframes of the different pixels mapped to each other. 3D position estimation, for example, triangulates the pixels mapped to each other, by estimating the 3D position of the scene portion considered as the intersection of the lines of sight associated respectively with the pixels mapped to each other (or at least the 3D position at which the lines of sight substantially intersect).
[0074] At the end of the 3D position estimation step S212, a plurality of 3D positions of different portions of the scene, relative to the camera 11 at the time of acquisition of the target image, have been estimated. These estimated 3D positions of the portions of the scene therefore describe the 3D geometry, in the frame of reference of the camera 11, of a part of the scene as seen by said camera 11 at the time of acquisition of the target image. These estimated 3D positions of the portions of the scene therefore make it possible to form the local 3D model. As indicated previously, the matching carried out during the step S211 is a dense matching, so that the local 3D model obtained represents the 3D geometry of this part of the scene with a 2D resolution of the order of the resolution of the target image.For example, the local 3D model may be in the form of a 2D image consisting of pixels associated with portions of the same dimensions (resolution) as the pixels of the target image. However, the value of a pixel of the 2D image of the local 3D model represents the third dimension, for example in the form of the distance between the camera 11 and the portion of the scene represented by this pixel, while the value of a pixel of the target image represents a physical quantity representative of the electromagnetic radiation coming from the portion of the scene represented by this pixel.
[0075] As illustrated by [Fig.2], the determination method 20 then comprises a step S22 of determining the absolute pose of the camera 11 at the time of acquisition of the target image, by comparing the local 3D model in the frame of reference of the camera 11 with a predetermined reference 3D model of said scene in the reference frame.
[0076] The reference 3D model represents the 3D geometry of the scene in the reference frame. It is therefore a priori information on the scene, which represents the 3D geometry of the scene correctly oriented, positioned and dimensioned in the reference frame. For example, the reference 3D model corresponds to a Digital Terrain Model (DTM) or, preferably, to a DEM correctly oriented, positioned and dimensioned in the reference frame (for example georeferenced in the case of a scene on the surface of the Earth). In a manner known per se, a DTM represents the 3D geometry of the ground of the scene without taking into account the various elements located above the ground (buildings, vegetation, etc.), unlike a DEM which represents the 3D geometry of the scene taking into account elements located above the ground.
[0077] Such a reference 3D model can be previously established according to any 3D mapping method known to those skilled in the art, and the choice of a particular method does not constitute a variant in any way limiting the implementation of the method 20 for determining absolute pose.
[0078] It should be noted that the reference 3D model can in particular be determined by applying the same steps as for determining the local 3D model, as a function of a sequence of images acquired by a camera whose position and attitude in the reference frame are known precisely during the acquisition of the sequence of images (for example determined by means of a GPS receiver, etc.). It should be noted that, in certain cases, the camera used to establish the reference 3D model of the scene may be the camera 11 of the machine 10. For example, in the case where the machine 10 makes a round trip flying over the same scene, and if the machine 10 is equipped for example with a GPS receiver, it is possible to determine the reference 3D model during the outward journey of the machine 10, and to use the method 20 for determining absolute pose during the return journey of the machine 10, for navigation purposes of the machine 10, for example to compensate for a failure of the GPS receiver.
[0079] The reference 3D model depends on the scene and the computing device 12 is therefore configured to retrieve the reference 3D model associated with the scene flown over, for example from a database which may be on board the space or air craft 10, or remote from said craft 10. For example, the database may store a single reference 3D model established specifically for a given mission of overflying the associated scene. In other examples, the database may store several reference 3D models associated respectively with different scenes which may be flown over. In such a case, the computing device 12 selects and retrieves the reference 3D model to be used as a function of information making it possible to identify the scene which must be flown over by the craft 10 (for example from approximate coordinates of the scene or the craft 10, etc.).In the case described above where the machine 10 makes a round trip and establishes the reference 3D model on the outward journey, it is sufficient during the return journey to retrieve from the database the reference 3D model which has just been established.
[0080] The comparison of the local 3D model with the reference 3D model aims to recalibrate the local 3D model with respect to the reference 3D model. This recalibration is carried out in position and attitude. It should be noted that the recalibration in position is carried out in three dimensions (3D position), that is to say that the recalibration aims to find not only the place (essentially 2D position) within the reference 3D model where the 3D geometry described by the local 3D model is found, but also a scale factor that describes the resizing required to locally match the local 3D model with the reference 3D model. The scale factor is induced by the distance between the camera 11 and the scene and therefore makes it possible to find the altitude component (third dimension) of the 3D position of the absolute pose to be determined. For example, the comparison (registration) is carried out by correlating the local 3D model with the reference 3D model, or according to any registration method known to the person skilled in the art, for example by a method for detecting and describing salient elements, by an “Iterative Closest Point” (ICP) type method, etc.
[0081] Given that the reference 3D model is established in the reference frame, the fact of realigning the local 3D model with respect to the reference 3D model makes it possible to determine the absolute pose of the camera 11 in the reference frame, at the time of acquisition of the target image, that is to say to determine both the position (3D) and the attitude of the camera 11 in the reference frame.
[0082] As indicated above, the registration of the local 3D model with respect to the reference 3D model aims to find the local 3D model within the reference model. In certain cases, such registration can perform a scan in a search domain which corresponds to ranges of possible registration values for the position (2D position and scale factor) and the attitude, in order to identify registration values which make it possible to optimize a predetermined resemblance function, representative of the resemblance between the local 3D model (registrated according to the registration values considered) and the reference 3D model. For example, the resemblance function corresponds to a correlation between the local 3D model (registrated according to the registration values considered) and the reference 3D model.
[0083] It should be noted that the scale factor may, in certain cases, be provided by other means, at least approximately. For example, the scale factor may be determined from measurements provided by an inertial unit of the craft 10 (for example predicted from these measurements by means of a navigation filter). If the scale factor thus determined is considered sufficiently precise, then it is not necessary to further rescale during step S22, for example it is not necessary to perform a scan in a range of possible values for the scale factor. Otherwise, the approximate scale factor may for example be used to reduce the search domain of the scale factor during step S22.
[0084] Generally, any a priori information about the absolute pose can be used to reduce the search domain (and to reduce the probability of false detection). In some cases, it is possible to obtain an approximate absolute pose of the camera 11 which can be used to reduce the search domain, and therefore to accelerate and improve the registration of the local 3D model with the reference 3D model. For example, the approximate absolute pose for the current target image can be determined from the absolute pose previously determined for the acquisition time of the previous target image and / or from measurements provided by an inertial unit of the machine 10, etc. In such a case, the step S22 of determining the absolute pose of the camera 11 aims to improve the accuracy of the determined absolute pose compared to the approximate absolute pose.
[0085] [Fig.7] schematically represents the main steps of an example of implementation of step S22 of determining the absolute pose of the camera 11 in the reference frame.
[0086] As illustrated by [Fig.7], the comparison step S22 comprises in this example : - a step S220 of obtaining an approximate absolute pose of the camera 11 at the time of acquisition of the target image, - a step S221 of projecting the local 3D model of a part of the scene at the time of acquisition of the target image towards the reference frame, as a function of the approximate absolute pose, - a step S222 of comparison, in the reference frame, of the projected local 3D model with the reference 3D model.
[0087] The projection step S221 essentially corresponds to a change of reference frame (with change of scale), from the frame of the camera 11 to the reference frame. This change of reference frame is carried out as a function of the approximate absolute pose which approximately describes the position (3D) and the attitude of the frame of the camera 11 in the reference frame. Starting for example from a local 3D model which corresponds to a 2D image describing the distances of the different portions of the scene relative to the camera 11, a local 3D model is obtained for example which corresponds to a local DEM (local in that it represents a part of the scene represented by the reference 3D model) in the reference frame. The comparison step S222 corresponds to the registration described above, carried out on a reduced search domain, limited to a neighborhood of the approximate absolute pose obtained during the step S220.As indicated above, if the scale factor of the approximate absolute pose is considered sufficiently accurate, there is no need to further rescale in step S222.
[0088] In particular modes of implementation, the local 3D model (possibly after projection into the reference frame) and the reference 3D model are filtered using a high-pass filter before being compared to determine the absolute pose of the camera 11. Such arrangements make it possible to facilitate the comparison between the local 3D model and the reference 3D model. For example, in the case where the model 3D local and the reference 3D model correspond to DEMs, then the application of high-pass filtering makes it possible to highlight high-frequency variations in elevation which correspond, for example, to the presence of buildings in the scene. Generally speaking, buildings, or more generally structures locally introducing rapid variations in elevation within the scene, and their respective positions are important for determining the absolute pose of the camera 11, and are in particular more relevant and more easily exploitable than absolute values of elevation. Conversely, slow variations in elevation within the scene (for example, the slope of the ground on which a building is built) carry little information and can also introduce noise during a comparison by correlation of the local 3D model with the reference 3D model.For example, the resolution of rapid elevation changes to be preserved is of the order of a pixel in the local 3D model (DEM) of the scene. The high-pass filter parameters are therefore advantageously selected to preserve such rapid elevation changes.
[0089] For example, [Fig.8] schematically represents an example of a local 3D model (projected into the reference frame using an approximate absolute pose of the camera 11) before high-pass filtering (part (a) of [Fig.8]) and after high-pass filtering (part (b) of [Fig.8]). In the example illustrated by [Fig.8], the local 3D model corresponds to a DEM. As illustrated by part (a) of [Fig.8], the elevation varies essentially with the slope of the scene and rapid variations in the elevation of the scene are sometimes negligible compared to the absolute elevation. In part (b) of [Fig.8], rapid variations in the elevation are highlighted using high-pass filtering and it can be seen that these carry more information for determining the absolute pose of the camera.We understand in fact that rapid variations in elevation within a scene characterize it more (for example in relation to another scene) than slow variations in elevation within this scene.
[0090] Another advantage of high-pass filtering is that it removes edge effects. Indeed, as shown in part (a) of [Fig.8], the local 3D model is only defined locally (the areas on the left, right, and top edges of part (a) are not defined), which introduces elevation discontinuities on the edges when comparing with the reference 3D model. These discontinuities, which could lead to the presence of artifacts during correlation, are removed by high-pass filtering, as shown in part (b) of [Fig.8].
[0091] More generally, it should be noted that the implementation and embodiment modes considered above have been described as non-limiting examples, and that other variants are consequently conceivable.
Claims
1.
2.
3. Claims Method (20) for determining an absolute pose of a camera (11), in a reference frame, said camera (11) being monocular and passive, and being on board an aerial or spacecraft (10) mobile relative to a known scene and mapped in the form of a 3D reference model, said method comprising: - obtaining (S20) a sequence of at least two consecutive images of said scene acquired by the camera (11) at respective consecutive times, each image comprising a plurality of pixels, each pixel of an image partially representing the scene seen by the camera at the time of acquisition of said image, - an estimation (S21), from the sequence of images, of a local 3D model in a reference frame centered and oriented relative to the camera (11), said local 3D model representing a part of the scene in three dimensions corresponding to the acquisition of a so-called target image among the sequence of images, the estimation (S21) of the local 3D model comprising a dense matching (S211) of all the pixels of the target image with pixels of other images among said sequence, - a provision (S220) of an approximate absolute pose of the camera, - a determination of the absolute pose of the camera (11) at the time of acquisition of the target image, by recalibration in position and attitude of the local 3D model with the reference 3D model, the recalibration in position and attitude being carried out in a research domain, from the approximate absolute pose of the camera. The method (20) of claim 1, wherein determining the local 3D model comprises determining relative poses of the camera for multiple images in the sequence. The method (20) of claim 2, wherein determining the relative poses of the camera for said plurality of images of the sequence comprises visual odometry (S2100), followed by updating (S2101) the relative poses by beam adjustment.
4. Method (20) according to any one of claims 2 and 3, wherein the estimation (S21) of the local 3D model comprises: - said estimation (S210) of relative poses of the camera (11) for said several images of the sequence, followed by - said dense matching (S211) of all the pixels of the target image with pixels of said several images of the sequence, so that the matched pixels are pixels of different images representing the same portion of the scene, - for each pixel of the target image actually matched: an estimation (S212) of a 3D position, in the frame of reference of the camera, of the portion of the scene represented by said pixel of the target image, as a function of the estimated relative poses, of the 2D position of said pixel in the target image and of the 2D positions in said several images of the pixels mapped to said pixel of the target image,in which the local 3D model is formed from the estimated 3D positions of the actually matched pixels of the target image.,
5. Method (20) according to claim 4, in which the dense matching (S211) of the pixels of the target image with the pixels of the other images of the sequence comprises, for each other image among the other images of the sequence: - a determination (S2110), from the estimated relative poses, of a registration transformation between the target image and said other image, - a registration (S2111), by means of the registration transformation, of the target image and said other image, - a comparison (S2112) of the target image and said other registered image.
6. The method (20) of claim 5, wherein the registration transformation is a homography.
7. Method (20) according to any one of claims 5 and 6, wherein the comparison of the target image and said other registered image implements a dense optical flow algorithm.
8. Method (20) according to any one of the preceding claims, in which the determination (S22) of the absolute pose of the camera (11) at the time of acquisition of the target image comprises: - a projection (S221) of the local 3D model of a part of the scene at the time of acquisition of the target image towards the reference frame, as a function of the approximate absolute pose, - a comparison (S222), in the reference frame, of the projected local 3D model with the reference 3D model.
9. A method (20) according to any preceding claim, wherein the approximate absolute pose is determined from an estimated absolute pose for the camera (11) during an acquisition of a previous target image.
10. A method (20) according to any preceding claim, wherein the local 3D model and the reference 3D model are filtered using a high-pass filter before being compared to determine the absolute pose of the camera (11).
11. Method (20) according to any one of the preceding claims, in which the images of the sequence of images obtained correspond to so-called key images selected from a sliding sequence of so-called initial images acquired successively by the camera.
12. The method (20) of claim 11, wherein an initial image is selected as a key image when a predetermined criterion of movement of the air or spacecraft (10) since the acquisition of the previous key image is satisfied.
13. A method (20) according to any preceding claim, wherein the camera (11) is sensitive in visible and / or infrared wavelengths.
14. A computer program product comprising instructions which, when executed by at least one processor, configure said at least one processor to implement a method (20) according to any one of the preceding claims.
15. Computing device (12) comprising at least one processor and at least one memory, said at least one processor being configured to implement a method (20) according to any one of claims 1 to 13.
16. Aerial or spacecraft (10), comprising a camera (11) and a computing device (12) according to claim 15.