Method and device for determining absolute pose of camera located on aircraft or spacecraft

By capturing image sequences with a monocular passive camera and matching and aligning the local 3D model with the reference 3D model, the pose determination error problem of the visual navigation system under cloud occlusion and seasonal changes is solved, and robustness and real-time performance under different conditions are achieved.

CN121532800AActive Publication Date: 2026-02-13AIRBUS DEFENCE AND SPACE(FR) +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202480029248.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-24
Filing Date
2024-05-24
Publication Date
2026-02-13
Estimated Expiration
2044-05-24

AI Technical Summary

Technical Problem

Existing visual navigation systems struggle to accurately determine the position of a vehicle relative to the scene when faced with cloud cover and seasonal changes, resulting in significant pose determination errors.

Method used

A monocular passive camera is used to capture scene image sequences. The absolute pose of the camera is determined in real time by matching and aligning the local 3D model with the reference 3D model. Dense matching and bundle adjustment are used to improve accuracy.

Benefits of technology

It achieves robust visual navigation under different observation conditions and seasonal changes, and can accurately determine the absolute pose of the camera in the reference coordinate system in real time, and is suitable for various types of cameras.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121532800A_ABST
    Figure CN121532800A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method (20) for determining the absolute pose of a camera (11) in a reference coordinate system, the camera (11) being located on a vehicle (10) movable relative to a scene, the method comprising:-acquiring (S20) a sequence of images of the scene, the images being respectively captured by the camera at different moments in time, the sequence being viewable as a video stream, -determining (S21) a local 3D model in the camera coordinate system from the sequence of images, said local 3D model representing a part of the scene at the time of image capture, referred to as the target image, in the sequence of images,-determining (S22) the target image by realigning the position and pose of the local 3D model in the camera coordinate system with a predetermined reference 3D model, the reference 3D model corresponds to the scene represented three-dimensionally in a reference coordinate system, determining (S22) an absolute pose of the camera at the moment of target image capture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the broad field of visual navigation (also referred to as VBN in the literature), and more specifically relates to a method for determining the absolute pose of a camera located on an aircraft or spacecraft in a reference coordinate system. Background Technology

[0002] In visual navigation systems, it is known to use a camera located on a vehicle moving relative to the scene to capture an image of the scene, and then compare this image with a reference image of the scene. This comparison aims to align the image captured by the camera with the reference image, i.e., to find and reposition the captured image within the reference image. Since the reference image is georeferenced, the position of the camera relative to the scene can be determined, thereby determining the position of the vehicle carrying the camera relative to the scene.

[0003] However, many factors can affect the accuracy of determining the launch vehicle's location.

[0004] For example, while reference images are typically captured under favorable viewing conditions, such as cloudless conditions and / or with the scene largely illuminated by sunlight (especially in the case of reference images captured at visible wavelengths), this may not be the case for images captured by a camera on a vehicle whose pose needs to be determined. If the camera captures an image when parts of the scene are obscured by clouds and / or receive very little sunlight, comparisons with reference images can become complicated, leading to significant errors in determining the vehicle's position relative to the scene.

[0005] Furthermore, the scene itself may be affected by seasonal changes. Therefore, if the reference image represents a summer scene, while the launch vehicle is flying over the same scene in winter, comparing the reference image with an image captured by a camera on the launch vehicle can become complicated. Summary of the Invention

[0006] The present invention aims to address some or all of the shortcomings of the prior art, particularly the aforementioned shortcomings, by proposing a vision-based navigation solution that is robust to changes in scene observation conditions and seasonal changes.

[0007] Therefore, and according to the first aspect, a real-time method (20) is proposed for determining the absolute pose of a camera (11) in a reference coordinate system, the camera (11) being monocular and passive, located on an aircraft or spacecraft (10) capable of moving relative to a known scene mapped in the form of a reference 3D model, the method comprising:

[0008] -Acquire (S20) a sequence of at least two consecutive images captured by the camera (11) at corresponding consecutive time points, each image comprising a plurality of pixels, each pixel of the image partially representing the scene observed by the camera at the time the image was captured.

[0009] - Determine (S21) a local 3D model from the image sequence in a coordinate system centered on the camera (11) and oriented relative to the camera, the local 3D model representing a three-dimensional portion of the scene corresponding to the capture of the image (referred to as the target image), the determination of the local 3D model (S21) including dense matching (S211) of all pixels of the target image with pixels of other images in the sequence,

[0010] - Provides an approximate absolute pose for the (S220) camera.

[0011] - By realigning the position and pose of the local 3D model with the reference 3D model, position and pose realignment is performed in the search domain based on the approximate absolute pose of the camera to determine the absolute pose of the camera (11) at the moment of target image capture.

[0012] Therefore, the proposed method aims to determine the absolute pose of the camera in a reference coordinate system of a 3D map. Here, "absolute pose" is understood as the camera's pose (position and orientation) in that reference coordinate system. Advantageously, the present invention allows the method for determining absolute pose to be used for visual navigation, particularly real-time navigation. Absolute pose differs from the camera's "relative pose," which corresponds to the camera's pose in any coordinate system, such as the camera's pose in the coordinate system at the time of the previous image capture in a sequence (in which case, relative pose corresponds to the change in camera pose between two image capture times).

[0013] Multiple images captured by a camera are used to determine a local 3D model of the scene. It should be noted that a passive camera is considered here (i.e., only passively measuring electromagnetic radiation from the scene, unlike active measurements such as those using lidar or radar that emit electromagnetic radiation into the scene and measure the electromagnetic radiation reflected by the scene). Furthermore, since the camera is a monocular camera, a single camera is used to capture a sequence of images, which are therefore captured at different times by the same camera (monocular vision system), which in principle has moved between the two capture moments, thus observing the scene from different viewpoints.

[0014] Furthermore, it should be noted that the term "3D model" here refers to a representation of the 3D geometry of a scene, and any type of 3D geometry representation can be envisioned in this disclosure. Thus, such a 3D model can be, for example, in the form of a simple digital elevation model (DEM) (sometimes referred to as a "pseudo-3D" or "2.5D" model), or in a more complex form (e.g., a mesh consisting of vertices, edges, and polygonal surfaces) that allows for a spatial representation of the volume formed by the scene.

[0015] Since the reference 3D model is built in a reference coordinate system, the absolute pose of the camera in the reference coordinate system can be determined by realigning the local 3D model with the reference 3D model (position and orientation). That is, the position and orientation of the camera in the reference coordinate system can be determined.

[0016] Furthermore, the primary focus is on the 3D geometry of the scene, which is independent of the viewing conditions and, in principle, varies minimally with the seasons. Similarly, the 3D geometry of the scene is independent of the type of physical properties measured within the scene. For example, a reference 3D model established using measurements taken at visible wavelengths (active or passive) can be used to determine the absolute pose of a passive camera measuring electromagnetic radiation at infrared wavelengths.

[0017] Therefore, this disclosure proposes a vision-based navigation solution that is particularly available in real time and advantageously applicable to any type of camera used for navigation. In fact, since the images captured by the camera are used to represent a local 3D model of the scene's 3D geometry (rather than the scene's electromagnetic radiation in a specific wavelength band), the same reference 3D model can be used to determine the absolute pose of any type of camera used for navigation.

[0018] In some specific implementations, the method for determining absolute pose may also optionally include one or more of the following features, individually or in any technically possible combination.

[0019] In some specific implementations, determining the local 3D model includes determining at least one relative pose of the camera for several consecutive images in a sequence.

[0020] In some specific implementations, determining the relative pose of the camera for several consecutive images in the sequence includes visual range measurement, followed by updating the relative pose via beam adjustment.

[0021] In some specific implementations, the determination of the local 3D model includes:

[0022] - Determine the relative pose of the camera for several consecutive images in the sequence, and then

[0023] - Perform dense matching of all pixels in the target image with pixels in several consecutive images in the sequence.

[0024] - For each pixel that is actually matched in the target image: Based on the relative pose, the 2D position of the pixel in the target image, and the 2D position of each matched pixel, determine the 3D position of the pixel in the camera coordinate system.

[0025] The local 3D model is formed based on the 3D positions of the actual matched pixels in the target image.

[0026] In some specific implementations, dense matching of pixels in the target image with pixels in other consecutive images in the sequence includes, for each of the other consecutive images in the sequence:

[0027] Based on the corresponding relative pose, determine the realignment transformation between the target image and the other images.

[0028] - Realign the other images using a realignment transformation.

[0029] - Determine the residual motion from the target image to other realigned images.

[0030] In some specific implementations, realignment is transformed into homography.

[0031] In some specific implementations, the residual motion of the target image and other realigned images is determined using a dense optical flow algorithm.

[0032] In some specific implementations, determining the absolute pose of the camera at the moment of target image capture includes:

[0033] - Project the local 3D model onto the reference coordinate system based on approximate absolute pose.

[0034] - Match the projected local 3D model with the reference 3D model in the reference coordinate system.

[0035] In some specific implementations, the approximate absolute pose is determined based on the absolute pose of the camera determined during previous target image capture or based on navigation instruments carried on the aircraft or spacecraft.

[0036] In some specific implementations, a high-pass filter is used to filter the local 3D model and the reference 3D model before determining the absolute pose of the camera.

[0037] In some specific implementations, the images in the acquired image sequence correspond to images selected from a sliding sequence of images (called initial images) (called key images), which are captured continuously by the camera.

[0038] In some specific implementations, an initial image is selected as a key image when it meets predetermined criteria for the motion of an aircraft or spacecraft since the previous key image was captured.

[0039] In some specific implementations, the camera is sensitive to visible and / or infrared wavelengths.

[0040] According to the second aspect, a real-time method for vision-based navigation is proposed, using a monocular and passive camera attached to an aircraft or spacecraft platform capable of moving relative to a known scene mapped in the form of a reference 3D model, the method comprising:

[0041] - Determine the absolute pose of the camera in the reference coordinate system using the method described in the first aspect, then

[0042] - Determine the position and attitude of the aircraft or spacecraft in the reference 3D model based on the absolute pose corrected according to the position of the camera relative to the aircraft or spacecraft platform.

[0043] According to a third aspect, a computer program product is provided, comprising instructions that, when executed by at least one processor, configure the at least one processor to implement a method according to any aspect and / or implementation of this disclosure.

[0044] According to a fourth aspect, a computing device is provided, comprising at least one processor and at least one memory, said at least one processor being configured to implement a method according to any aspect and / or embodiment of this disclosure.

[0045] According to the fifth aspect, an aircraft or spacecraft is provided, comprising a platform carrying a camera and a computing device according to any embodiment. It should be noted that "spacecraft" is understood to mean any launch vehicle operating outside the Earth's atmosphere, including launch vehicles (satellites, space shuttles, rovers, etc.) operating on the ground of other celestial bodies outside the Earth, or launch vehicles flying over said celestial bodies. Attached Figure Description

[0046] The invention will be better understood by reading the following description, which is given by way of non-limiting example and with reference to the accompanying drawings, wherein:

[0047] -[ Figure 1 ] Figure 1 This illustration schematically shows an example of a vehicle carrying a camera, which is used to capture images of a scene in which the vehicle may be moving relative to it.

[0048] -[ Figure 2 ] Figure 2 : A diagram illustrating the main steps of an example implementation of a method for determining the absolute pose of a camera located on a carrier.

[0049] -[ Figure 3 ] Figure 3 : A diagram illustrating the main steps of an example implementation of a method for determining a local 3D model of a scene in determining absolute pose.

[0050] -[ Figure 4 ] Figure 4: A diagram illustrating the main steps of an example implementation for determining the relative pose of a camera with respect to different images captured by the camera.

[0051] -[ Figure 5 ] Figure 5 : A diagram illustrating the main steps of an example implementation of the pixel matching step in the process of determining a local 3D model of a scene.

[0052] -[ Figure 6 ] Figure 6 The illustration shows pixels from two images, representing the same element in the scene.

[0053] -[ Figure 7 ] Figure 7 : A diagram illustrating the main steps of an example implementation of the process for determining the absolute pose.

[0054] -[ Figure 8 ] Figure 8 The illustration shows the example local 3D model before and after high-pass filtering.

[0055] In these figures, the same reference numerals in different figures denote the same or similar elements. For clarity, the elements shown are not drawn to scale unless otherwise stated.

[0056] Furthermore, the order of steps shown in these figures is given only as a non-limiting example of this disclosure, and can be applied where the same steps are performed in a different order. Detailed Implementation

[0057] As described above, the present invention relates to a method 20 for determining the absolute pose of a camera 11 in a reference coordinate system, the camera 11 being mounted on an aircraft or spacecraft 10 capable of moving relative to a scene. It should be noted that "spacecraft" refers to any vehicle operating outside the Earth's atmosphere, including vehicles operating on the ground over other celestial bodies (satellites, space shuttles, rovers, etc.). An aircraft can be any vehicle flying within the Earth's atmosphere (airplanes, helicopters, drones, etc.).

[0058] "Absolute pose" refers to the pose of camera 11 in the reference coordinate system under consideration, that is, the position (3D) and orientation (orientation) of camera 11 in the reference coordinate system.

[0059] The camera is preferably calibrated, which provides better accuracy and repeatability.

[0060] To measure the position of a camera relative to the platform of an aircraft or spacecraft (10), a coordinate system linked to one or more other sensors, particularly an inertial measurement unit (IMU) or GPS, is used. Inter-sensor calibration is performed, for example, by simultaneously acquiring measurements from all sensors and estimating the geometric transformations that allow for the optimal superposition of these different measurements. In general, this involves, for example, performing a maneuver that maximizes the observability of all sensors.

[0061] Preferably, a single unit assembly exists, which is detachable from the platform and integrates multiple sensors fixed to each other. The single unit assembly can then be manually oriented in multiple directions while observing distant objects, such as buildings, the horizon, or the ground as seen from the top of a building.

[0062] The reference coordinate system can be of any type suitable for identifying the position and attitude of camera 11 in three-dimensional space, and is typically defined by an origin and three non-coplanar axes (e.g., orthogonal). In some cases, the reference coordinate system under consideration depends on the scene being flown over. For example, if the scene being flown over is a scene on the Earth's surface, the reference coordinate system can be a geocentric coordinate system or a coordinate system whose origin is located at a specific point on the Earth's surface. If spacecraft 10 is flying over other types of celestial bodies, such as another planet or asteroid, the reference coordinate system is, for example, centered on that celestial body or located at a specific point on the surface of that celestial body.

[0063] Figure 1 An example implementation of an aircraft or spacecraft 10 is illustrated schematically. Figure 1 As shown, the carrier 10 includes a camera 11 and a computing device 12.

[0064] Camera 11 can be of any type suitable for passively capturing two-dimensional (2D) images of a scene (i.e., measuring only electromagnetic radiation from the scene without first emitting electromagnetic radiation into the scene). Camera 11 is configured to measure electromagnetic radiation in one or more defined wavelength bands. For example, camera 11 is configured to measure electromagnetic radiation in visible wavelengths (i.e., wavelengths from 380 nanometers (nm) to 780 nm). Additionally or alternatively, camera 11 may be configured to measure electromagnetic radiation in infrared wavelengths (i.e., wavelengths from 780 nm to 5 millimeters (mm)). For example, camera 11 is sensitive to near-infrared (NIR, i.e., wavelengths from 780 nm to 3 micrometers (μm)) and / or mid-infrared (MIR, i.e., wavelengths from 3 μm to 50 μm).

[0065] The image captured by camera 11 is, for example, a pixel matrix that provides physical information about the scene region located within the field of view of camera 11. The image is, for example, N... x ×N y pixel matrix, N x and N yFor example, each image has a resolution of several hundred to tens of thousands of pixels. Images are typically captured by camera 11 in a cyclical manner, for example, periodically at frequencies ranging from a few hertz (Hz) to several hundred hertz.

[0066] It should be noted that the monocular vision system (especially in contrast to the stereo vision system) includes a single camera 11, which is sufficient to implement the method 20 for determining the absolute pose. The camera 11 is used to capture a sequence of images of the scene, which are thus captured at different times by the same camera (monocular vision system), which in principle has moved between the two capture moments and thus observes the scene from different viewpoints in principle.

[0067] It should be noted that the carrier 10 may include a visual sensor other than the camera 11, but the present invention makes it possible to determine the absolute pose of the camera in a reference coordinate system by using the image captured by the camera 11.

[0068] The computing device 12 is configured to implement all or part of the steps of the method 20 for determining the absolute pose of the camera 11. Figure 1 In the non-limiting example shown, the computing device 12 is mounted on the spacecraft or aircraft 10. However, in other examples, the computing device 12 may be located away from the launch vehicle 10 and not mounted on it. Where appropriate, images captured by the camera 11 may be sent to the computing device 12, and information related to the absolute pose of the camera may then be sent back via any suitable type of communication.

[0069] The computing device 12 includes, for example, one or more processors (CPU, DSP, GPU, FPGA, ASIC, etc.). In the case of multiple processors, these processors may be integrated in the same device and / or integrated in a separate hardware device. The computing device 12 also includes one or more memories (magnetic hard disk, electronic storage, optical disk, etc.), which, for example, store a computer program product, which is stored in the form of a set of program code instructions to be executed by the processor to implement all or part of the steps of the method 20 for determining the absolute pose of the camera 11.

[0070] Figure 2 The main steps of a method 20 for determining the absolute pose of camera 11 in a reference coordinate system are illustrated schematically. It should be noted that since the position and orientation of camera 11 relative to the carrier 10 carrying it are known or can be determined at any time, the absolute pose of camera 11 in the reference coordinate system can be used, for example, to determine the absolute pose of spacecraft or aircraft 10 in the reference coordinate system, for example, for vision-based navigation purposes, possibly in conjunction with navigation measurements provided by navigation sensors (GPS receiver, accelerometer, odometer, gyroscope, etc.), which may be mounted on carrier 10.

[0071] like Figure 2 As shown, the method 20 for determining absolute pose includes a step S20 of acquiring an image sequence of the scene. The image sequence acquired during step S20 consists of images captured by camera 11 at corresponding consecutive moments, and therefore may represent the scene from different viewpoints. For example, the image sequence acquired during step S20 includes a predetermined number N. c N images, where N c ≥2. Preferably, N c ≥3 or N c ≥5. For example, 6≤N c ≤10. In practice, it's important to track pixels to measure significant differences corresponding to movement, thereby improving the accuracy of 3D perception. However, the greater the difference between images, the more difficult it is to accurately track pixels. This is why using intermediate images to track pixels through intermediate images can be advantageous.

[0072] According to the first example, all images captured continuously (e.g., periodically) by camera 11 can be used, such that the acquired image sequence corresponds to N images captured continuously by camera 11. c One image.

[0073] According to another example, to determine the absolute pose of camera 11, all images captured by the camera may not be used. If "initial image" refers to all images continuously captured by the camera, then N obtained during step S20 c Each image (hereinafter referred to as a "key image") is an image selected from the initial images captured sequentially by camera 11.

[0074] Generally, this selection of key images (which is equivalent to discarding some initial images) aims to increase the probability of obtaining a sequence of key images representing the scene from practically different viewpoints. In practice, since the time difference between the capture times of two consecutive initial images can be small, the viewpoint change between two consecutive initial images can sometimes be negligible, and this may also depend on the motion of the carrier 10 relative to the scene. For example, it could be every n... c One initial image is selected as the key image. In this case, if the initial image has a period T... I If captured, the key image is captured in a period of n. c ·T I Capture, n c For example, greater than or equal to 10, or for example, greater than or equal to 100.

[0075] According to another example, a predetermined criterion for the motion of the vehicle 10 since the previous key image was captured is selected, and the initial image is selected as the key image when the motion criterion is met. The evaluation of the motion of the vehicle 10, even if approximate, can utilize any method known to those skilled in the art. For example, the motion estimate of the vehicle 10 provided by a navigation filter, or a measurement provided by the inertial measurement unit of the vehicle 10, can be used to evaluate whether the motion criterion is met (i.e., whether the estimated motion is greater than the required minimum motion). In a preferred embodiment, motion evaluation is performed using the content of the initial image. For example, using a given key image (the first key image can be arbitrarily chosen), the content of the initial image can be compared with the content of the key image. If the content of the initial image differs significantly from that of the relevant key image, it means that the vehicle 10 has moved, and the relevant initial image can be selected as another key image. For example, feature patterns (which correspond to a set of pixels representing scene feature elements) can be identified in the key images under consideration, and these feature patterns can be tracked in subsequent initial images (“tracking” in the literature). Where appropriate, for example, if a predetermined percentage of the reference pattern cannot be tracked in the initial image, the motion criterion is considered met.

[0076] The remainder of this section considers a continuous sequence of images to be processed, which may, for example, consist of key images.

[0077] In the image sequence acquired during step S20, one of these images is designated as the "target image" and corresponds to the image relative to which the absolute pose of the camera 11 to be determined is located. In other words, the absolute pose to be determined corresponds to the absolute pose at the moment the target image was captured. The target image can be any image in the image sequence. Preferably, the target image corresponds to the last captured image in the sequence, i.e., the image with the most recent capture time.

[0078] To determine the continuous absolute pose of camera 11, an image sliding sequence can be considered, for example. For instance, during step S20, the image sequence previously used to determine the absolute pose of camera 11 (i.e., N forming the image sequence) can be updated. c (Images), by removing the oldest image from the sequence and adding a new image selected from the latest initial image captured by camera 11 to the sequence.

[0079] like Figure 2As shown, the method 20 for determining absolute pose includes a step S21 of determining a local 3D model based on an image sequence. The local 3D model represents the 3D geometry of a portion of the scene as observed by the camera 11 at the moment the target image is captured. The local 3D model represents the 3D geometry of this portion of the scene in the coordinate system of the camera 11, i.e., for this coordinate system, the origin and orientation are defined relative to the camera 11 (and are different from the reference coordinate system involved).

[0080] Therefore, a sequence of images partially representing a scene and viewed from different viewpoints is used to reconstruct the 3D geometry of that portion of the scene. Assuming most of the elements of the scene are static, various images captured at corresponding consecutive moments and representing the scene from their respective viewpoints can actually be used to reconstruct the 3D geometry of the portion of the scene visible at the moment the target image is captured. This is achieved by utilizing multi-view... Figure 3 The D-reconstruction method (based on the "SFM" principle of "motion structure recovery" in the literature) is used. Generally speaking, any multi-view reconstruction method known to those skilled in the art is... Figure 3 All 3D reconstruction methods can be implemented during step S21 of determining the local 3D model, and the choice of a particular method constitutes only a non-limiting variant implementation of method 20 for determining absolute pose.

[0081] Figure 3 The schematic illustration shows the main steps of an example implementation of step S21 for determining a local 3D model of the scene in the camera coordinate system.

[0082] like Figure 3 As shown, in this example, step S21 of determining the local 3D model includes step S210 of determining one or more relative poses of camera 11 for some or all of the images in the image sequence.

[0083] As described above, the relative pose of an image corresponds to the pose of camera 11 in an arbitrary coordinate system whose position and / or orientation are initially unknown (or at least not sufficiently precise) in a reference coordinate system. For example, the arbitrary coordinate system could be the coordinate system of camera 11 at the moment an image in the sequence is captured. The scale is obtained, for example, by realigning a path obtained based on the camera's relative pose with a hypothetical path obtained using, for example, its approximate absolute pose provided by an inertial measurement unit. The arbitrary coordinate system can also vary from one image to another; for example, it could be the coordinate system of camera 11 at the moment a previous image in the sequence is captured.

[0084] For example, step S210, which determines one or more relative poses, can utilize a visual range method that allows, for example, determining pose changes from one image to another by analyzing image content.

[0085] Figure 4The schematic illustration shows the main steps of an example implementation of step S210 for determining the relative pose of camera 11. For example... Figure 4 As shown, step S210 includes:

[0086] - Step (S2100) to determine the relative pose of camera 11 for some or all images in an image sequence using visual range measurement.

[0087] - Step (S2101) to update the relative pose determined by bundle adjustment.

[0088] Therefore, in this example, the relative pose is determined in at least two stages.

[0089] First, during step S2100, a first estimate of the relative pose is obtained via visual measurement, for example, by tracing pixels (or pixel patterns) from one image to another, where these pixels are assumed to represent the same part of the scene. Several pixels in the same image are respectively matched with several pixels in one or more other images in the sequence. It should be noted that in some example embodiments, step S2100 of determining the relative pose of camera 11 may be implemented to determine the relative pose of camera 11 for each initial image captured by camera 11.

[0090] However, in some cases, the accuracy of relative pose determined by visual range may be limited due to possible error accumulation (the relative pose error estimated in one image may propagate from one image to another).

[0091] Following visual measurement, we obtain, for example, the estimated relative pose for each image in the sequence. The coordinates of pixels in an image represent the line of sight, i.e., a half-line in 3D space originating from the camera's optical center and passing through the pixel's center. Knowing the camera's pose, this half-line can be repositioned into a coordinate system representing the pose.

[0092] The update step S2101 using bundle adjustment aims to improve the accuracy of the relative pose determined using visual range in the previous step S2100.

[0093] Bundle adjustment is based on the principle that for two images captured by camera 11 at two different capture times, lines of sight from camera 11 that are associated with pixels representing the same part of the scene in both images must intersect at that same part. Pixels in different images representing the same part of the scene are called “matches”.

[0094] During the update step S2101 of the bundle adjustment, the relative pose is updated to ensure that lines of sight (pixels that have been matched) associated with pixels in different images substantially intersect at the same part (point) of the scene, and this applies to multiple parts of the scene represented in several images. In practice, it is generally impossible to make the lines of sight intersect strictly; "substantially intersecting" is understood as the lines of sight at least nearly intersecting.

[0095] For example, the Levenberg-Marquardt algorithm, using pixel-level epipolar error as a metric, can be used to update the relative poses of the first and last images in a sequence. This allows for the determination of approximate 3D positions of scene portions in arbitrary coordinate systems, identified as visible in both the first and last images of the sequence (approximate 3D positions corresponding to coordinates where the respective lines of sight substantially intersect). The relative poses of other images in the sequence can then be updated based on these approximate 3D positions using an n-point perspective (PNP) problem-solving algorithm, such as the SolvePnP function from the OpenCV library.

[0096] After bundle adjustment, we obtain, for example, the relative camera pose determined for each image in the sequence.

[0097] like Figure 3 As shown, in this example, the determination of the local 3D model further includes step S211 for matching pixels of the target image with pixels of other images in the sequence. As described above, pixels that match each other correspond to pixels in different images, which are considered to represent the same part (or point) of the scene, which is therefore, in principle, observed from different viewpoints. It should be noted that the matching performed during step S211 is advantageously a “dense” matching, that is, all pixels of the target image are matched with pixels of other images in the sequence. Obviously, for a given pixel of the target image, it is not always possible to identify a pixel representing the same part of the scene in another image (since the carrier 10 is movable, the part is not necessarily within the field of view of the camera 11 when other images are captured). However, preferably, a search for corresponding pixels in other images is performed for each pixel of the target image in order to identify a large number of scene parts visible in several images.

[0098] Specifically, if pixel matching is performed during step S210 to determine the relative pose, the number of pixels involved is far fewer than the number of pixels matched during the dense matching step S211. These arrangements enable improved accuracy and resolution of the local 3D model, and ultimately improved accuracy in determining the absolute pose of camera 11. Using a reduced set of points to measure pose allows for rapid estimation, compatible with the constraints of real-time systems, such as vision-based navigation systems. However, the geometric information subsequently allowed for matching with the 3D map cannot be extracted from this reduced set of points. Therefore, the geometric information subsequently allowed for matching with the 3D map is obtained through the next step S211, which performs dense matching of the target image with pixels from several consecutive images in the sequence. It should be noted that all these processing operations are applied incrementally to the image stream: it is not necessary to have all the images to begin processing. Each new image allows for additional measurements, but these are not necessary for the aforementioned measurements.

[0099] The matching step S211 can utilize any dense matching method known to those skilled in the art, and the choice of a particular method constitutes only a non-limiting variant implementation for determining the local 3D model. For example, the matching step S211 can utilize a dense optical flow algorithm. After matching the pixels of the target image with the pixels of other images in the sequence, we actually have the pixels of the target image, each pixel associated with one or more pixels of one or more other images.

[0100] Figure 5 This illustration schematically shows the main steps of an example implementation of the step of matching pixels of a target image with pixels of other images in the sequence. In this example, the relative pose determined during the step of determining the relative pose is used.

[0101] For example, to match pixels of a target image with pixels of another image in the sequence, a realignment transformation between the target image and the other image can be determined during step S2110 based on the relative poses of the two images. The realignment transformation aims to bring the target image and the other image into similar acquisition geometry, for example, by bringing the other image into acquisition geometry close to that of the target image (or vice versa). For example, the realignment transformation makes it possible to predict pixels in another image that theoretically represent the same part of the scene. In reality, the positions of pixels representing the same part of the scene may vary from image to image, especially due to changes in the pose of the camera 11 relative to the scene.

[0102] Figure 6The diagram schematically illustrates two consecutive images and the pixels in these images that represent the same part of a scene. More specifically, pixels p1 and p'1 represent the same part of the scene, pixels p2 and p'2 represent the same part, and pixels p3 and p'3 represent the same part. However, the position of pixel p1 in the left image is different from the position of pixel p'1 in the right image (and the same is true for pixels p2 and p'2, and pixels p3 and p'3). The realignment transform aims to attempt to predict the position of pixels representing the same part of the scene in the other image based on the position of pixels in the target image (or vice versa).

[0103] In some cases, the realignment transformation can be determined by making simplifying assumptions, particularly about the scene geometry. For example, in some cases, the realignment transformation corresponds to homography. As is known by itself, homography models motion by assuming that the parts of the scene represented by pixels lie in the same plane, i.e., assuming that the scene is generally a flat surface. This homography is easy to model and remains sufficient in many cases.

[0104] However, in other examples, it is possible to model the approximate geometry of the scene in a different way to determine the realignment transformation. It should be noted that this modeling of the scene's approximate geometry is intended only to establish a realignment transformation to facilitate pixel matching, and is therefore different from a local 3D model.

[0105] After step S2110, the realignment transformation between the two images is determined.

[0106] Then, during the next step S2111, a realignment transformation is used to align the target image with another image, that is, to make the target image and the other image as close as possible to the same acquisition geometry. The appearance of scene elements in the image varies depending on the viewpoint at which the image is captured. This realignment thus makes it possible to obtain a target image that is closer to the other image. For example, making the target image as close as possible to the acquisition geometry of the other image means that the target image is used to predict (by applying the realignment transformation to the target image) a desired image representing the scene as viewed from the viewpoint of the other image, which is therefore closer to the other image than the target image.

[0107] During the subsequent dense matching step S2112, the realigned images are paired to identify all pixels that can match each other. This dense matching can be performed using any method known to those skilled in the art. For example, it can utilize the correlation of pixel patterns between the target image and the realigned image, or a dense optical flow algorithm. After dense matching, residual motion between the pre-compensated image and the target image is obtained.

[0108] Generally, pre-aligning the relevant images using a realignment transform before matching improves the robustness and accuracy of the matching. Different types of realignment transforms can be used, which in particular can model the scene geometry in different ways. However, using homography has the advantage of being easy to determine and apply, while still achieving good results for dense pixel matching in many cases. Simplifying the pixel-to-pixel matching problem in this way allows for simpler (and therefore less costly in terms of verification, confirmation, and authentication), faster algorithms that better meet the real-time execution constraints of airborne systems.

[0109] like Figure 3 As shown, after step S211, which matches the pixels of the target image with the pixels of other images in the sequence when determining the local 3D model, step S212 is to determine the 3D position of each pixel of the target image in the coordinate system of camera 11 at the time of target image capture. For each part of the scene represented by the matched pixels in different images, the 3D position relative to camera 11 at the time of target image capture is determined, taking into account the relative pose and the positions of various matched pixels in the image. For example, the determination of the 3D position is achieved by triangulation of the 3D position (or at least the 3D position where the lines of sight substantially intersect) of the intersection points between the lines of sight corresponding to each mapped pixel in the scene.

[0110] At the end of step S212, which determines the 3D positions, the 3D positions of a large number of different parts of the scene relative to camera 11 at the moment of target image capture have been determined. These 3D positions of the scene parts thus describe the 3D geometry of the scene parts as observed by camera 11 in the camera 11 coordinate system at the moment of target image capture. These 3D positions of the scene parts thus form a local 3D model.

[0111] As described above, the matching performed during step S211 is dense matching, so the resulting local 3D model represents the 3D geometry of the scene portion at a 2D resolution on the order of the target image's resolution. For example, the local 3D model can be in the form of a 2D image, where pixels are associated with portions of the target image that have the same pixel size (resolution). However, the pixel values ​​of the 2D image of the local 3D model represent a third dimension, for example, in the form of the distance between camera 11 and the scene portion represented by that pixel, while the pixel values ​​of the target image represent a physical quantity of electromagnetic radiation from the scene portion represented by that pixel.

[0112] like Figure 2 As shown, after determining the local 3D model in step S21, the next step is step S22, which involves realigning the elements of the local 3D model in position and orientation to a predetermined reference 3D model of the scene to determine the absolute pose of the camera 11 at the moment of target image capture.

[0113] A reference 3D model represents the 3D geometry of a scene in a reference coordinate system. Therefore, it is the initial information about the scene, representing its 3D geometry, correctly oriented, positioned, and sized in the reference coordinate system. For example, a reference 3D model corresponds to a digital terrain model (DTM), or preferably to a digital elevation model (DEM) correctly oriented, positioned, and sized in the reference coordinate system (e.g., georeferenced in the case of a scene on the Earth's surface). As is known per se, a DTM represents the 3D geometry of the scene's ground without considering various elements (buildings, vegetation, etc.) located above the ground, while a DEM represents the 3D geometry of the scene considering elements located above the ground.

[0114] This reference 3D model can be established in advance using any 3D mapping method known to those skilled in the art, and the choice of a particular method constitutes only a non-limiting variant implementation of the method 20 for determining the absolute pose.

[0115] It should be noted that the reference 3D model can be determined, in particular, through the same steps as determining the local 3D model, based on a sequence of images captured by a camera whose position and orientation in the reference coordinate system are precisely known during the image sequence capture (e.g., determined using a GPS receiver). It should also be noted that, in some cases, the camera used to establish the scene reference 3D model can be camera 11 of vehicle 10. For example, in the case where vehicle 10 is making a round trip over the same scene, if vehicle 10 is equipped with, for example, a GPS receiver, the reference 3D model can be determined during the outbound journey of vehicle 10, and the method 20 for determining absolute pose can be used during the return journey of vehicle 10 for navigation purposes, such as to compensate for GPS receiver failure. Even without a GPS sensor, a non-metric map that allows for path tracing can still be generated.

[0116] The reference 3D model depends on the scenario, so computing device 12 is configured to retrieve, for example, reference 3D models associated with the scenario being flown over from a database, which may be mounted on the spacecraft or aircraft 10 or located remotely from the launch vehicle 10. For example, the database may store a single reference 3D model specifically created for a given mission to fly over an associated scenario. In other examples, the database may store several reference 3D models, each associated with a different scenario that can be flown over. In this case, computing device 12 selects and retrieves the reference 3D model to use based on information that allows identification of the scenario the launch vehicle 10 will fly over (e.g., based on the scenario or approximate coordinates of the launch vehicle 10). In the case where the launch vehicle 10 makes a round trip and creates a reference 3D model during the outbound journey, it is sufficient to retrieve the newly created reference 3D model from the database during the return journey.

[0117] The realignment of the local 3D model with the reference 3D model aims to align the local 3D model relative to the reference 3D model. This realignment is performed in both position and pose. It should be noted that the position realignment is done in three dimensions (3D position), meaning that the realignment aims not only to find the position (essentially a 2D position) within the reference 3D model that describes the 3D geometry of the local 3D model, but also to find the orientation and scale factors necessary to resize the local 3D model to locally match the reference 3D model. The scale factor is derived from the distance between the camera 11 and the scene, thus allowing the finding of the height component (third dimension) of the 3D position for determining the absolute pose. Realignment can be performed, for example, through correlation processing between the local 3D model and the reference 3D model, preferably block correlation, or according to any realignment method known to those skilled in the art, such as a "Iterative Closest Point" (ICP) type method by detecting and describing prominent elements.

[0118] Since the reference 3D model is built in the reference coordinate system, the absolute pose of the camera 11 in the reference coordinate system can be determined by realigning the local 3D model to the reference 3D model. That is, the position (3D) and pose of the camera 11 in the reference coordinate system can be determined at the moment of target image capture.

[0119] As mentioned above, realigning the position and orientation of the local 3D model with the reference 3D model aims to locate the local 3D model within the reference model.

[0120] In some cases, this realignment can be scanned within a search domain corresponding to a range of possible realignment values ​​for position (2D position and scale factor) and pose, in order to identify realignment values ​​that allow optimization of a predetermined similarity function representing the similarity between the local 3D model (recalibrated according to the considered realignment values) and the reference 3D model. For example, the similarity function corresponds to correlation processing between the local 3D model (recalibrated according to the considered realignment values) and the reference 3D model.

[0121] It should be noted that, in some cases, the scaling factor can be provided at least approximately by other means. For example, the scaling factor can be determined based on measurements provided by the inertial measurement unit of the vehicle 10 (e.g., predictions made by a navigation filter based on these measurements). If the scaling factor determined in this way is considered sufficiently accurate, further scaling is not required during step S22; for example, it is not necessary to scan the range of possible values ​​for the scaling factor. Otherwise, an approximate scaling factor can be used, for example, to reduce the search domain of the scaling factor in step S22.

[0122] Generally, any initial information about the absolute pose can be used to reduce the search domain (and lower the probability of false detection). In some cases, an approximate absolute pose of the camera 11 can be obtained to reduce the search domain, thereby accelerating and improving the realignment of the local 3D model with the reference 3D model. For example, the approximate absolute pose of the current target image can be determined based on the absolute pose previously determined for the target image capture moment and / or based on measurements provided by the inertial measurement unit of the carrier 10. In this case, step S22, which determines the absolute pose of the camera 11, aims to improve the accuracy of the determined absolute pose relative to the approximate absolute pose.

[0123] Figure 7 This schematically illustrates the main steps of an example implementation of step S22 for determining the absolute pose of camera 11 in a reference coordinate system, which includes:

[0124] - Step S220: Obtain the approximate absolute pose of camera 11 at the moment of target image capture.

[0125] - Step S221: Projecting the local 3D model of the scene portion at the moment of target image capture onto the reference coordinate system based on the approximate absolute pose.

[0126] - Step S222: Align the projected local 3D model with the reference 3D model in the reference coordinate system.

[0127] Projection step S221 essentially corresponds to changing the coordinate system (accompanied by a scaling change) from the camera 11's coordinate system to a reference coordinate system. This coordinate system change is based on an approximate absolute pose, which approximately describes the position (3D) and orientation of the camera 11's coordinate system in the reference coordinate system. For example, starting with a local 3D model corresponding to a 2D image describing the distances of various parts of the scene relative to the camera 11, we obtain a local 3D model, for example, corresponding to a local DEM in the reference coordinate system (local in that it represents the scene portion represented by the reference 3D model). Realignment step S222 is performed over a reduced search domain, limited to the neighborhood of the approximate absolute pose obtained in the previous step S220. As mentioned above, if the scaling factor of the approximate absolute pose is considered sufficiently accurate, no further scaling is required during step S222.

[0128] In certain implementations, a high-pass filter is used to filter both the local 3D model (possibly after projection onto a reference coordinate system) and the reference 3D model before realigning the position and pose to determine the absolute pose of camera 11. This arrangement facilitates realignment of the position and pose between the local 3D model and the reference 3D model. For example, in cases where the local 3D model and the reference 3D model correspond to a DEM, applying a high-pass filter highlights, for example, high-frequency changes in elevation corresponding to the presence of buildings in the scene. Generally, buildings or more general structures in the scene that introduce rapid elevation changes locally, along with their respective locations, are important in determining the absolute pose of camera 11 and are more relevant and readily available than absolute elevation values. Conversely, slow changes in elevation within the scene (e.g., the slope of the ground where buildings are located) carry little information. Slow changes in elevation within the scene often contain most of the differences between the two maps due to initial alignment errors, as the algorithm first attempts to realign the average slope, for example, by using translation, even if that average slope originates from initial pose errors or pose drift used for 3D reconstruction. Slow changes in elevation within a scene can also introduce more noise in processing related to the local 3D model and the reference 3D model. For example, the resolution of a rapid elevation change to be preserved is on the order of pixels in the scene's local 3D model (DEM). Therefore, the parameters of the high-pass filter are advantageously chosen to preserve such rapid elevation changes. Completely eliminating low frequencies is also unnecessary, as this allows for alignment with hills or other large natural structures.

[0129] When applying a high-pass filter, it's necessary to consider both what needs to be preserved, such as buildings, and what is being eliminated by the filter. In effect, a high-pass filter allows for the removal of the most incorrectly reconstructed 3D information that particularly interferes with realignment.

[0130] For example, Figure 8 This schematically illustrates a local 3D model (projected onto a reference coordinate system using the approximate absolute pose of camera 11) before high-pass filtering. Figure 8 (a) and the processed ( Figure 8 Examples of (b)). In Figure 8 In the example shown, the local 3D model corresponds to the DEM. For example... Figure 8 As shown in section (a), the elevation changes substantially with the scene slope, and the rapid changes in scene elevation are sometimes negligible compared to the absolute elevation. Figure 8 In section (b), high-pass filtering is used to highlight rapid elevation changes, and it can be seen that they provide more information for determining the camera's absolute pose. It is indeed understandable that rapid changes in elevation within a scene better characterize a scene (e.g., compared to another scene) than slow changes in elevation within the same scene.

[0131] Another advantage of high-pass filtering is the removal of edge effects. In fact, as... Figure 8 As shown in section (a), the local 3D model is only defined locally (the left, right, and top regions of section (a) are undefined), which introduces elevation discontinuities at the edges when realigning the position and pose with the reference 3D model. These discontinuities can cause artifacts during correlation, which are removed by high-pass filtering, such as... Figure 8 As shown in part (b).

[0132] This method advantageously utilizes dense 3D representations. In practice, for visual navigation, identifiable elements are typically small when viewed from the sky (buildings, streets, trees, roads), so our dense method allows for the reconstruction of relevant salient elements for matching, namely building edges / walls, subtle elevation differences between the center and edge of roads, and trees.

[0133] More generally, it should be noted that the implementations and embodiments considered above have been described as non-limiting examples, and therefore other variations are possible.

Claims

1. A real-time method (20) for determining the absolute pose of a camera (11) in a reference coordinate system, the camera (11) being monocular and passive, located on an aircraft or spacecraft (10) capable of moving relative to a known scene mapped in the form of a reference 3D model, the method comprising: - Acquire (S20) a sequence of at least two consecutive images captured by the camera (11) at corresponding consecutive time points, each image comprising a plurality of pixels, each pixel of the image partially representing the scene observed by the camera at the time the image was captured. - Determine (S21) a local 3D model from the image sequence in a coordinate system centered on the camera (11) and oriented relative to the camera, the local 3D model representing a three-dimensional portion of the scene corresponding to the capture of an image in the image sequence called the target image, the determination of the local 3D model (S21) including dense matching (S211) of all pixels of the target image with pixels of other images in the sequence. - Provides the approximate absolute pose of the camera described in (S220), - By realigning the position and pose of the local 3D model with the reference 3D model, position and pose realignment is performed in the search domain based on the approximate absolute pose of the camera to determine the absolute pose of the camera (11) at the moment of target image capture.

2. The method (20) according to claim 1, characterized in that, Determining the local 3D model involves determining at least one relative pose of the camera for several consecutive images in a sequence.

3. The method (20) according to claim 2, characterized in that, Determining the relative pose of the camera for several consecutive images in the sequence includes visual range measurement (S2100), followed by updating the relative pose through bundle adjustment (S2101).

4. The method (20) according to claim 2 or 3, characterized in that, The determination of the local 3D model (S21) includes: - Determine the relative pose of the camera for several consecutive images in the sequence (S210), then - Perform dense matching of all pixels in the target image with pixels in several consecutive images in the sequence (S211). - For each pixel that is actually matched in the target image: Based on the relative pose, the 2D position of the pixel in the target image, and the 2D position of each matched pixel, determine (S212) the 3D position of the pixel in the target image in the camera coordinate system. The local 3D model is formed based on the 3D positions of the actual matched pixels in the target image.

5. The method (20) according to claim 4, characterized in that, The dense matching of pixels in the target image with pixels in other consecutive images in the sequence (S211) includes, for each other consecutive image in the sequence: -Based on the corresponding relative pose, determine (S2110) the realignment transformation between the target image and the other images. - The other images are realigned using a realignment transformation (S2111). - Determine the residual motion from the target image to other realigned images (S2112).

6. The method (20) according to claim 5, characterized in that, Realign and transform to homography.

7. The method (20) according to claim 5 or 6, characterized in that, The residual motion of the target image and other realigned images is determined using a dense optical flow algorithm.

8. The method (20) according to any one of the preceding claims, characterized in that, Determining the absolute pose of the camera (11) at the target image capture time (S22) includes: - Project the local 3D model (S221) onto the reference coordinate system based on the approximate absolute pose. - Match the projected local 3D model with the reference 3D model in the reference coordinate system (S222).

9. The method (20) according to any one of the preceding claims, characterized in that, The approximate absolute pose is determined based on the absolute pose of the camera (11) determined during previous target image capture or based on navigation instruments carried on an aircraft or spacecraft.

10. The method (20) according to any one of the preceding claims, characterized in that, Before determining the absolute pose of the camera (11), a high-pass filter is used to filter the local 3D model and the reference 3D model.

11. The method (20) according to any one of the preceding claims, characterized in that, The images in the acquired image sequence correspond to images called key images selected from a sliding sequence of images called initial images, which are captured continuously by the camera.

12. The method (20) according to claim 11, characterized in that, An initial image is selected as a key image when it meets predetermined criteria for the motion of the aircraft or spacecraft (10) since the previous key image was captured.

13. A real-time method for vision-based navigation, using a monocular and passive camera (11) connected to an aircraft or spacecraft (10) platform capable of moving relative to a known scene mapped in the form of a reference 3D model, the method comprising: - According to any one of claims 1 to 12, the absolute pose of the camera (11) in the reference coordinate system is determined, and then... - Determine the position and attitude of the aircraft or spacecraft in the reference 3D model based on the absolute pose corrected according to the position of the camera relative to the aircraft or spacecraft platform.

14. A computer program product comprising instructions that, when executed by at least one processor, configure the at least one processor to perform the method according to any one of claims 1 to 13.

15. A computing device (12) comprising at least one processor and at least one memory, said at least one processor being configured to implement the method according to any one of claims 1 to 13.

16. An aircraft or spacecraft (10) comprising a platform carrying a camera (11) and a computing device (12) according to claim 15.

Citation Information

Patent Citations

  • Monocular stereo vision relative position / pose measuring method

    CN103528571A

  • Aircraft attitude determination method using event camera as star sensor

    CN114663492A

  • Method of calibrating a computer-based vision system onboard a craft

    US20140267608A1