Systems and methods for visual navigation

The visual navigation system processes video frames to determine aircraft geolocation using image-based transformations and Kalman filters, addressing GPS reliance and enabling accurate navigation in GPS-denied areas.

JP2026516079APending Publication Date: 2026-05-19PALANTIR TECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PALANTIR TECHNOLOGIES INC
Filing Date
2024-05-06
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Conventional navigation systems rely heavily on GPS and fail to function in GPS-denied environments, limiting their effectiveness in areas where satellite signals are unavailable.

Method used

A visual navigation system utilizing image-based transformations, georegistration, and nonlinear Kalman filters to determine aircraft geolocation by processing video frames from onboard sensors, incorporating metadata and pixel motion analysis to track and correct positional errors.

Benefits of technology

Enables accurate aircraft navigation and geolocation in GPS-denied environments by integrating visual data and motion data, providing precise location information for remote control and improving navigation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026516079000001_ABST
    Figure 2026516079000001_ABST
Patent Text Reader

Abstract

Systems and methods for visual navigation are provided. An exemplary method includes receiving a plurality of video frames from an image sensor positioned on an aircraft and generating an image-based transformation based on the plurality of video frames. In some examples, the image-based transformation is associated with the motion of one or more image features and the motion of the image sensor. In some examples, the method further includes determining an image-based movement associated with the aircraft based on the image-based transformation, generating a georegistration transformation based on at least one video frame from the plurality of video frames and a reference image, determining a georegistration-based geolocation associated with the aircraft based on the georegistration transformation, and determining the aircraft geolocation by applying a nonlinear Kalman filter to the image-based movement and the georegistration-based geolocation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 465,392, filed on May 10, 2023, entitled "SYSTEMS AND METHODS FOR VISUAL NAVIGATION", the entire disclosure of which is incorporated herein by reference for all purposes.

[0002] Technical Field

[0002] Certain embodiments of the present disclosure relate to navigation. More specifically, some embodiments of the present disclosure relate to visual navigation.

Background Art

[0003] Background

[0003] Navigation includes various types, including visual navigation and device navigation. In some examples, device navigation may include navigation performed with the assistance of a Global Positioning System (GPS). Exemplary contexts of navigation include navigating an aircraft.

[0004]

[0004] Therefore, it is desirable to improve the technology for visual navigation.

Summary of the Invention

Means for Solving the Problems

[0005] Summary

[0005] Certain embodiments of the present disclosure relate to navigation. More specifically, some embodiments of the present disclosure relate to visual navigation.

[0006]

[0006] At least some aspects of the present disclosure relate to methods for visual navigation. In some examples, the method includes receiving a plurality of video frames from an image sensor positioned on an aircraft and generating an image-based transformation based on the plurality of video frames. In some examples, the image-based transformation is associated with the motion of one or more image features and the motion of the image sensor. In some examples, the method further includes determining an image-based movement associated with the aircraft based on the image-based transformation, generating a georegistration transformation based on at least one video frame from the plurality of video frames and a reference image, determining a georegistration-based geolocation associated with the aircraft based on the georegistration transformation, and determining the aircraft geolocation by applying a nonlinear Kalman filter to the image-based movement and the georegistration-based geolocation. In some examples, the method is performed using one or more processors.

[0007]

[0007] At least some aspects of the present disclosure relate to systems for visual navigation. In some examples, the system includes at least one processor and at least one memory for storing instructions, which, when an instruction is executed by the at least one processor, causes the system to perform a set of operations. In some examples, the set of operations includes receiving a plurality of video frames from an image sensor positioned on an aircraft and generating an image-based transformation based on the plurality of video frames. In some examples, the image-based transformation is associated with the motion of one or more image features and the motion of the image sensor. In some examples, the set of operations further includes determining an image-based movement associated with the aircraft based on the image-based transformation, generating a georegistration transformation based on at least one video frame from the plurality of video frames and a reference image, determining a georegistration-based geolocation associated with the aircraft based on the georegistration transformation, and determining the aircraft geolocation by applying a nonlinear Kalman filter to the image-based movement and the georegistration-based geolocation.

[0008]

[0008] At least some aspects of the present disclosure relate to methods for visual navigation. In some examples, the method includes receiving a plurality of video frames from an image sensor positioned on an aircraft and generating an image-based transformation based on the plurality of video frames. In some examples, the image-based transformation is associated with the motion of one or more image features and the motion of the image sensor. In some examples, the method further includes determining an image-based movement associated with the aircraft based on the image-based transformation, generating a georegistration transformation based on at least one video frame from the plurality of video frames and a reference image, determining a georegistration-based geolocation associated with the aircraft based on the georegistration transformation, receiving metadata associated with the aircraft, and estimating the aircraft geolocation by applying a nonlinear Kalman filter to the image-based movement, georegistration-based geolocation, and metadata. In some examples, the method is performed using one or more processors.

[0009]

[0009] Depending on the embodiment, one or more benefits may be achieved. These benefits of the present disclosure, as well as various additional objectives, features and advantages, can be fully understood by referring to the following detailed description and accompanying drawings. [Brief explanation of the drawing]

[0010] Brief explanation of the drawing [Figure 1]

[0010] This is an illustrative diagram of a visual navigation environment or workflow according to a particular embodiment of the present disclosure. [Figure 2]

[0011] This is a simplified diagram illustrating a method for visual navigation according to a particular embodiment of the present disclosure. [Figure 3]

[0012] This is a simplified diagram illustrating a method for visual navigation according to a particular embodiment of the present disclosure. [Figure 4]

[0013] This is a simplified diagram illustrating a software architecture for a video georegistration system according to a specific embodiment of the present disclosure. [Figure 5]

[0014] This is a simplified diagram illustrating a method for video georegistration according to a specific embodiment of the present disclosure. [Figure 6]

[0015] This is a simplified diagram illustrating a method for generating an image transformation according to a specific embodiment of the present disclosure. [Figure 7]

[0016] A simplified diagram illustrating a computing system according to a specific embodiment of this disclosure is shown. [Modes for carrying out the invention]

[0011] Detailed explanation

[0017] Unless otherwise indicated, all figures used in this specification and the claims to represent the size, quantity, and physical properties of features should be understood in all cases to be modified by the term "approximately." Therefore, unless otherwise indicated, the numerical parameters described in the foregoing specification and the appended claims are approximations that may vary depending on the desired properties to be obtained by a person skilled in the art using the teachings disclosed herein. The use of numerical ranges by endpoints includes all numbers within that range (for example, 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.80, 4, and 5), as well as any range within that range.

[0012]

[0018] Exemplary methods may be represented by one or more drawings (e.g., flow charts, communication flows, etc.), but the drawings should not be construed as implying any requirement of any of the various steps disclosed herein, or any particular order between them. However, some embodiments may require certain steps and / or a particular order between certain steps, as can be expressly described herein and / or understood from the nature of the steps themselves (e.g., the execution of some steps may depend on the results of previous steps). In addition, a “set,” “subset,” or “group” of items (e.g., inputs, algorithms, data values, etc.) may include one or more items, and similarly, a subset or subgroup of items may include one or more items. “Multiple” means two or more.

[0013]

[0019] As used herein, the term “based on” is not limited to, but rather indicates that a decision, identification, prediction, calculation, etc., is performed by using at least the term preceding “based on” as input. For example, predicting an outcome based on certain information may, additionally or alternatively, be making the same decision based on other information. As used herein, the term “receive” or “receiving” means to obtain from a data repository (e.g., a database), from another system or service, from another software, or from another software component within the same software. In certain embodiments, the term “access” or “accessing” means to retrieve data or information and / or generate data or information.

[0014]

[0020] Conventional navigation systems and methods are often unable to navigate when GPS (Global Positioning System) is unavailable. Conventional systems and methods typically use GPS information for navigation, and as a result, they cannot navigate in areas where GPS is unavailable (also known as GPS-denied environments).

[0015]

[0021] Various embodiments of this disclosure can achieve benefits and / or improvements by using, for example, visual data (e.g., video) and / or motion data for navigation through a computing system. In some embodiments, benefits include significant improvements, such as performing aircraft navigation and generating aircraft location information even when the aircraft is located in an area where GPS is unavailable. In some embodiments, location information may be provided to one or more displays to enable a user to control the aircraft's navigation via remote control, for example. In some embodiments, location information is provided to one or more control systems, which can perform certain technical measures based on the location information, such as, but not limited to, remotely controlled navigation of the aircraft. In certain embodiments, other benefits include improved accuracy of navigation, for example, by using visual data and / or motion data. In some embodiments, benefits further include the ability to process visual data from two or more sensor types and use the visual data for navigation. In certain embodiments, systems and methods are configured to use visual data, motion data, and / or georegistration for navigation.

[0016]

[0022] In certain embodiments, a system (e.g., a navigation system) may utilize one or more unmanned aerial vehicles (UAs) (e.g., drones) having camera feeds to monitor and analyze an area of ​​interest. In some embodiments, as used herein, aircraft refers to aircraft, unmanned aerial vehicles, drones, etc. In some use cases, it is important (e.g., extremely important) for the user to know the geolocation and camera image of the drone with high accuracy and precision. In some conventional systems, the problem of geographically positioning the drone can be solved by performing GPS (Global Positioning System) measurements, but aircraft often fly in areas where GPS measurements are unavailable. To solve this problem, in certain embodiments, the system and method may use a visual navigation solution (e.g., visual navigation software, visual navigation software module, visual navigation module, visual navigation system), which includes, for example, the integration of software and algorithmic techniques for tracking the movement of the aircraft.

[0017]

[0023] According to some embodiments, the system and method may include visual navigation using one or more Sensor Inference Platform (SIP) processors. In certain embodiments, the SIP performs an adjustment between input sensor data and an output feed. As an example, the SIP is a model orchestrator for one or more models and / or one or more sensors (e.g., sensor feeds), and is also referred to as a model and / or sensor orchestrator. In some embodiments, a model, referred to as a computing model or algorithm, includes a model for processing data. In some embodiments, the model includes, for example, an AI model, a machine learning (ML) model, a deep learning (DL) model, an image processing model, a physical model, simple heuristics, rules, a mathematical model, other computing models, and / or combinations thereof. For example, one or more components of the SIP utilize open standard formats (e.g., input data format, output data format). As an example, the SIP processes the decoding of input data, the orchestration between the processor and an artificial intelligence (AI) model, and packages the result in an open output format for downstream consumers. According to some embodiments, the system includes one or more SIPs for coordinating one or more sensors, one or more edge devices, one or more user devices, and / or one or more models. In certain embodiments, at least some of the one or more sensors, one or more edge devices, one or more user devices, and one or more models are each associated with a SIP.

[0018]

[0024] According to certain embodiments, a visual navigation module integrated with the SIP enables the visual navigation module to be deployed in a variety of settings, including, for example, on the edge, and operate without depending on the format of the incoming video stream. In some embodiments, such an implementation may distinguish this solution from other solutions.

[0019]

[0025] According to some embodiments, a user can stream an aircraft (e.g., a drone) video feed via SIP, which associates metadata related to the aircraft (e.g., speed, velocity, orientation, direction, altitude, etc.) with each video frame and passes the information along a visual navigation module. In certain embodiments, the metadata enables tracking of the aircraft (e.g., initial tracking) along with an initial estimate of the UA position, although in some cases such a solution may be insufficient to obtain long-term accuracy.

[0020]

[0026] According to certain embodiments, through the process of feature extraction and feature matching, a visual navigation system can use pixel motion in frame-to-frame analysis to determine the movement of an image sensor (e.g., a camera, video camera). In some embodiments, the visual navigation system can apply a georegistration algorithm to a video frame to find its position and then use that to determine the position of the aircraft.

[0021]

[0027] According to some embodiments, a georegistration implementation uses image matching techniques that function within and across two or more sensor types, including, for example, electro-optical (EO) sensor types, infrared (IR) sensor types, synthetic aperture radar (SAR) sensor types, etc. In certain embodiments, such a georegistration implementation has one or more advantages over other techniques that function with only a single type of data. In some embodiments, visual navigation according to certain embodiments also functions well in natural terrain with omitted details, which is a common problem area in image matching. In a specific example, this is an important component of this technology.

[0022]

[0028] In certain embodiments, the two inputs (e.g., metadata and pixel motion) do not provide actual geolocation information, resulting in an accumulation of errors. In some embodiments, incorporating georegistration results allows this solution to eliminate this accumulation of errors.

[0023]

[0029] According to some embodiments, three pieces of information are fed into an unscented Kalman filter (UKF), which, once an initial estimate of the aircraft's geolocation is fed into the UKF, can continuously track the aircraft's position throughout its flight time. In some examples, at least two pieces of information are fed into the UKF, for example, to continuously track the aircraft's position throughout its flight time. In certain embodiments, the implementation of this solution involves the integration of SIP and the design of the UKF. In some embodiments, the filter design involves determining the relevant information to track and / or model how the aircraft moves and how its sensors behave.

[0024]

[0030] According to certain embodiments, the system involves the integration of these various components, which include two parts. Firstly, in some embodiments, the visual navigation system may take in one or more video streams (e.g., arbitrary video streams). In certain embodiments, the system may combine metadata (e.g., speed, velocity, heading, orientation, altitude of the aircraft system) with images on a frame-by-frame basis. In some embodiments, a video frame, also called an image frame or frame, is an image in a sequence of images or an image in a video. Secondly, in certain embodiments, the system may include an image matching georegistration solution (e.g., an advanced image matching georegistration solution) for matching with a wide range of data as input. In some embodiments, a solution (e.g., a system) that utilizes georegistration can obtain consistently accurate geolocation measurements during flight time, even while operating in GPS denial (e.g., no GPS information) settings. Some conventional navigation systems may not have access to or be able to integrate with georegistration solutions. Some conventional navigation systems may not have access to or be able to integrate with image matching georegistration solutions. In some embodiments, this solution can be integrated with one or more sensors present in existing workflows and / or aircraft.

[0025]

[0031] According to some embodiments, the visual navigation system includes and / or is integrated with a georegistration system (e.g., georegistration for video, georegistration for full-motion video). In certain embodiments, the georegistration system (e.g., georegistration service) is configured to receive (e.g., acquire) video from a video recording platform (e.g., imaging sensor on a VA) having location information (e.g., telemetry data), and to align the received video with a reference frame (e.g., dynamic reference frame, reference image, map), thereby enabling the georegistration system to determine the geospatial coordinates (also called geographic coordinates) of the video.

[0026]

[0032] According to certain embodiments, the georegistration system incorporates one or more techniques, including, for example, geometric correction, orthorectification, etc. In some embodiments, geometric correction refers to assigning geographic coordinates to an image. In certain embodiments, orthorectification refers to distorting an image to match a top-down view. In some examples, orthorectification includes reshaping hill slopes so that the image appears to have been taken directly from above rather than from the side. In some embodiments, georegistration refers to improving the geographic coordinates of a video based on, for example, reference data.

[0027]

[0033] In certain embodiments, image registration (e.g., georegistration) refers to finding a transformation that maps an input image to corresponding portions of one or more reference images, given an input image and one or more reference images. In some embodiments, video registration (e.g., georegistration) refers to finding a transformation or sequence of transformations that maps an input video, containing one or more video frames, to corresponding portions of one or more reference images, given an input video and one or more reference images, and using this transformation to generate a aligned video. In certain embodiments, image / video georegistration may have one or more challenges, namely: 1) the image / video may have visual variations, e.g., changes in lighting, temporal variations (e.g., seasonal variations), sensor modes (e.g., electro-optical (EO), infrared (IR), synthetic aperture radar (SAR), etc.); 2) the image / video may have minimal structured content (e.g., forests, fields, water, etc.); 3) the image / video may have noise (e.g., image noise in SAR images); and 4) the image / video may have rotation, scaling, and / or changes in viewpoint.

[0028]

[0034] According to some embodiments, a georegistration system is configured to receive video (e.g., streaming video) and select one or more video frames (e.g., video images) and one or more selected derivatives within the video frames (e.g., a composite derived from multiple video frames, a pixel grid of video frames, etc.), also called a template (e.g., 60x60 pixels). In certain embodiments, video georegistration uses selected video frames (e.g., every second) and a template that can shorten the time. In some embodiments, the georegistration system performs georegistration of the template, collects desired matches, calculates image transformations, and generates a sequence of aligned video frames (e.g., video frames aligned to a geographic coordinate system) and an aligned video (e.g., video aligned to a geographic coordinate system).

[0029]

[0035] In certain embodiments, the georegistration system computes an image representation (e.g., one or more feature descriptors) of a template for georegistration. In some embodiments, the georegistration system computes an angle-weighted directional gradient (AWOG) representation of the template for georegistration. In certain embodiments, the georegistration system compares the AWOG representation of the template to a reference image (e.g., a reference image) to determine a match and / or match score, e.g., if the template matches the reference image well (e.g., 100%, 80%). In some embodiments, the georegistration system iterates through the process to find a well-matched template. In certain embodiments, the georegistration system uses the matched template to perform georegistration of an image or video frame. In some embodiments, the matched template may be noisy and / or irregular.

[0030]

[0036] According to some embodiments, video georegistration is achieved by a collection of individual components of a georegistration system. In certain embodiments, these are individual SIP (sensor inference platform) (e.g., software orchestrator, model, and / or sensor orchestrator) processors (e.g., computing units implementing SIP) operating in parallel with and / or behind an aggregate filter processor (e.g., a computing unit implementing SIP). In some embodiments, a processor refers to a computing unit implementing a model (e.g., a computing model, algorithm, AI model, etc.). In certain embodiments, a model, also called a computing model, includes a model for processing data. Models include, for example, AI models, machine learning (ML) models, deep learning (DL) models, image processing models, algorithms, rules, other computing models, and / or combinations thereof.

[0031]

[0037] Figure 1 is an illustrative diagram of a visual navigation environment or workflow 100 according to a particular embodiment of the present application. Figure 1 is merely an example. Those skilled in the art will recognize many variations, alternatives, and modifications. For example, some of the components may be extended, integrated, and / or combined. Other components may be inserted into those described above. Depending on the embodiment, the arrangement of components may be replaced with others. Further details of these components are found throughout this disclosure.

[0032]

[0038] In certain embodiments, the visual navigation environment or workflow 100 includes one or more aircraft 105, a visual navigation system 102, and one or more output systems 130. In some embodiments, the aircraft 105 may be an aircraft, an unmanned aerial vehicle, a drone, etc. In some embodiments, the visual navigation system 102 includes SIP 120A (e.g., a software orchestrator, a model and / or a sensor orchestrator), a visual navigation module 110 (e.g., visual navigation software, a visual navigation system, etc.), and SIP 120B. In certain embodiments, the visual navigation module 110 includes an optical flow processor 112, metadata 114, a georegistration processor 116, and / or an estimation processor 118. In some embodiments, different components within the visual navigation module 110 are expected to run at different FPS (frames per second) values ​​based on their respective computational demands. In certain embodiments, SIP 120A receives one or more videos 122, for example, via a video stream. In some embodiments, one or more videos 122 include a sequence of images (e.g., frames).

[0033]

[0039] According to some embodiments, the visual navigation system 102 and / or the visual navigation module 110 include several technical components that work together to provide accurate PNT (positioning, navigation, and timing) in GPS-denied environments. In certain embodiments, the visual navigation module 110 takes an input video stream or a series of image frames having metadata 114 (e.g., telemetry data). In some embodiments, the visual navigation module 110 outputs geolocation information 126 (e.g., PNT information, location information and timing information, three-dimensional (3D) location information, etc.). In certain embodiments, the PNT information 126 is output, for example, via SIP 120B, to one or more output systems structured for integration into a positioning system (e.g., an APNT (guaranteed positioning navigation and timing) system, including, for example, tracking and telemetry, active control, or a PNT open architecture (e.g., pntOS).

[0034]

[0040] According to certain embodiments, SIP 120A and / or 120B enable the delivery of data through a user- or system-configurable data processing and analysis pipeline. In some embodiments, SIP 120A is configured to ingest various videos 122 (e.g., video streams, arbitrary video streams) and associate metadata (e.g., location metadata, aircraft metadata) with each frame of the video. In certain embodiments, this enables rapid cross-platform deployment rather than being tied to a specific sensor / platform integration.

[0035]

[0041] According to some embodiments, an optical flow processor 112 (e.g., a processing unit implementing an optical flow software module) can measure pixel motion between two video frames (e.g., image frames, frames) through a feature extraction and feature matching process. In some embodiments, the optical flow processor 112 can determine image-based transformations and image-based motions (e.g., pixel motion, image feature motion, etc.). In certain embodiments, a first frame is analyzed for relatively distinctly different areas within the image. In some embodiments, a second frame is analyzed to find features that are the same as or similar to those extracted from the first frame. In certain embodiments, each successful match generates a motion vector (e.g., how that small area moved from the first frame to the second frame). In some embodiments, one or more vectors are combined to determine an image-based transformation (e.g., inter-frame transformation) at least partially based on the pixel motion (e.g., transformation from the first frame to the second frame). In certain embodiments, the image-based transformation includes the motion vector. In certain embodiments, this transformation can provide a transformation from pixel space to camera space, and thus give a measurement of camera movement (e.g., movement of an image sensor placed on an aircraft). In some embodiments, the transformation from pixel space to camera space is computed using metadata associated with the aircraft.

[0036]

[0042] In certain embodiments, metadata 114 includes, for example, speed, heading, altitude, and orientation (e.g., of and / or associated with aircraft 105). In certain embodiments, metadata 114 includes motion vectors. In some embodiments, the visual navigation system 102 includes measurements of metadata 114, including, for example, the speed, heading, altitude, and orientation of aircraft 105. In certain embodiments, even if consistent GPS data is not available, the visual navigation system 102 includes measurements of metadata 114, including, for example, the speed, heading, altitude, and orientation of aircraft 105. In some embodiments, the visual navigation system 102 can track the aircraft's position. In some embodiments, lossy / unreliable GPS information and externally supplied initial settings (e.g., initial position) can be integrated into metadata 114 to improve the accuracy and reliability of the solution.

[0037]

[0043] According to some embodiments, the georegistration processor 116 is configured to reduce location-based errors in the initial correction by performing a process of aligning the geometrically corrected image to a reference image. In certain examples, the georegistration processor 116 uses a geometric correction algorithm to generate the geometrically corrected image. In certain embodiments, the georegistration processor 116 uses an algorithm that enables registration across multiple image modalities (EO, IR, SAR) and registration in a wide range of environments that outperform conventional registration algorithms. In some embodiments, since georegistration can determine the position of an image (e.g., image frame, video frame) with high precision and accuracy, and assuming that the transformation between the image and the camera itself is determinable, the visual navigation module 110 can determine the position of the aircraft by determining the georegistration transformation. In some embodiments, the georegistration transformation is the transformation between the geographical location of a first image and the geographical location of a second image.

[0038]

[0044] In certain embodiments, the visual navigation module 110 and / or the georegistration processor 116 (e.g., a georegistration system) can use the beam angle and / or beam angle metadata of an imaging sensor located on the aircraft 105 to determine a georegistration transformation, also known as geolocation correction or geolocation transformation. In some embodiments, the georegistration processor 116 uses a reference image 117 in the georegistration process. In certain embodiments, the reference image refers to a set of pre-aligned images for one or more georegistrations to align a new image (e.g., a new image, an input image, a new image frame, an input image frame, etc.). In some embodiments, the visual navigation module 110 and / or the georegistration processor 116 can retrieve the reference image 117 from components, external components, third-party datasets (e.g., a custom online map dataset for a website or application, a custom dataset), a third-party system, etc.

[0039]

[0045] According to some embodiments, the visual navigation module 110 includes an estimation processor 118 that can receive one or more inputs, including one or more image-based movements (e.g., pixel motion) from an optical flow processor 112, metadata 114 (e.g., motion metadata), and / or georegistration-based geolocation (e.g., geolocation correction) from a georegistration processor 116. In certain embodiments, the estimation processor 118 can implement one or more estimation techniques and / or machine learning applications, including, for example, a nonlinear Kalman filter, an unscented Kalman filter, or an extended Kalman filter.

[0040]

[0046] In certain embodiments, the estimation processor 118 combines one or more inputs to generate an aircraft position estimate (e.g., a high-accuracy estimate). In some embodiments, the estimation processor 118 uses an unscented Kalman filter to receive one or more inputs (e.g., image-based motion, motion metadata, georegistration-based geolocation) to generate an aircraft position estimate. In certain embodiments, the Kalman filter includes a Bayesian state estimation technique that combines prediction of how a given process will behave with measurement of its current state. In some embodiments, the use of prediction and measurement improves the accuracy of the final state estimate compared to, for example, other techniques. In certain embodiments, the unscented Kalman filter is an extension of the Kalman filter configured to handle nonlinearity. In some embodiments, the estimation processor 118 can incorporate a behavior model (e.g., an aircraft behavior model) that predicts how the aircraft will behave over time.

[0041]

[0047] According to some embodiments, the visual navigation system 102 may provide one or more outputs 126 via SIP 120B, for example, aircraft position. In certain embodiments, one or more outputs 126 may be integrated with or input to one or more output systems 130 (e.g., external systems), including, for example, one or more user systems 132, one or more control systems 134, and one or more position systems 136 (e.g., positioning, navigation, and timing-based operating systems such as pntOS from the PNT Open Architecture).

[0042]

[0048] According to certain embodiments, the georegistration processor 116 includes a calibration module (e.g., a calibration processor) for performing calibration on video frames at a selected FPS. In some embodiments, the calibration module can perform calibration on video frames at full FPS, for example, each video frame being calibrated. In certain embodiments, the calibration module uses historical telemetry data (e.g., historical telemetry) and / or optional corrections (e.g., built-in corrections). In some embodiments, the calibration module requires a low computational cost (e.g., a few milliseconds, 2-3 milliseconds).

[0043]

[0049] According to some embodiments, the optical flow processor 112 processes video frames at a low FPS (e.g., 5 FPS, adaptive FPS). In certain embodiments, the visual navigation system 102 and / or the optical flow processor 112 compute an optical flow-based motion model to provide an alternative smoothed estimate of the movement of one or more objects in the video. In some embodiments, depending on the performance profile, the visual navigation system 102 moves the computation kernel of the optical flow processor 112 to a specific library (e.g., a C++ library). In certain embodiments, a DEM (digital elevation model) or similar data model (e.g., a digital terrain model) can be used to translate visual movement into estimated physical movement. In some embodiments, the optical flow processor requires a moderate computation cost (e.g., tens of milliseconds or more). In certain embodiments, the optical flow processor extracts the relative movement of objects from the video frame to perform corrections.

[0044]

[0050] In certain embodiments, the georegistration processor 116 periodically references georegistration (e.g., less than 1 FPS). In some embodiments, the reference georegistration processor 116 aligns selected video frames or several derived composites with respect to a reference image. In certain embodiments, the visual navigation system 102 and / or the georegistration processor 116 can acquire more data using the video frames themselves or composites of multiple frames. In some embodiments, the visual navigation system 102 and / or the georegistration processor 116 can compare a reference image from above or pre-project a reference or input image based on a predicted line-of-sight angle. In certain embodiments, the visual navigation system 102 and / or the georegistration processor 116 may use various algorithms (e.g., algorithm classes) including algorithms to support multimodal (e.g., EO (electro-optical) and IR (infrared)) results. In some embodiments, depending on the performance profile, the visual navigation system 102 moves the computation kernel for the georegistration processor 116 to a specific library (e.g., a C++ library).

[0045]

[0051] In some embodiments, the georegistration processor 116 requires a weight calculation cost (e.g., less than 1 second). In certain embodiments, the calibration module processes video frames at a first frame rate (e.g., once per frame, once every other frame, once every J frames). In some embodiments, the frame rate refers to the frequency of video frames being used (e.g., frames per second (FPS)) and / or the frequency with which video frames are used in a sequence of video frames. In some embodiments, the optical flow processor 112 processes video frames at a second frame rate (e.g., once every M frames). In certain embodiments, the georegistration processor 116 processes video frames at a third frame rate (e.g., once every N frames). In some embodiments, the first frame rate is higher than the second frame rate. In certain embodiments, the second frame rate is higher than the third frame rate. In some embodiments, J is relative to the frame rate. <M<Nである。

[0046]

[0052] According to certain embodiments, the georegistration processor 116 processes video frames for reference georegistration at a dynamic frame rate. In some embodiments, the georegistration processor 116 performs georegistration at a first frame rate in a first time period. In certain embodiments, the georegistration processor 116 performs georegistration at a second frame rate in a second time period, where the first frame rate is different from the second frame rate. In some embodiments, the georegistration processor 116 is configured to perform georegistration when processing resources (e.g., computing processing units (CPU), graphics processing units (GPU)) are available.

[0047]

[0053] According to some embodiments, the estimation processor 118 can process video frames at full FPS (e.g., per frame). In certain embodiments, the estimation processor 118 can integrate various feeds into an estimation (e.g., a global estimation). In some embodiments, the estimation processor 118 implements a Kalman filter algorithm and / or similar algorithms to synthesize an estimation of geographic coordinates (e.g., true geographic coordinates) based on various observations provided by other processors (e.g., a calibration module, an optical flow processor 112, a georegistration processor 116, etc.). In certain embodiments, the visual navigation system 102 is adapted to the variable availability of different streams, including dropouts or missing processors, to continue providing estimations (e.g., the best available estimation) and predicted confidence. In some embodiments, the estimation processor 118 requires a lightweight computation cost (e.g., a few milliseconds, 2-3 milliseconds).

[0048]

[0054] In certain embodiments, the visual navigation system 102 and / or the georegistration processor 116 include a projection processor for projection. In some embodiments, the visual navigation system 102 projects onto a DEM (Digital Elevation Model). In certain embodiments, the projection processor is a standalone processor. In some embodiments, the projection processor is part of the georegistration processor 116 and / or part of the estimation processor 118.

[0049]

[0055] In some embodiments, the visual navigation environment 100 includes a repository (not shown) that can store and / or contain video, video frames, metadata, geolocation information, reference images, georegistration transformations, image-based transformations, etc. The repository 430 can be implemented using any one of the configurations described below. The data repository may include random access memory, flat files, XML files, and / or one or more database management systems (DBMS) running on one or more database servers or data centers. The database management system may be a relational (RDBMS), hierarchical (HDBMS), multidimensional (MDBMS), object-oriented (ODBMS or OODBMS), or object-relational (ORDBMS) database management system, etc. The data repository may be, for example, a single relational database. In some cases, the data repository may include multiple databases from which data can be exchanged and aggregated by a data integration process or software application. In exemplary embodiments, at least a portion of the data repository may be hosted in a cloud data center. In some cases, the data repository may be hosted on a single computer, server, storage device, cloud server, etc. In other cases, a data repository may be hosted on a set of networked computers, servers, or devices. In some cases, a data repository may be hosted on a hierarchy of data storage devices, including local, regional, and central storage.

[0050]

[0056] In some cases, various components within the visual navigation environment 100 may execute software or firmware stored on a non-temporary computer-readable medium to perform various processing steps. The various components and processors of the visual navigation environment 100 may be implemented by one or more computing devices, including but not limited to circuits, computers, cloud-based processing units, processors, processing units, microprocessors, mobile computing devices, and / or tablet computers. In some cases, various components of the visual navigation environment 100 (e.g., the visual navigation system 102, the visual navigation module 110, SIP 120A / 120B, the user system 132, the control system 134, the position system 136, etc.) may be implemented on a shared computing device. Alternatively, the components of the visual navigation environment 100 may be implemented on multiple computing devices. In some implementations, the various modules and components of the visual navigation environment or workflow 100 may be implemented as software, hardware, firmware, or a combination thereof. In some cases, the various components of the visual navigation environment or workflow 100 may be implemented as software or firmware executed by a computing device.

[0051]

[0057] Various components of the visual navigation environment or workflow 100 can communicate or connect via communication interfaces, such as wired or wireless interfaces. These communication interfaces include, but are not limited to, any wired or wireless short-range and long-range communication interfaces. Short-range communication interfaces may be, for example, a local area network (LAN), an interface conforming to a known communication standard, such as Bluetooth®, IEEE 802 (e.g., IEEE 802.11), ZigBee®, or similar specifications, such as those based on IEEE 802.15.4, or other public or proprietary wireless protocols. Long-range communication interfaces may be, for example, a wide area network (WAN), a cellular network interface, or a satellite communication interface. Communication interfaces may be located within a private computer network, such as an intranet, or on a public computer network, such as the Internet.

[0052]

[0058] Figure 2 is a simplified diagram illustrating Method 200 for Visual Navigation according to a particular embodiment of the present disclosure. This figure is merely an example. Those skilled in the art will recognize many variations, alternatives, and modifications. Method 200 for Visual Navigation includes processes 210, 215, 220, 225, 230, 235, and 240. The above is shown using a selected group of processes for Method 200 for Visual Navigation, but many alternatives, modifications, and variations may exist. For example, some of the processes may be extended and / or combined. Other processes may be inserted into the processes described above. Depending on the embodiment, the sequence of processes may be replaced with others. Further details of these processes are found throughout the present disclosure.

[0053]

[0059] In some embodiments, some or all processes (e.g., steps) of Method 200 are executed by a system (e.g., a computing system 700). In certain examples, some or all processes (e.g., steps) of Method 200 are executed by a computer and / or processor directed by code. For example, the computer includes a server computer and / or client computers (e.g., a personal computer). In some examples, some or all processes (e.g., steps) of Method 200 are executed according to instructions contained by a non-temporary computer-readable medium (e.g., in a computer program product such as a computer-readable flash drive). For example, the non-temporary computer-readable medium is readable by a computer including a server computer and / or client computers (e.g., a personal computer and / or a server rack). As an example, the instructions contained by the non-temporary computer-readable medium are executed by a processor including the processor of the server computer and / or the processor of the client computer (e.g., a personal computer and / or a server rack).

[0054]

[0060] According to some embodiments, in process 210, the system receives an input video containing multiple video frames. In certain embodiments, the multiple video frames are a sequence of video frames. In some embodiments, in process 215, the system obtains a first geolocation of an aircraft (e.g., a drone, UAV, etc.) at a first time t1. In certain embodiments, the first geolocation is received from a GPS system.

[0055]

[0061] In certain embodiments, in process 220, the system determines one or more pixel motions between frames for each video frame within a plurality of video frames. In some embodiments, the pixel motion is image-based motion, which is movement determined based on the video frame, e.g., image features within the video frame. In certain embodiments, in process 225, the system estimates the movement (e.g., displacement) of the aircraft from a first time t1 to a second time t2, at least in part on one or more pixel motions and metadata associated with the aircraft (e.g., speed, orientation, direction, velocity, etc.). In some embodiments, the system estimates the geolocation of the aircraft based on a first geolocation and the estimated motion.

[0056]

[0062] According to some embodiments, for each frame, the system estimates the aircraft's geolocation based on a first geolocation and estimated motion. For example, in a three-dimensional vector space, "aircraft's geolocation at time t2" = "aircraft's geolocation at time t1" + "aircraft's motion from time t1 to time t2".

[0057]

[0063] In certain embodiments, in process 235, once every multiple frames, the system generates a georegistration transformation (e.g., a georegistration technique using image matching, a georegistration technique using AWOG, etc.). In some embodiments, in process 240, the system refines the estimated geolocation of the aircraft based on the georegistration transformation. In some embodiments, the georegistration transformation may be a transformation from the geographic location of a first image to the geographic location of a second image, so that given a first image having the first geographic location, the georegistration transformation can be used to calculate the second geographic location of the second image. In certain embodiments, the system determines whether the use of the georegistration technique is accurate enough to justify the consumption of additional computing power. In some embodiments, if yes, process 240 is performed; otherwise, process 240 is not performed. In some examples, the use of the georegistration technique in areas with few topographic features (e.g., "above water") is not accurate enough to justify the consumption of additional computing power.

[0058]

[0064] Figure 3 is a simplified diagram illustrating Method 300 for Visual Navigation according to a particular embodiment of the present disclosure. This figure is merely an example. Those skilled in the art will recognize many variations, alternatives, and modifications. Method 300 for Visual Navigation comprises processes 310, 315, 320, 325, 330, 335, and 340. The above is shown using a selected group of processes for Method 300 for Visual Navigation, but many alternatives, modifications, and variations may exist. For example, some of the processes may be extended and / or combined. Other processes may be inserted into the processes described above. Depending on the embodiment, the sequence of processes may be replaced with others. Further details of these processes are found throughout the present disclosure.

[0059]

[0065] In some embodiments, some or all processes (e.g., steps) of Method 300 are executed by a system (e.g., a computing system 700). In certain examples, some or all processes (e.g., steps) of Method 300 are executed by a computer and / or processor directed by code. For example, the computer includes a server computer and / or client computers (e.g., a personal computer). In some examples, some or all processes (e.g., steps) of Method 300 are executed according to instructions contained by a non-temporary computer-readable medium (e.g., in a computer program product such as a computer-readable flash drive). For example, the non-temporary computer-readable medium is readable by a computer including a server computer and / or client computers (e.g., a personal computer and / or a server rack). As an example, the instructions contained by the non-temporary computer-readable medium are executed by a processor including the processor of the server computer and / or the processor of the client computer (e.g., a personal computer and / or a server rack).

[0060]

[0066] In certain embodiments, in process 310, the system receives multiple video frames (e.g., images, image frames, etc.) from image sensors positioned on an aircraft (e.g., a drone, UAV, UA, etc.). In some embodiments, the aircraft may move into a GPS-denied environment, such as an area with GPS information. In certain embodiments, the aircraft's geolocation cannot be retrieved via a GPS system. In some embodiments, the system includes two or more image sensors. In certain embodiments, the system includes two or more image sensors positioned on an aircraft. In some embodiments, the two or more image sensors include two or more types of sensors. In certain embodiments, the system can receive two or more types of video frames from two or more types of sensors (e.g., EO sensors, infrared sensors, SAR sensors). In some embodiments, the system includes SIPs (e.g., SIP 120A, SIP 120B) for tuning the sensor feed and / or computing model. In certain embodiments, the SIPs can perform one or more image transformations.

[0061]

[0067] According to some embodiments, in process 315, the system generates an image-based transformation based on a plurality of video frames. In certain embodiments, the image-based transformation is associated with the motion of one or more image features and the motion of an image sensor. For example, in some embodiments, one or more features of a first image may be matched to one or more corresponding features of a second image such that the movement of the position of one or more features in the first image to the position of the corresponding one or more features in the second image is captured by the image-based transformation. In some embodiments, the transformation corresponds to the motion of features between images (for example, calculated based on a visual measurement of the motion of features across images). In some embodiments, in process 320, the system determines the image-based motion associated with the aircraft based on the image-based transformation.

[0062]

[0068] In certain embodiments, the system analyzes a first video frame of a plurality of video frames to identify one or more first image features in the first video frame. In some embodiments, the system analyzes a second video frame of a plurality of video frames to identify one or more second image features in the second video frame, where each of the one or more second image features matches one of the one or more first image features, and the second video frame is located after the first video frame of the plurality of video frames. In certain embodiments, the system generates one or more motion vectors based on one or more first image features and one or more second image features, where each motion vector corresponds to one of the one or more first image features and the matched second image feature.

[0063]

[0069] In some embodiments, the system determines the motion of an image sensor based on image-based transformations. In certain embodiments, the system generates image-based transformations based on one or more motion vectors (e.g., by combining them). Using the determined image-based motion, the system can determine and / or estimate the geolocation of an aircraft when it is located in a GPS denial area. In certain embodiments, the system is configured to determine image-based transformations and motion at a first frame rate (e.g., per video frame, every other video frame, every M video frames, etc.). In some embodiments, the system is configured to determine image-based transformations and / or motion in all video frames of a plurality of received video frames.

[0064]

[0070] According to some embodiments, in process 325, the system generates a georegistration transform based on at least one video frame from a plurality of video frames and a reference image (e.g., a reference image from reference image 117, a received reference image). Further details regarding the georegistration transform are provided throughout this disclosure. In some embodiments, the georegistration transform is a transformation between the geographic location of a first image and the geographic location of a second image. In certain embodiments, in process 330, the system determines a georegistration-based geolocation associated with an aircraft based on the georegistration transform. In some embodiments, the georegistration transform can be associated with a transformation from the geographic location of a first image to the geographic location of a second image, so that given a first image having a first geographic location, the georegistration transform can be used to calculate a second geographic location of a second image. In some embodiments, the system determines a georegistration transform and / or georegistration-based geolocation at a second frame rate.

[0065]

[0071] In certain embodiments, the second frame rate is different from the first frame rate. In some embodiments, the second frame rate is lower than the first frame rate. In certain embodiments, the system determines image-based transformations and translations for a first subset of video frames in a plurality of video frames. In some embodiments, the system determines georegistration transformations and / or georegistration-based geolocations for a second subset of video frames in a plurality of video frames. In certain embodiments, the second subset of video frames is a smaller subset than the first subset of video frames. In some embodiments, at least one video frame in the first subset of video frames is not in the second subset of video frames. In some embodiments, the second frame rate is a dynamic frame rate that changes over time. In certain embodiments, the second frame rate is a dynamic frame rate that depends on computing resources, for example, having sufficient computing resources to perform georegistration.

[0066]

[0072] In certain embodiments, in process 335, the system receives and / or extracts metadata associated with the aircraft. In some embodiments, the metadata is associated with the aircraft's movement, including but not limited to speed, heading, altitude, and orientation. In certain embodiments, the system (e.g., SIP) is configured to extract metadata from video frames. In some embodiments, the system is configured to extract at least a portion of the metadata from at least one video frame out of a plurality of video frames. In certain embodiments, the system associates at least a portion of the metadata with at least one video frame out of a plurality of video frames.

[0067]

[0073] According to some embodiments, in process 340, the system determines and / or estimates the aircraft geolocation by applying a nonlinear Kalman filter to image-based movement, georegistration-based geolocation, and / or metadata. In certain embodiments, the nonlinear Kalman filter includes an extended Kalman filter and / or an unscented Kalman filter. In some embodiments, the system determines and / or estimates the aircraft geolocation by applying a trained machine learning model (e.g., an estimation machine learning model) to image-based movement, georegistration-based geolocation, and / or metadata.

[0068]

[0074] According to certain embodiments, the system can process video frames received from two or more image sensors to determine the aircraft's position. In some embodiments, the system can process video frames received from two or more types of image sensors to determine the aircraft's position. In certain embodiments, the system determines two or image-based transformations and / or motions for two or more sets of video frames received from two or more image sensors. In some embodiments, the system determines the aircraft's position based at least in part on a second image-based motion.

[0069]

[0075] According to some embodiments, the system can receive aircraft position from a GPS system at one or more specific times (e.g., periodically when the aircraft is in a GPS-enabled area). In certain embodiments, the system determines the aircraft geolocation at a later time based at least in part on the received aircraft position. In some embodiments, the system returns to process 310 to receive additional video frames (e.g., via a video stream) and continues to determine the aircraft geolocation based at least in part on the additional video frames.

[0070]

[0076] Figure 4 is a simplified diagram showing a software architecture 400 for a video georegistration system according to a particular embodiment of the present disclosure. This figure is merely an example. Those skilled in the art will recognize many variations, alternatives, and modifications. The architecture 400 for a video georegistration system includes components and processes 410, 420, 423, 425, 427, 430, 435, 440, 443, 445, 447, 450, 453, 455, 457, 460, 465, 470, 475, and 480. The above is shown using a selected group of components and processes for the software architecture 400 for a video georegistration system, but many alternatives, modifications, and variations may exist. For example, some of the components and / or processes may be extended and / or combined. Other components and / or processes may be inserted into those described above. Depending on the embodiment, sequences of processes may be replaced with others that are substituted. Further details of these components and processes are found throughout the present disclosure.

[0071]

[0077] According to some embodiments, the video georegistration system receives one or more videos or video streams containing one or more video frames 405 (e.g., 30 frames per second (FPS), 60 FPS, etc.). In certain embodiments, in process 410, the video georegistration system is configured to determine whether to perform a reference georegistration based on the available processor time and the expected reference georegistration execution time. In some embodiments, if the available processor time is shorter than the expected reference georegistration execution time, the video georegistration system does not perform the reference georegistration. In certain embodiments, if the available processor time is longer than the expected reference georegistration execution time, the video georegistration system continues to perform the reference georegistration.

[0072]

[0078] According to certain embodiments, the video georegistration system generates corrected telemetry 425 based on raw telemetry 420, calibration results 423, and / or previously filtered results (e.g., Kalman filter aggregate results) 427. In some embodiments, the raw telemetry 420 is extracted from video frames of the received video, video stream, and / or one or more video frames 405. In certain embodiments, calibration is essentially cleaning up the video telemetry based on common failure modes. In some embodiments, calibration includes interpolating missing frames in the telemetry. In certain embodiments, calibration includes aligning video frames if they came out of sync. In some embodiments, calibration includes the ability to incorporate basically known corrections, such as previously filtered results 427. In certain embodiments, calibration can correct several types of video feeds that exhibit systematic errors. In some examples, systematic errors include errors in the field of view (e.g., lens angle). For example, a lens angle of 5.1 degrees may actually be 5.25 degrees, and such deviations may be used when performing calibration.

[0073]

[0079] According to some embodiments, in process 430, the video georegistration system generates a candidate grid of geopoints (e.g., a grid of pixel (latitude, longitude) pairs) using corrected telemetry to generate an unaligned grid 435. In certain embodiments, in process 445, the video georegistration system is configured to retrieve a reference image based on the unaligned grid. In some embodiments, the video georegistration system retrieves a reference image using a reference image service 440, a local reference image cache 443, and / or a previously aligned frame 447. In certain embodiments, the video georegistration system can retrieve or generate a reference image in synchronization with the input video (e.g., within one hour) using, for example, the reference image service 440. In some embodiments, the video georegistration system can retrieve a reference image using the local reference image cache 443. In certain embodiments, the video georegistration system can use a previously aligned frame 447 (e.g., a reference image used in a previously aligned frame).

[0074]

[0080] According to certain embodiments, a reference image (e.g., a reference image) can be generated based on the geographic coordinates of an unaligned grid. In some embodiments, the reference image is retrieved from a local reference image cache 443, for example, on the same edge device where at least part of the video georegistration system is running. In certain embodiments, the reference image is generated, stored, and / or retrieved from the same edge environment (e.g., on a physical plane). In some embodiments, the georegistration system avoids sending usage requests over high-latency connections. In certain embodiments, the georegistration system can use a pre-bundled set of tiles or a shifted pre-bundled set of tiles as the reference image. In some embodiments, the georegistration system and / or another system supports the generation of a local basemap for reference image creation (e.g., rapid generation).

[0075]

[0081] According to some embodiments, a georegistration system can combine a platform (e.g., a platform that utilizes satellite technology for autonomous decision-making) with other available satellite image sources to automatically extract satellite imagery of an area (e.g., an area associated with an input video, an area associated with a video frame, an area associated with an unaligned grid) and automatically construct a basemap within that area in temporal proximity (e.g., within 4 hours, within 1 hour).

[0076]

[0082] In certain embodiments, in process 450, the video georegistration system selects a template pattern and generates a video frame template 453 (e.g., a template slice). In some embodiments, in process 455, the video georegistration system distorts a reference image around the template to match the video frame angle and generates a distorted template 457 (e.g., a template slice) of the reference image. In certain embodiments, in process 460, the video georegistration system performs AWOG matching and / or other matching algorithms on the template and the distorted template, also called a template pair, to generate a calculated registration shift 465 of the template pair.

[0077]

[0083] According to some embodiments, the video georegistration system recursively generates template registrations. In certain embodiments, in process 470, the video georegistration system combines the template registrations to generate frame registrations (e.g., image transformations) for video frames. In some embodiments, in process 475, the video georegistration system uses the frame registrations to update the grid to generate a grid 480 aligned to a geographic coordinate system, which can then be used downstream as geographic information for video frames. In certain embodiments, the video georegistration system can use the frame registrations to generate video frames aligned to a geographic coordinate system.

[0078]

[0084] Figure 5 is a simplified diagram illustrating a method 500 for video georegistration according to a particular embodiment of the present disclosure. This figure is merely an example. Those skilled in the art will recognize many variations, alternatives, and modifications. The method 500 for sorting templates and / or generating a template queue includes processes 510, 515, 520, 525, 530, 535, and 540. The above is illustrated using a selected group of processes for the method 500 for sorting templates and / or generating a template queue, but many alternatives, modifications, and variations may exist. For example, some of the processes may be extended and / or combined. Other processes may be inserted into the processes described above. Depending on the embodiment, the sequence of processes may be replaced with others. Further details of these processes are found throughout the present disclosure.

[0079]

[0085] In certain embodiments, in process 510, the georegistration system receives an input video containing multiple video frames. In certain embodiments, in process 515, the video georegistration system is configured to start a reference georegistration process to the video frames at a certain frame rate. In some embodiments, the frame rate is a fixed frame rate. In certain embodiments, the frame rate is a dynamic frame rate (e.g., not a fixed frame rate). In some embodiments, the video georegistration system decides whether to perform reference georegistration based on the available processor time and the expected reference georegistration execution time. In some embodiments, if the available processor time is shorter than the expected reference georegistration execution time, the video georegistration system does not perform reference georegistration. In certain embodiments, if the available processor time is longer than the expected reference georegistration execution time, the video georegistration system continues to perform reference georegistration.

[0080]

[0086] In certain embodiments, in process 520, the video georegistration system identifies geographic information associated with video frames. In some embodiments, the video georegistration system generates corrected telemetry based on raw telemetry, calibration results, and / or previous filtered results (e.g., Kalman filtered results). In some embodiments, raw telemetry is extracted from video frames of the received video, video stream, and / or one or more video frames. In certain embodiments, calibration is essentially cleaning up the video telemetry based on common failure modes. In some embodiments, calibration includes interpolating missing frames in the telemetry. In certain embodiments, calibration includes aligning video frames if they come in staggered order. In some embodiments, calibration includes the ability to incorporate essentially known corrections, such as previous filtered results. In certain embodiments, calibration can correct several types of video feeds that exhibit systematic errors.

[0081]

[0087] According to some embodiments, the video georegistration system generates a candidate grid of geopoints (e.g., a grid of pixel (latitude, longitude) pairs) using corrected telemetry and generates an unaligned grid. In certain embodiments, in process 525, the video georegistration system is configured to generate or select a reference image based at least in part on geographic information associated with video frames (e.g., an unaligned grid). In some embodiments, the video georegistration system retrieves a reference image using a reference image service, a local reference image cache, and / or previously aligned frames. In certain embodiments, the video georegistration system can retrieve or generate a reference image in sync with the input video (e.g., within one hour) using, for example, a reference image service. In some embodiments, the video georegistration system can retrieve a reference image using a local reference image cache. In certain embodiments, the video georegistration system can use previously aligned frames (e.g., reference images used in previously aligned frames). In some embodiments, the video georegistration system can generate a reference image by combining multiple images.

[0082]

[0088] According to certain embodiments, the reference image can be generated based on the geographic coordinates of an unaligned grid. In some embodiments, the reference image is retrieved from a local reference image cache, for example, on the same edge device where at least part of the video georegistration system is running. In certain embodiments, the reference image is generated, stored, and / or retrieved from the same edge environment (e.g., on a physical plane). In some embodiments, the georegistration system avoids sending usage requests over high-latency connections. In certain embodiments, the georegistration system can use a pre-bundled set of tiles or a shifted pre-bundled set of tiles as the reference image. In some embodiments, the georegistration system and / or another system supports the generation of a local basemap for reference image creation (e.g., rapid generation).

[0083]

[0089] According to some embodiments, a georegistration system can be coupled with a metaconstellation platform having other available satellite image sources to automatically extract satellite imagery of an area (e.g., an area associated with an input video, an area associated with a video frame, an area associated with an unaligned grid) and automatically construct a basemap within that area in temporal proximity (e.g., within 4 hours, within 1 hour).

[0084]

[0090] In certain embodiments, in process 530, the video georegistration system generates a georegistration transformation based at least partially on a reference image. In some embodiments, the video georegistration system selects a template pattern and generates a template of video frames (e.g., a template slice). In some embodiments, the video georegistration system distorts the reference image around the template to match the video frame angle and generates a distorted template of the reference image (e.g., a template slice). In certain embodiments, the video georegistration system performs AWOG matching and / or other matching algorithms on the template and the distorted template, also called a template pair, to generate a calculated registration shift of the template pair.

[0085]

[0091] According to some embodiments, the video georegistration system recursively generates template registrations. In certain embodiments, the video georegistration system combines the template registrations to generate frame registrations (e.g., image transformations) for video frames. In some embodiments, the video georegistration system can use the frame registrations to update a grid to generate a grid aligned to a geographic coordinate system, which can then be used downstream as geographic information for video frames.

[0086]

[0092] In certain embodiments, in process 535, the video georegistration system applies a georegistration transformation to video frames to generate aligned video frames (e.g., video frames aligned to a geographic coordinate system). In some embodiments, in process 540, the video georegistration system outputs aligned video frames. In certain embodiments, the video georegistration system recursively performs steps 515-540 to successively generate video frames aligned to a geographic coordinate system and / or videos aligned to a geographic coordinate system.

[0087]

[0093] Figure 6 is a simplified diagram illustrating a method 600 for generating a transformation (e.g., image transformation) according to a particular embodiment of the present disclosure. This figure is merely an example. Those skilled in the art will recognize many variations, alternatives, and modifications. The method 600 for generating a transformation (e.g., image transformation) comprises processes 610, 615, 620, 625, 630, 635, and 640. The above is shown using a selected group of processes for the method 600 for generating a transformation (e.g., image transformation), but many alternatives, modifications, and variations may exist. For example, some of the processes may be extended and / or combined. Other processes may be inserted into the processes described above. Depending on the embodiment, the sequence of processes may be replaced with others. Further details of these processes are found throughout the present disclosure.

[0088]

[0094] In a particular embodiment, in process 610, the georegistration system performs the image transformation operation N times. In some embodiments, the georegistration system performs the image transformation operation iteratively. In a particular embodiment, in process 615, the georegistration system randomly selects several points (e.g., three points), and in process 620, the georegistration system calculates a transformation that matches the selected points. In some embodiments, the georegistration system randomly selects a predetermined number of points and calculates a transformation (e.g., an affine transformation) that matches those selected points. In a particular embodiment, the georegistration system selects one point for translation.

[0089]

[0095] According to some embodiments, in process 625, the georegistration system applies a nonlinear algorithm (e.g., the Levenberg-Marquardt nonlinear algorithm) to determine the error associated with the transformation. In certain embodiments, the georegistration system applies the nonlinear algorithm to the sum of the distances (e.g., Lorentz distances) between the shift values ​​(e.g., preferred shift values) of all points. In certain embodiments, the shift value of each point is weighted by the intensity value of each point when determining the error.

[0090]

[0096] In certain embodiments, in process 630, if the error is lower than all previous transformations, the transformation is designated as a candidate transformation (e.g., best candidate). In some embodiments, in process 635, the georegistration system determines whether N iterations have been completed. In certain embodiments, if N iterations have not been completed, the georegistration system returns to process 610. In some embodiments, if N iterations have been completed, the georegistration system proceeds to process 640. In certain embodiments, in process 640, the georegistration system returns a candidate transformation (e.g., best candidate transformation).

[0091]

[0097] Figure 7 is a simplified diagram showing a computing system for implementing System 700 for georegistration, according to at least one example described herein. This diagram is merely an example and should not unduly limit the scope of the claims. Those skilled in the art will recognize many variations, alternatives, and modifications.

[0092]

[0098] The computing system 700 includes a bus 702 or other communication mechanism for communicating information, a processor 704, a display 706, a cursor control component 708, an input device 710, main memory 712, read-only memory (ROM) 714, a storage unit 716, and a network interface 718. In some embodiments, some or all of the processes (e.g., steps) of methods / processes 200, 300, 400, 500, and / or 600 are performed by the computing system 700. In some examples, the bus 702 is coupled to the processor 704, the display 706, the cursor control component 708, the input device 710, the main memory 712, the read-only memory (ROM) 714, the storage unit 716, and / or the network interface 718. In certain examples, the network interface is coupled to a network 720. For example, the processor 704 includes one or more general-purpose microprocessors. In some examples, main memory 712 (e.g., random access memory (RAM), cache, and / or other dynamic storage devices) is configured to store information and instructions executed by processor 704. In certain examples, main memory 712 is configured to store temporary variables or other intermediate information during the execution of instructions by processor 704. For example, instructions, once stored in a storage unit 716 accessible to processor 704, render computing system 700 into a dedicated machine customized to perform the operations specified by the instructions. In some examples, ROM 714 is configured to store static information and instructions for processor 704. In certain examples, storage unit 716 (e.g., magnetic disk, optical disk, or flash drive) is configured to store information and instructions.

[0093]

[0099] In some embodiments, a display 706 (e.g., a cathode ray tube (CRT), an LCD display, or a touchscreen) is configured to display information to a user of the computing system 700. In some examples, an input device 710 (e.g., alphanumeric and other keys) is configured to communicate information and commands to a processor 704. For example, a cursor control component 708 (e.g., a mouse, trackball, or cursor directional keys) is configured to communicate additional information and commands to the processor 704 (e.g., to control the movement of a cursor on the display 706).

[0094]

[0100] According to a particular embodiment, a method for video georegistration is provided. The method includes receiving an input video comprising a plurality of video frames; calibrating a first set of video frames selected from the plurality of video frames to generate a first set of calibrated video frames using a calibration transform; performing one or more reference georegistrations to a second set of video frames selected from the plurality of video frames to generate a video georegistration transform using the second set of video frames, wherein the second set of video frames has fewer video frames than the first set of video frames; and generating an output video using the calibration transform and the video georegistration transform, the method being performed using one or more processors. For example, the method is carried out according to at least Figures 1, 2, 3, 4, and / or 5.

[0095]

[0101] In some embodiments, the video transformation includes one or more video frame transformations corresponding to a second set of video frames. In certain embodiments, the method further includes applying optical flow estimation to a third set of video frames among a plurality of video frames, wherein the third set of video frames has fewer video frames than the first set of video frames and more video frames than the second set of video frames.

[0096]

[0102] According to several embodiments, a method for visual navigation is provided. The method includes receiving a plurality of video frames from an image sensor positioned on an aircraft; generating an image-based transformation based on the plurality of video frames, wherein the image-based transformation is associated with the motion of one or more image features and the motion of the image sensor; determining an image-based movement associated with the aircraft based on the image-based transformation; generating a georegistration transformation based on at least one video frame from the plurality of video frames and a reference image; determining a georegistration-based geolocation associated with the aircraft based on the georegistration transformation; and determining the aircraft geolocation by applying a nonlinear Kalman filter to the image-based movement and the georegistration-based geolocation, the method being performed using one or more processors. For example, the method is carried out according to at least Figures 1, 2, 3, and / or 4.

[0097]

[0103] In certain embodiments, the method further includes receiving metadata associated with the movement of an aircraft, the metadata including at least one of speed, bearing, altitude, and orientation, and determining the aircraft position includes determining the aircraft position by image-based motion, georegistration-based geolocation, and applying a nonlinear Kalman filter to the metadata. In some embodiments, the method further includes extracting at least a portion of the metadata from at least one video frame of a plurality of video frames. In certain embodiments, the nonlinear Kalman filter is an unscented Kalman filter. In some embodiments, generating an image-based transformation includes analyzing a first video frame of a plurality of video frames to identify one or more first image features in the first video frame; analyzing a second video frame of a plurality of video frames to identify one or more second image features in the second video frame, wherein each of the one or more second image features matches a first image feature from the one or more first image features, and the second video frame is located after the first video frame of the plurality of video frames; generating one or more motion vectors based on the one or more first image features and the one or more second image features, wherein each motion vector corresponds to one of the one or more first image features and the matched second image feature; and generating an image-based transformation based on the one or more motion vectors.

[0098]

[0104] In certain embodiments, generating image-based transformations includes generating one or more image-based transformations for a first set of video frames at a first frame rate, and generating georegistration transformations includes generating georegistration transformations for a second set of video frames at a second frame rate, where the first frame rate is different from the second frame rate. In some embodiments, the first frame rate is higher than the second frame rate, and the second set of video frames is a subset of multiple video frames. In certain embodiments, the first frame rate is higher than the second frame rate, and the first set of video frames includes each video frame of multiple video frames. In some embodiments, the method further includes determining the motion of an image sensor based on the image-based transformations.

[0099]

[0105] In certain embodiments, the plurality of video frames are a first plurality of video frames, the image sensor is a first image sensor, and the method includes receiving a second plurality of video frames from a second image sensor, the second image sensor being different from the first image sensor, generating a second image-based transformation based on the second plurality of video frames, determining a second image-based movement associated with an aircraft based on the second image-based transformation, and determining the aircraft geolocation based at least in part on the second image-based movement. In some embodiments, the method further includes receiving a first aircraft position, and determining the aircraft geolocation includes determining the aircraft position based at least in part on the first aircraft position.

[0100]

[0106] According to several embodiments, a system for visual navigation is provided. In some embodiments, the system includes at least one processor and at least one memory for storing instructions, which, when executed by the at least one processor, cause the system to perform a series of operations. In some embodiments, the series of operations includes receiving a plurality of video frames from an image sensor positioned on an aircraft and generating an image-based transformation based on the plurality of video frames. In some embodiments, the image-based transformation is associated with the motion of one or more image features and the motion of the image sensor. In some embodiments, the series of operations further includes determining an image-based movement associated with the aircraft based on the image-based transformation, generating a georegistration transformation based on at least one video frame from the plurality of video frames and a reference image, determining a georegistration-based geolocation associated with the aircraft based on the georegistration transformation, and determining the aircraft geolocation by applying a nonlinear Kalman filter to the image-based movement and the georegistration-based geolocation.

[0101]

[0107] In some embodiments, a series of operations includes receiving metadata associated with the movement of an aircraft, the metadata including at least one selected from the group consisting of speed, bearing, altitude, and orientation, and determining the aircraft's position includes determining the aircraft's position by image-based motion, georegistration-based geolocation, and by applying a nonlinear Kalman filter to the metadata. In some embodiments, a series of operations includes extracting at least a portion of the metadata from at least one video frame among a plurality of video frames.

[0102]

[0108] In some embodiments, the nonlinear Kalman filter is an unscented Kalman filter. In some embodiments, generating an image-based transformation includes analyzing a first video frame of a plurality of video frames to identify one or more first image features in the first video frame, and analyzing a second video frame of a plurality of video frames to identify one or more second image features in the second video frame. In some embodiments, each of the one or more second image features matches one of the one or more first image features. In some embodiments, the second video frame is located after the first video frame of the plurality of video frames. In some embodiments, generating an image-based transformation further includes generating one or more motion vectors based on one or more first image features and one or more second image features. In some embodiments, each motion vector of the one or more motion vectors corresponds to one of the one or more first image features and the matched second image feature. In some embodiments, generating an image-based transformation further includes generating an image-based transformation based on one or more motion vectors.

[0103]

[0109] In some embodiments, generating image-based transformations includes generating one or more image-based transformations for a first set of video frames at a first frame rate. In some embodiments, generating georegistration transformations includes generating georegistration transformations for a second set of video frames at a second frame rate. In some embodiments, the first frame rate is different from the second frame rate. In some embodiments, the sequence of operations further includes determining the motion of an image sensor based on the image-based transformations.

[0104]

[0110] According to several embodiments, a method for visual navigation is provided. In some embodiments, the method includes receiving a plurality of video frames from an image sensor positioned on an aircraft and generating an image-based transformation based on the plurality of video frames. In some embodiments, the image-based transformation is associated with the motion of one or more image features and the motion of the image sensor. In some embodiments, the method further includes determining an image-based movement associated with the aircraft based on the image-based transformation, generating a georegistration transformation based on at least one video frame from the plurality of video frames and a reference image, determining a georegistration-based geolocation associated with the aircraft based on the georegistration transformation, receiving metadata associated with the aircraft, and estimating the aircraft geolocation by applying a nonlinear Kalman filter to the image-based movement, georegistration-based geolocation, and metadata. In some embodiments, the method is performed using one or more processors.

[0105]

[0111] In some embodiments, the metadata includes the aircraft's speed, heading, and altitude.

[0106]

[0112] For example, some or all components of the various embodiments of this disclosure may be implemented individually and / or in combination with at least another component, using one or more software components, one or more hardware components, and / or one or more combinations of software and hardware components. In another example, some or all components of the various embodiments of this disclosure may be implemented individually and / or in combination with at least another component, in one or more circuits, such as one or more analog circuits and / or one or more digital circuits. In yet another example, while the embodiments described above refer to specific features, the scope of this disclosure also includes embodiments having different combinations of features and embodiments that do not necessarily include all of the features described. In yet another example, the various embodiments and / or examples of this disclosure may be combined.

[0107]

[0113] In addition, the methods and systems described herein can be implemented on many different types of processing devices by program code, which includes program instructions that can be executed by a device processing subsystem. The software program instructions may include source code, object code, machine code, or any other stored data that can be operated to cause a processing system (e.g., one or more components of a processing system) to execute the methods and operations described herein. However, other implementations, such as firmware or more appropriately designed hardware, configured to execute the methods and systems described herein, may also be used.

[0108]

[0114] System and method data (e.g., associations, mappings, data inputs, data outputs, intermediate data results, final data results, etc.) may be stored and implemented in one or more different types of computer-implemented data stores, such as different types of storage devices and programming structures (e.g., RAM, ROM, EEPROM, flash memory, flat files, databases, programming data structures, programming variables, IF-THEN (or similar) statement structures, application programming interfaces, etc.). Note that data structures describe a format for use when organizing and storing data in databases, programs, memory, or other computer-readable media for use by computer programs.

[0109]

[0115] The systems and methods described herein may be provided on many different types of computer-readable media, including computer storage mechanisms (e.g., CD-ROMs, diskettes, RAM, flash memory, computer hard drives, DVDs, etc.), which include instructions (e.g., software) for performing the operations described herein and for use when executed by a processor to implement the systems. The computer components, software modules, functions, data stores, and data structures described herein may be connected to one another directly or indirectly to enable the flow of data necessary for their operation. It should also be noted that a module or processor may include a unit of code that performs software operations and may be implemented, for example, as a subroutine unit of code, as a software function unit of code, as an object (such as in the object-oriented paradigm), as an applet, in a computer scripting language, or as another type of computer code. The software components and / or functions may be located on a single computer or distributed across multiple computers, depending on the circumstances at hand.

[0110]

[0116] A computing system may include client devices and servers. Client devices and servers are generally geographically separated and typically interact via a communication network. The relationship between client devices and servers arises from computer programs running on each computer, and they have a client-device-server relationship with one another.

[0111]

[0117] This specification includes many details relating to specific embodiments. Certain features described herein in relation to separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in relation to a single embodiment may also be implemented in multiple embodiments, separately or in any suitable partial combination. Furthermore, features may be described above as acting in specific combinations, but one or more features from a combination may, in some cases, be removed from the combination, and the combination may, for example, be a partial combination or a variation of a partial combination.

[0112]

[0118] Similarly, although the operations are shown in a specific order in the drawings, this should not be understood as meaning that such operations must be performed in a specific or sequential order shown, or that all illustrated operations must be performed, in order to achieve the desired result. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged in multiple software products.

[0113]

[0119] While specific embodiments of this disclosure have been described, those skilled in the art will understand that there are other embodiments equivalent to those described. Therefore, it should be understood that the present invention should not be limited by the specifically illustrated embodiments. Various modifications and variations of the disclosed embodiments will be apparent to those skilled in the art. The embodiments described herein are illustrative. Unless otherwise indicated, features of one disclosed embodiment may apply to all other disclosed embodiments. It should also be understood that all U.S. patents, patent application publications, and other patents and non-patent literature referenced herein are incorporated by reference, insofar as they do not conflict with the foregoing disclosures.

Claims

1. A method for visual navigation, Receiving multiple video frames from image sensors placed on an aircraft, The process involves generating an image-based transformation based on the plurality of video frames, wherein the image-based transformation is associated with the motion of one or more image features and the motion of the image sensor. Determining the image-based movement associated with the aircraft based on the aforementioned image-based transformation, A georegistration transformation is generated based on at least one video frame from the plurality of video frames and a reference image. Based on the aforementioned georegistration conversion, determine the georegistration-based geolocation associated with the aircraft, The aircraft geolocation is determined by applying a nonlinear Kalman filter to the image-based movement and the georegistration-based geolocation. Includes, A method that is executed using one or more processors.

2. Receiving metadata associated with the movement of the aforementioned aircraft It further includes, The metadata includes at least one selected from the group consisting of speed, bearing, altitude, and orientation. The method according to claim 1, wherein determining the aircraft position includes determining the aircraft position by applying the image-based movement, the georegistration-based geolocation, and the metadata to the nonlinear Kalman filter.

3. Extracting at least a portion of the metadata from at least one video frame among the plurality of video frames. The method according to claim 2, further comprising:

4. The method according to claim 1, wherein the nonlinear Kalman filter is an unscented Kalman filter.

5. The generation of the image-based transformation is Analyzing the first video frame of the plurality of video frames to identify one or more first image features in the first video frame, Analyzing a second video frame of the plurality of video frames to identify one or more second image features in the second video frame, wherein each of the one or more second image features matches a first image feature among the one or more first image features, and the second video frame is located after the first video frame in the plurality of video frames. The method involves generating one or more motion vectors based on the one or more first image features and the one or more second image features, wherein each motion vector corresponds to one of the one or more first image features and the matching second image feature. To generate the image-based transformation based on the one or more motion vectors mentioned above. The method according to claim 1, including the method described in claim 1.

6. The generation of image-based transformations includes generating one or more image-based transformations for a first set of video frames at a first frame rate. Generating the georegistration transform includes generating the georegistration transform for a second set of video frames at a second frame rate, The method according to claim 1, wherein the first frame rate is different from the second frame rate.

7. The method according to claim 6, wherein the first frame rate is higher than the second frame rate, and the second set of video frames is a subset of the plurality of video frames.

8. The method according to claim 6, wherein the first frame rate is higher than the second frame rate, and the first set of video frames includes each of the plurality of video frames.

9. The movement of the image sensor is determined based on the aforementioned image-based transformation. The method according to claim 1, further comprising:

10. The method according to claim 1, wherein the plurality of video frames are a first plurality of video frames, and the image sensor is a first image sensor, The method involves receiving a second set of video frames from a second image sensor, wherein the second image sensor is different from the first image sensor. To generate a second image-based transformation based on the second set of video frames, Determining a second image-based movement associated with the aircraft based on the second image-based transformation, Determining the aircraft geolocation based at least partially on the second image-based movement and Methods that further include the above.

11. Receiving the first aircraft position It further includes, The method according to claim 1, wherein determining the aircraft geolocation includes determining the aircraft position based at least in part on the first aircraft position.

12. At least one processor, At least one memory for storing instructions, wherein when the instructions are executed by the at least one processor, the memory causes the system to perform a series of operations. A system for visual navigation, comprising the series of operations, Receiving multiple video frames from image sensors placed on an aircraft, The process involves generating an image-based transformation based on the plurality of video frames, wherein the image-based transformation is associated with the motion of one or more image features and the motion of the image sensor. Determining the image-based movement associated with the aircraft based on the aforementioned image-based transformation, A georegistration transformation is generated based on at least one video frame from the plurality of video frames and a reference image. Based on the aforementioned georegistration conversion, determine the georegistration-based geolocation associated with the aircraft, The aircraft geolocation is determined by applying a nonlinear Kalman filter to the image-based movement and the georegistration-based geolocation. A system that includes this.

13. The aforementioned series of operations Receiving metadata associated with the movement of the aforementioned aircraft The system according to claim 12, further comprising: The metadata includes at least one selected from the group consisting of speed, bearing, altitude, and orientation. A system for determining the aircraft position, comprising determining the aircraft position by applying the nonlinear Kalman filter to the image-based movement, the georegistration-based geolocation, and the metadata.

14. The aforementioned series of operations Extracting at least a portion of the metadata from at least one video frame among the plurality of video frames. The system according to claim 13, further comprising:

15. The system according to claim 12, wherein the nonlinear Kalman filter is an unscented Kalman filter.

16. The generation of the image-based transformation is Analyzing the first video frame of the plurality of video frames to identify one or more first image features in the first video frame, Analyzing the second video frame of the plurality of video frames to identify one or more second image features in the second video frame, wherein each second image feature of the one or more second image features matches one first image feature of the one or more first image features, and the second video frame is located after the first video frame of the plurality of video frames. The method involves generating one or more motion vectors based on the one or more first image features and the one or more second image features, wherein each motion vector corresponds to one of the one or more first image features and the matching second image feature. To generate the image-based transformation based on the one or more motion vectors mentioned above. The system according to claim 12, including the above.

17. The generation of image-based transformations includes generating one or more image-based transformations for a first set of video frames at a first frame rate. Generating the georegistration transform includes generating the georegistration transform for a second set of video frames at a second frame rate, The first frame rate is different from the second frame rate. The system according to claim 12.

18. The aforementioned series of operations The movement of the image sensor is determined based on the aforementioned image-based transformation. The system according to claim 12, further comprising:

19. A method for visual navigation, Receiving multiple video frames from image sensors placed on an aircraft, The process involves generating an image-based transformation based on the plurality of video frames, wherein the image-based transformation is associated with the motion of one or more image features and the motion of the image sensor. Determining the image-based movement associated with the aircraft based on the aforementioned image-based transformation, A georegistration transformation is generated based on at least one video frame from the plurality of video frames and a reference image. Based on the aforementioned georegistration conversion, determine the georegistration-based geolocation associated with the aircraft, Receiving metadata associated with the aforementioned aircraft, Estimating aircraft geolocation by applying a nonlinear Kalman filter to the image-based movement, georegistration-based geolocation, and metadata. Includes, A method that is executed using one or more processors.

20. The method according to claim 19, wherein the metadata includes at least one selected from the group consisting of the speed, heading, and altitude of the aircraft.