Neural radiation field-based camera alignment
The system uses neural radiation fields to align camera coordinates with vehicle coordinates by capturing multiple images and optimizing orientation parameters, addressing misalignment issues and improving detection accuracy for vehicle safety systems.
Patent Information
- Application Number
- DE102024104219
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-02-15
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2044-02-15
AI Technical Summary
Existing camera systems in vehicles often have misaligned coordinate systems between the camera and the vehicle, leading to inaccuracies in object detection and positioning, which are not effectively addressed by current calibration methods.
A system and method using neural radiation fields (NeRF) to estimate camera orientation relative to a platform by capturing multiple images at different timestamps, generating a rendered portion of the environment, and updating orientation parameters to align the camera coordinate system with the vehicle's coordinate system, utilizing a computer to determine differences and optimize alignment parameters.
Improves the accuracy of camera alignment by correcting orientation errors, enabling precise 3D reconstruction and enhancing vehicle safety features like visibility, lane detection, and pedestrian detection.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical FieldThe present disclosure relates to a system and method for aligning cameras based on neural radiation fields.IntroductionThe orientation of camera and vehicle is useful in vehicles for vision, redundant lane recognition, and pedestrian recognition. A forward facing camera is typically mounted high at the front of the vehicle and is in fixed relationship to the vehicle. The forward facing camera often has a different coordinate system than the vehicle. Therefore, objects detected in the field of view of the look-ahead camera are offset relative to the vehicle coordinate system.Accordingly, those skilled in the art continue their research and development work in the field of real-time detection of the orientation of cameras to vehicles.JEONG, Yoonwoo, et al. Self-Calibrating Neural Radiation Fields. In: 2021 IEEE / CVF International Conference on Computer Vision (ICCV). IEEE Computer Society, 2021. S5826-5834 describes a method for automatic calibration of neural radiance fields (NeRFs) that corrects inaccuracies in camera poses to enable improved 3D reconstruction from inaccurately registered images.DESCRIPTION OF THE INVENTIONThe object of the invention is to provide a system for estimating orientation parameters of a camera relative to a platform. This object is achieved by the subject matter according to claim 1.A system is provided herein. The system includes a platform, a navigation system, a camera, and a computer. The platform is capable of moving through an environment. The navigation system is mounted on the platform and serves to measure a plurality of platform positions of the platform in the environment. The camera is mounted to the platform and generates a plurality of images of the environment into a plurality of time stamps. The computer is mounted on the platform and may receive a first image of the plurality of images from the camera. The first image is captured from a first camera position relative to the environment to a first time stamp of the plurality of time stamps. The computer is further capable of receiving a second image of the plurality of images from the camera. The second image is captured from a second camera position relative to the environment to a second time stamp of the plurality of time stamps. The second time stamp is different from the first time stamp. The computer is further capable of estimating a gaze direction of the first camera pose based on a first platform pose of the plurality of platform poses to the first timestamp, a second platform pose of the plurality of platform poses to the second timestamp, and a plurality of orientation parameters of the camera relative to the platform; generating a rendered portion of the environment using a neural radiation field technique based on the second image and the plurality of orientation parameters; generating a predicted image of the environment as observed along the gaze direction by the rendered portion of the environment; determining one or more differences between the first image and the predicted image; and updating one or more of the plurality of orientation parameters based on the one or more differences.In one or more embodiments of the system, generating the rendered portion includes generating an image conditioned neural radiation field based solely on the second image. The generation of the predicted image is based on the rendered portion with a volume rendering of the image conditioned neural radiation field.In one or more embodiments of the system, the computer is further capable of determining one or more platform maneuver degeneracy conditions before the plurality of platform positions are used. The estimation of the gaze direction is further based on the one or more maneuver degeneracy conditions.In one or more embodiments of the system, the computer is further capable of determining one or more image enable conditions before the second image is used. The generation of the rendered portion is further based on the one or more image release conditions.In one or more embodiments of the system, the computer is further capable of determining one or more orientation parameters as a reduction in a loss function based on a plurality of differences between the first image and the predicted image.In one or more embodiments of the system, the loss function is based on one or more color differences, a two-dimensional positional difference of one or more pairs of features, and a maturation of multiple images from the plurality of images within a time window.In one or more embodiments of the system, reducing the loss function further comprises aligning a camera coordinate system of the camera with a platform coordinate system of the platform.In one or more embodiments of the system, reducing the loss function includes locating the platform over multiple ones of the multiple platform positions.In one or more embodiments of the system, the platform is part of one or more of a land vehicle, a watercraft, or an aircraft.A method for camera alignment based on a neural radiation field is provided herein. The method includes measuring a plurality of platform locations of a platform in an environment with a navigation system; generating a plurality of images of the environment at a plurality of times with a camera mounted on the platform; and receiving a first image of the plurality of images from the camera at a computer. The first image is captured from a first camera position relative to the environment to a first time stamp of the plurality of time stamps. The method includes receiving, at the computer, a second image of the plurality of images from the camera. The second image is captured from a second camera position relative to the environment to a second time stamp of the plurality of time stamps. The second time stamp is different from the first time stamp. The method further comprises estimating a gaze direction of the first camera pose based on a first platform pose of the plurality of platform poses to the first timestamp, a second platform pose of the plurality of platform poses to the second timestamp, and a plurality of orientation parameters of the camera relative to the platform; generating a rendered portion of the environment using a neural radiation field technique based on the second image and the plurality of orientation parameters; generating a predicted image of the environment as observed along the gaze direction by the rendered portion of the environment; determining one or more differences between the first image and the predicted image; and updating one or more of the plurality of orientation parameters based on the one or more differences.In one or more embodiments of the method, generating the rendered portion includes generating an image conditioned neural radiation field based solely on the second image. The generation of the predicted image is based on the rendered portion with a volume rendering of the image conditioned neural radiation field.In one or more embodiments, the method includes determining one or more platform maneuver degradation conditions prior to using the plurality of platform positions. The estimation of the gaze direction is further based on the one or more maneuver degeneracy conditions.In one or more embodiments, the method includes determining one or more image release conditions prior to using the second image. Generating the rendered portion is further based on the one or more image release conditions.In one or more embodiments, the method includes determining one or more orientation parameters as a reduction of a loss function based on a plurality of differences between the first image and the predicted image.In one or more embodiments of the method, the loss function is based on one or more color differences, a two-dimensional position difference of one or more pairs of features, and a maturation of multiple images from the plurality of images within a time window.In one or more embodiments of the method, reducing the loss function is to align a camera coordinate system of the camera with a platform coordinate system of the platform.In one or more embodiments of the method, reducing the loss function includes locating the platform across a plurality of the plurality of platform positions.A vehicle is provided herein. The vehicle includes a navigation system, a camera, and a computer. The navigation system is capable of measuring a plurality of vehicle postures of the vehicle in an environment. The camera may generate a plurality of images of the environment at a plurality of times. The computer is capable of receiving a first image of the plurality of images from the camera. The first image is captured from a first camera position relative to the environment to a first time stamp of the plurality of time stamps. The computer is capable of receiving a second image of the plurality of images from the camera. The second image is captured from a second camera position relative to the environment to a second time stamp of the plurality of time stamps. The second time stamp is different than the first time stamp. The computer is further operable to: estimate a gaze direction of the first camera pose based on a first vehicle position from the plurality of vehicle positions to the first timestamp, a second vehicle position from the plurality of vehicle positions to the second timestamp, and a plurality of orientation parameters of the camera relative to the vehicle; generate a rendered portion of the environment using a neural radiation field technique based on the second image and the plurality of orientation parameters; generate a predicted image of the environment as observed along the gaze direction through the rendered portion of the environment; determine one or more differences between the first image and the predicted image; and update one or more of the plurality of orientation parameters based on the one or more differences.In one or more embodiments, the vehicle includes circuitry coupled to the computer and capable of using the plurality of orientation parameters and the plurality of images.In one or more embodiments of the vehicle, the circuitry is capable of performing one or more automatic driving functions based on the plurality of orientation parameters and the plurality of images.The above features and advantages, as well as other features and advantages of the present disclosure, will be readily apparent from the following detailed description of the best modes for carrying out the disclosure when taken in conjunction with the accompanying drawings.Brief Description of the DrawingsFIG. 1 is a schematic illustration of a system for aligning cameras based on the neural radiation field, according to one or more example embodiments FIG. 2 is a functional flowchart of operations within the system according to one or more example embodiments. FIG. 3 is a perspective diagram of a rendered portion of an environment, according to one or more example embodiments. FIG. 4 is a flowchart of a gaze direction estimate in accordance with one or more example embodiments. FIG. 5 is a flowchart of a single-time neural radiation field technique in accordance with one or more example embodiments. FIG. 6 is a flow diagram of a volume rendering ring in accordance with one or more example embodiments.Detailed DescriptionEmbodiments of the disclosure provide a system and method for an online camera-to-vehicle alignment technique. The camera is generally mounted at a specific position of the vehicle at a fixed height above the ground. For alignment, a neural radiation field (NeRF) and predicted images are used. A first image is captured when the camera is at a first position and in a first camera position. A second image is then recorded at a second position and in a second camera position. A one-shot neural radiation field method uses the second image to create a three-dimensional volume of a portion of the space in front of the camera and the vehicle. A volume rendering method generates a two-dimensional prediction image from the three-dimensional volume.Based on a global positioning system (GPS) / inertial measurement unit (IMU) of the vehicle, changes in an estimated gaze direction of the camera are estimated. The estimated gaze direction is based on the orientation parameters of the camera to the vehicle. A predicted image is generated from the rendered space in the estimated gaze direction. The predicted image is compared to the first image to determine errors in the orientation parameters of the camera. The alignment parameters are then updated to reduce alignment errors.Referring to FIG. 1, a schematic plan diagram illustrating a system 100 for the orientation of cameras based on the neural radiation field is shown, according to one or more example embodiments. The system 100 generally includes an environment 102, a floor 104, and a platform 110. The environment 102 may be an atmosphere through which the platform 110 moves. The ground 104 may be a roadway upon which the platform 110 rests, water upon which the platform 110 floats, and / or the ground.The platform 110 implements a movable machine. The platform 110 may be part of a land vehicle 110 a, an aircraft 110 band / or a watercraft 110 c. The platform generally defines a platform coordinate system 112. Platform 110 generally includes a navigation system 120, a camera 130, a computer 140, and additional circuitry 150 that communicates with each other via a communication bus 160. One or more optional dedicated connections 170 may be included in platform 110 to transmit low latency and / or high speed data.The navigation system 120 includes a positioning system and an inertial system. The navigation system 120 is rigidly mounted to the platform 110. The navigation system 120 is capable of measuring a three-dimensional or two-dimensional position of the platform 110 with respect to the ground 104. The position is reported on the communication bus 160. The navigation system 120 is also capable of measuring a three-dimensional location of the platform 110 relative to the ground 104 or the environment 102. The position is reported on the communication bus 160. In various embodiments, the positioning system includes a global positioning system (GPS) 122. Global positioning system 122 provides geolocation and time information. Other positioning systems may be implemented to meet the criteria of a particular application. In some embodiments, the inertial system is an inertial measurement unit 124. An inertial measurement unit 124 is an electronic device that measures and reports force, angular velocity, and sometimes also orientation with a combination of accelerometer, gyro, and sometimes also magnetometers. Other inertial systems may also be used to meet the design criteria of a particular application.The camera 130 is equipped with a forward camera sensor. In various embodiments, the camera 130 is rigidly attached to or near the front end 114 of the platform 110. In addition to the forward facing camera 130, other cameras may be installed on the platform 110, e.g., cameras on the left side, on the right side, on the rear side, etc. The coordinate systems of the other cameras may also be different from those of the platform (or the vehicle) 110 and may be similarly oriented. The camera 130 and / or the other cameras may be installed at other locations of the platform 110 to cover other fields of view of the environment 102. The camera 130 is capable of capturing a sequence of images of the environment 102 and / or the ground 104 around (e.g., in front of) the platform 110 in a camera direction 134. The image sequence may be associated with a sequence of positions and a sequence of positions of the platform 110. The images relate to a camera coordinate system 132. The camera 130 may be an optical camera operating in the visible spectrum and / or near infrared spectrum. In some embodiments, the camera 130 may include a high speed aperture to limit blurring of the images due to movement of the platform 110. In various embodiments, the images may be transmitted on the communication bus 160. In other embodiments, the images may be transmitted to the computer 140 and / or the additional circuitry 150 via specific connections 170.The computer 140 implements one or more processing circuits. The computer 140 is capable of receiving the poses from the navigation system 120 and the sequence of multiple images from the camera 130. The image sequence includes a first image captured from a first camera pose relative to environment 102 to a first timestamp of a plurality of timestamps and a second image captured from a second camera pose relative to environment 102 to a second timestamp of the plurality of timestamps. The second timestamp is different (e.g., later) than the first timestamp. The timestamps are generated by the computer 140, the navigation system 120, and / or the camera 130. From the images, poses, and positions, the computer 140 estimates a gaze direction of each pixel in the first camera pose based on (i) a first platform pose among multiple platform poses to the first timestamp, (ii) a second platform pose among multiple platform poses to the second timestamp, and (iii) multiple orientation parameters of the camera 130 relative to the platform 110. The viewing direction may be in the coordinates of the given image. A rendered portion of the environment 102 is then generated by the computer 140 using a neural radiation field technique based on the second image and the alignment parameters. The computer 140 then generates a predicted image of the environment 102 as viewed along the gaze direction through the rendered portion of the environment. One or more differences between the first image and the predicted image are calculated by the computer 140. One or more of the alignment parameters are updated based on the one or more differences.In various embodiments, the computer 140 generally includes at least one microcontroller. The at least one microcontroller may include one or more processors, each of which may be embodied as a separate processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or a dedicated electronic control unit. The at least one microcontroller may be an electronic processor (implemented in hardware, software executing on hardware, or a combination of both). The at least one microcontroller may also include tangible, non-volatile memory (e.g., read-only memory in the form of optical, magnetic, and / or flash memory). The at least one microcontroller may include, for example, application-compliant amounts of random access memory, read only memory, flash memory, and other types of electrically erasable programmable read only memory, as well as accompanying hardware in the form of a high speed clock or timer, analog-to-digital and digital-to-analog circuits and input / output circuits and devices, as well as suitable signal conditioning and buffer circuits.Computer readable and executable instructions embodying the present method may be recorded (or stored) in memory and executed as described herein. The executable instructions may be a series of instructions used to execute applications on the at least one microcontroller (either foreground or background). The at least one microcontroller may receive commands and information in the form of one or more input signals from various controllers or components and transmit instructions to the other electronic components.The additional circuit 150 implements a circuit for driver assistance. The additional circuitry 150 assists the driver of a vehicle 110 a- 110 cin speed and / or direction based on the images received from the camera 130 and the orientation parameters of the camera 130 with respect to the platform 110. In some embodiments, the additional circuitry 150 may implement automatic braking functions that respond to obstacles appearing in the images in front of a land vehicle 110 aand then slow and / or stop the land vehicle 110 a. In other embodiments, the additional circuitry 150 may implement steering assist functions that help keep the land vehicles 110 ain the lanes centered on the ground 104. In yet other embodiments, the additional circuitry 150 may implement semi-automatic and / or autonomous driving functions. Other functions, such as perception, localization, and / or mapping, may be implemented in the additional circuitry 150 to meet the design criteria of a particular application. In other embodiments, the circuit 150 and the computer 140 may be implemented as hardware to perform the tasks.The communication bus 160 is a multi-node digital bus. The communication bus 160 is used to transmit the position and posture data from the navigation system 120 to the computer 140. In some embodiments, the communication bus 160 may also transmit the images from the camera 130 to the computer 140 and / or the additional circuitry 150. The orientation parameters of the camera 130 are transmitted over the communication bus 160 from the computer 140 and the additional circuits 150.The dedicated connections 170 are comprised of cables and / or optical cables. The dedicated connections 170 enable the transmission of the images from the camera 130 to the computer 140 and / or to the additional low latency circuitry 150. In various embodiments, dedicated connections 170 may transfer data from computer 140 to additional circuitry 150.Referring to FIG. 2 and back to FIG. 1, a functional flowchart of an example implementation of operations 180 within the system 100 is shown in accordance with one or more example embodiments. Operations 180 receive as input data a first image 182, a second image 184, and dynamics 186 of platform 110. The first image 182 and the second image 184 are received by the computer 140 from the camera 130. Dynamics 186 are transmitted from navigation system 120 to computer 140. Dynamics 186 may include, but are not limited to, the speed, location, direction, position, and the like of platform 110. Operations 180 generate output data (e.g., updated alignment parameters 222 a). The output data / updated orientation parameters 222a include the various parameters for the orientation of the camera to the platform. The output data / updated alignment parameters 222 aare transmitted to the additional circuit 150. Operations 180 generally include steps 190- 214, as shown. The sequence of steps is shown as a representative example. Other step sequences may be implemented to meet the criteria of a particular application.In step 190, the computer 140 may receive the first image 182 from the camera 130. In step 192, the computer 140 may receive the second image 184 from the camera 130. The first image 182 and the second image 184 may be fields or frames. The dynamic data 186 (e.g., speed, position, direction, and pose) is received from the navigation system 120 by the computer 140 at step 194.In step 196, the first image 182 is cached in the computer 140. In step 198, multiple image release conditions of the second image 184 are checked. In step 198 of the image release conditions, it is checked whether the image features (e.g., sharpness and brightness) of the second image 184 are sufficient to be used for further processing. If the second image 184 is too blurred, too dark, and / or washed (too light), the second image 184 may be discarded and the subsequent images checked.The current dynamics 186 of the platform 110 are checked in step 200 by a plurality of motion enabling conditions. The motion enable condition step 200 determines whether the current dynamics 186 represents realistic movements of the platform 110. In the example, the platform 110 may move at a location 221 and may change from a first platform position 223 to a second platform position 225. If the platform 110 has moved too far, too close, and / or too fast since a previous test and / or has rotated too far or too close in pitch, yaw (or direction), and / or roll to be usable (e.g., test for maneuver degradation conditions), the current dynamics 186 is discarded and the subsequent dynamics are tested.In step 202, the computer 140 uses a change 220 in the pose data and uses multiple current (e.g., initial) orientation parameters 222 to calculate a gaze direction of a predicted image 224. The changes 220 in the pose data (e.g., changes in up to six dimensions) generally represent movement of the platform 110 from the first platform pose 223 to the first timestamp to a second platform pose 225 to the second timestamp. The combination of the changes 220 and the orientation parameters 222 generally yields the camera pose changes that the camera 130 represents from a first camera pose 226 associated with the first platform pose 223 for the first time stamp to a second camera pose 228 associated with the second platform pose 225 for the second time stamp. The calculated gaze direction may be similar to or coincident with the camera direction 134 while in the first camera position 226 at the first time stamp. The changes in camera pose are used in step 204.In step 204, the computer 140 predicts an image conditioned neural radiation field from the previously activated second image 184 using a one-shot technique for neural radiation fields. The image conditioned neural radiation field is used in step 206 to generate a rendered portion of the environment 102. The rendered portion is a three-dimensional estimate of the environment 102 based solely on the second image 184. The predicted image 224 is an estimate of what the camera 130 could have recorded while being viewed in the environment 102 from the first camera pose 226 to the first time stamp. The predicted image 224 may be a field or a frame.In step 208, the first image 182 is read from the buffer and compared to the predicted image 224. The comparison generally aims to minimize the differences in six degrees of freedom 209 (e.g., pitch, roll, yaw, translation in x-direction (tx), translation in y-direction (ty), and translation in vertical z-direction (tz)) between the first image 182 and the predicted image 224. To minimize the differences between the two images, existing methods such as Gaussian-Newton minimization may be used. Other minimization techniques may be implemented to meet the design criteria of a particular application. The six difference values in the six degrees of freedom are used in step 210.In step 210, the computer 140 uses the six difference values in the six degrees of freedom to update the orientation parameters 222 used to estimate the gaze direction in step 202. The updated alignment parameters 222a are stored in a calibration memory within the computer 140 at step 212. A next pass through step 202 accesses the updated orientation parameters 222a from the calibration memory and generates the next gaze direction. In step 214, the updated alignment parameters 222a are presented to the additional circuits 150.Referring to FIG. 3 and back to FIGS. 1 and 2, a perspective diagram 240 of an example rendered portion 242 of the environment 102 is shown in accordance with one or more example embodiments. The diagram 240 illustrates a transformation of the second image 184 into the predicted image 224 as seen in the rendered portion 242 of the environment 102 when viewed along a vector d in the gaze direction 246. The transformation generally learns a neural radiation field from each original pixel (u,v) 248 on a corresponding ray 244 through the second image 184 captured with the second camera pose 228, and predicts each pixel [u,v] T250 on a corresponding vector d through the predicted image 224. The transformation is based on the neural radiation field, which requires the determination of the vector d using a first transfer function f 1 as follows:Here, vector d is gaze direction 246, represents an initial pose of platform (p) relative to environment (e) to first time stamp τ 1 represents a subsequent pose of platform (p) relative to environment (e) to second time stamp τ 2, and represents the orientation parameters of camera (c) relative to platform (p).A second transfer function f 2 may determine a color (C) of the predicted pixel [u,v] T250 in the predicted image 224 (I 2) by integrating 262 the density (σ) and the NeRF color (c) of the points (r(t)) in the ray (r) with distances (t) along the vector d as follows:Here, I 2 is the second image 184. The figure shows an example curve 264 of the density σ, which changes as a function of the beam spacing t.Referring now to FIG. 4 and back to FIGS. 1 and 2, a flowchart of an example implementation of the estimation of the gaze direction of step 202 is shown, in accordance with one or more example embodiments. The flowchart shows an example implementation of the transfer function f 1. Step 202 generally includes steps 272 through 282 as shown. The order of the steps is shown as a representative example. Other step sequences may be implemented to meet the criteria of a particular application.In step 272, the computer 140 receives from the navigation system 120 the platform poses 223, 225 (e.g., ) relative to the environment 102 at the first time stamp and the second time stamp. The computer 140 calculates a relative camera pose from the first time stamp to the second time stamp in step 274 as follows:Wherein the camera's orientation parameters represent on the platform.In step 276, each pixel (u,v) in a second image coordinate is converted as follows:Here, [x,y,1] T is a point on a view beam associated with the pixel (u,v), K -1 is the inverse of the camera's own matrix, and p represents a point in the camera coordinate system of the second camera pose 228.In step 278, a vector a parallel to vector d is calculated as follows:The vector d may be calculated in step 280 as follows:The estimation ends in step 282.Referring to FIG. 5 and back to FIGS. 1 and 2, a flowchart of an example implementation of the one-shot neural radiation field method of step 204 is shown in accordance with one or more example embodiments. The flowchart in FIG. 5 and the flowchart in FIG. 6 show an example implementation of the transfer function f 2. Step 204 generally includes steps 302 through 310 as shown. The order of the steps is shown as a representative example. Other step sequences may be implemented to meet the criteria of a particular application.In step 302, each three-dimensional point r(t) on vector d is scanned as follows:Step 304 may extract a feature W from the second image 184 using a deep neural network. In various embodiments, step 302 and step 304 may be performed in parallel. In some embodiments, step 302 and step 304 may be performed sequentially. In step 306, a feature query W(πr(t)) may be performed. The image conditioned neural radiation field is generated in step 308 as follows:Here, y is a position encoding for r(t). The one-shot neural radiation field method may end in step 310.Referring now to FIG. 6, and back to FIGS. 1 and 2, a flowchart of an example implementation of the volume representation of step 206 is shown, in accordance with one or more example embodiments. The flowchart in FIG. 6 and the flowchart in FIG. 5 show an example implementation of the transfer function f 2. Step 206 generally includes steps 322-326, as shown. The order of the steps is shown as a representative example. Other step sequences may be implemented to meet the criteria of a particular application.In step 322, the computer 140 may perform an initial volume rendering by computing C for each point r(t) in the rendered portion 242 as follows:whereinEach pixel in the predicted image 224 may be concatenated in step 324 to generate the predicted image 224 1 as follows:Volume rendering ends with step 326.The alignment parameters may be optimized in step 208 to achieve minimum differences between the predicted images 224 and the corresponding first images 182. In various embodiments, a loss function may be implemented for minimization. The loss function is generally based on the differences between multiple predicted images 224 and multiple first images 182 in a time window for maturation (e.g., averaging, consensus, and the like). In various embodiments, the pixel comparisons may be implemented as follows:Here, k is the predicted image 224 and I k is the first image 182. In other embodiments, feature comparisons may be performed for optimization. The various feature comparisons may be explicit features including, but not limited to, Harris Corner, Scale Invariant Feature Transform (SIFT), Speeded Up Robust Feature (SURF), Kaze, and SuperPoint. The explicit feature comparisons may be implemented as follows:Here, D is the feature descriptor function for extracting features from the images. In some embodiments, the feature comparisons may be implicit features, such as the Siamese neural network. The implicit feature comparisons can be implemented as follows:Here, S is a neural network model that determines the similarity between two images.In some implementations, the minimization may be based on optimization. The optimization can be based solely on the alignment as follows:In other embodiments, the matching with the localization can be carried out as follows:So that constrEmbodiments of the system include loading navigation data and selected images into the computer. It is checked whether the image features are sufficient and the conditions maneuver degeneration conditions are fulfilled. The vehicle positions and the initial orientation may be used to calculate the gaze direction of the predicted image. A camera position transformation is calculated from the vehicle positions and the initial orientation. For any given pixel in the predicted image, the gaze direction is calculated in the coordinates of the respective image.A one-shot NeRF is used to generate a rendered portion of the environment from the second image. Volume rendering is used to generate a predicted image. Three-dimensional points on view rays through the rendered portion of the environment may be scanned. Features are extracted from the given image. The features are then interrogated on the basis of the scanned three-dimensional points. For each three-dimensional point on each viewing beam, a color and density are produced. Visual rendering is used to integrate the color and density on the same ray to predict the color of a pixel in the predicted image.The alignment parameters may be optimized to achieve minimum differences between the predicted images and the corresponding first images. Loss functions based on differences between the predicted images and the first images may be based on color differences between the individual pixels, two-dimensional position differences between the individual feature pairs, explicit features such as Harris Corner, SIFT, SURF, Kaze, SuperPoint, etc., and implicit features such as Siamese Neural Networks.Maturation within a time window may be performed to smooth the alignment parameter adjustment. Maturation techniques include, but are not limited to, mean, mean, random sample consensus (RANSAC), M estimator sample consensus (MSAC), and similar methods. In some cases, optimization may be performed to minimize the loss function. For a particular vehicle position, the optimization may be limited to the orientation of the camera and platform. The orientation may be optimized based on platform / vehicle position, where vehicle position is constrained by B-spline, graph model, etc. The alignment parameters are updated and a coordinate transformation matrix is created. The updated alignment parameters may be stored in a calibration memory. The updated alignment parameters may also be passed to other circuitry within a vehicle to apply the alignment parameters in the analysis and processing of the video images.Embodiments of the disclosure generally provide a system that includes a platform, a navigation system, a camera, and a computer. The platform is capable of moving through an environment. The navigation system is mounted on the platform and serves to measure a plurality of platform positions of the platform in the environment. The camera is mounted to the platform and may generate multiple images of the environment at different times. The computer is mounted to the platform and may receive a first image from the camera, the first image captured from a first camera pose relative to the environment to a first timestamp of the plurality of timestamps, and receive a second image from the camera, the second image captured from a second camera pose relative to the environment to a second timestamp of the plurality of timestamps, and the second timestamp is different than the first timestamp.The computer is further operable to: estimate a gaze direction of the first camera pose based on a first platform pose of the plurality of platform poses to the first timestamp, a second platform pose of the plurality of platform poses to the second timestamp, and a plurality of orientation parameters of the camera relative to the platform; generate a rendered portion of the environment using a neural radiation field technique based on the second image and the plurality of orientation parameters; generate a predicted image of the environment as observed along the gaze direction by the rendered portion of the environment; determine one or more differences between the first image and the predicted image; and update one or more of the plurality of orientation parameters based on the one or more differences.
Claims
A system (100) comprising: a platform (110) operable to move through an environment; a navigation system (120) mounted on the platform (110) and operable to measure a plurality of platform locations of the platform (110) in the environment; a camera (130) mounted on the platform (110) and operable to generate a plurality of images of the environment at a plurality of time stamps; and a computer (140) mounted on the platform (110) and operable to: receive a first image (182) of the plurality of images from the camera (130), wherein the first image (182) is captured from a first camera location relative to the environment at a first time stamp of the plurality of time stamps; receiving a second image (184) of the plurality of images from the camera (182), wherein the second image (184) is captured from a second camera pose relative to the environment to a second timestamp of the plurality of timestamps, and the second timestamp is different from the first timestamp; estimating a gaze direction of the first camera pose based on a first platform pose of the plurality of platform poses to the first timestamp, a second platform pose of the plurality of platform poses to the second timestamp, and a plurality of orientation parameters of the camera relative to the platform; generating a rendered portion of the environment with a radiation field neural technique based on the second image (184) and the plurality of orientation parameters; generating a predicted image (224) of the environment as viewed along the gaze direction through the rendered portion of the environment; determining one or more differences between the first image (182) and the predicted image (224); and updating one or more of the plurality of alignment parameters based on the one or more differences.The system (100) of claim 1, wherein: the generation of the rendered portion comprises generating an image conditioned neural radiation field solely based on the second image; and the generation of the predicted image on the rendered portion is based on a volume rendering of the image conditioned neural radiation field.The system (100) of claim 1, wherein the computer (140) is further operable to: determine one or more maneuver degeneracy conditions of the platform (110) prior to use of the plurality of platform positions, wherein the estimate of the gaze direction is further based on the one or more maneuver degeneracy conditions.The system (100) of claim 3, wherein the computer (140) is further operable to: determine one or more image release conditions prior to use of the second image (184), wherein the generation of the rendered portion is further based on the one or more image release conditions.The system (100) of claim 1, wherein the computer (140) is further operable to: determine one or more alignment parameters as a reduction in loss function based on a plurality of differences between the first image (182) and the predicted image.The system (100) of claim 5, wherein the loss function is based on one or more of a color difference, a two-dimensional position difference of one or more pairs of features, and a maturation of multiple of the plurality of images within a time window.The system (100) of claim 5, wherein reducing the loss function further comprises aligning a camera coordinate system of the camera (130) with a platform coordinate system of the platform (110).The system (100) of claim 7, wherein the reduction in loss function includes locating the platform over a plurality of the plurality of platform positions.The system (100) of claim 1, wherein the platform is part of one or more of land, water, or aircraft.A method for camera alignment based on a neural radiation field, comprising: measuring a plurality of platform locations of a platform (110) in an environment with a navigation system; generating a plurality of images of the environment to a plurality of time stamps with a camera (130) mounted on the platform; receiving a first image (182) of the plurality of images from the camera at a computer, wherein the first image (182) is captured from a first camera location relative to the environment to a first time stamp of the plurality of time stamps; receiving a second image (184) of the plurality of images from the camera on the computer, wherein the second image is captured from a second camera location relative to the environment to a second time stamp of the plurality of time stamps, and the second time stamp is different than the first time stamp; estimating a gaze direction of the first camera pose based on a first platform pose of the plurality of platform poses to the first timestamp, a second platform pose of the plurality of platform poses to the second timestamp, and a plurality of orientation parameters of the camera relative to the platform; generating a rendered portion of the environment using a neural radiation field technique based on the second image (184) and the plurality of orientation parameters; generating a predicted image of the environment as observed along the gaze direction by the rendered portion of the environment; determining one or more differences between the first image (182) and the predicted image; and updating one or more of the plurality of orientation parameters based on the one or more differences.