INVERTING NEURAL RADIATION FIELDS FOR PART AND SCENE POSE APPRAISAL

DE602024003577T2Active Publication Date: 2026-04-01RTX CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-06-12
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

Existing pose estimation frameworks struggle with complex, non-standard part geometries in smart factories due to the lack of labeled training data and environmental variations, making off-the-shelf networks non-transferrable and unsupervised methods ineffective.

Method used

The method employs Inverted Neural Radiance Fields (iNeRF) and Bundle-Adjusting Neural Radiance Fields (BARF) to perform application-specific pose estimation by refining camera, scene, and object poses using an iterative comparison of observed and rendered images, enabling robust estimation without CAD models and extensive data collection.

Benefits of technology

Enables accurate pose estimation for complex scenes and objects in smart factories by optimizing neural scene representations, reducing data requirements, and adapting to imperfect camera poses, thus facilitating defect mapping to 3D CAD models.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] Neural Radiance Fields (NeRFs) is a technique for generating a 3D representation of an object / scene. NeRF learns this 3D representation of an object / scene from a plurality of 2D images using advanced machine learning. This technique encodes multiple 2D views of an object or scene into an artificial neural network, where, given viewpoint parameters as input, it can predict the light intensity, or radiance, at any point along imaginary rays emitted from the view point position. This allows for the generation of realistic images of the object or scene from different angles and positions. Essentially, from the limited set of multi-view images the NeRF learns how an object would look like if a camera pose were supplied. The iNeRF is actually repeatedly applying NeRF given a camera pose, and then generates renders for that given pose. The iNeRF will continue comparing the render with an actual observation and updating the pose estimate to try to minimize the error. Thus, understanding the pose of a moving camera (6DoF, including translational and rotational motions) in the scene, or the pose of an object part, is essential for better situational awareness. Many existing pose estimation frameworks require large amount of labeled training data, such as through large scale labeling of key points on images for use in supervised training (e.g., deep learning).

[0002] Unfortunately, however, while unsupervised pose estimation exists, its success has been limited to a certain extent, such as fitting a silhouette of CAD model of an object / scene / assembly over the segmented objects and scenes of interest. This is not always feasible when the scene is cluttered and subject to many environmental variations such as illumination, transient objects, and noise. In particularly for smart factory applications, part geometries are highly specialized and typically nonstandard (i.e., not common objects as simple as tables, chairs, etc.) which makes using off-the-shelf pose-estimation pretrained networks difficult (non-transferrable weights, as well as lack of training datasets), if not impossible. Acknowledged are: LIN YEN-CHEN ET AL: "iNeRF: Inverting Neural Radiance Fields for Pose Estimation", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 10 December 2020 (2020-12-10), XP081834080, Yen-Chen Lin: "iNeRF: Inverting Neural Radiance Fields for Pose Estimation", , 8 December 2020 (2020-12-08), XP093214193, Retrieved from the Internet: URL:https: / / www.youtube.com / watch?v=eQuCZaQN0tI&ab_channel=Yen-ChenLin CHEN-HSUAN LIN ET AL: "BARF: Bundle-Adjusting Neural Radiance Fields", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 19 August 2021 (2021-08-19), XP091024204, KARL STELZNER ET AL: "Decomposing 3D Scenes into Objects via Unsupervised Volume Segmentation", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 2 April 2021 (2021-04-02), XP081931740, BRIEF DESCRIPTION

[0003] Disclosed is a method for generating an application specific pose estimation responsive to a camera pose, a scene pose and an object pose according to claim 1.BRIEF DESCRIPTION OF THE DRAWINGS

[0004] The following descriptions should not be considered limiting in any way. With reference to the accompanying drawings, like elements are numbered alike: FIG. 1 is a block diagram illustrating the operation of Inverted Neural Radiance Fields (iNeRF), in accordance with the prior art; and FIG. 2 is an operational block diagram illustrating a method for generating an application specific pose estimation responsive to a camera pose, a scene pose and an object pose, in accordance with an embodiment of the invention; DETAILED DESCRIPTION

[0005] A detailed description of one or more embodiments of the disclosed apparatus and method are presented herein by way of illustration and not limitation with reference to the Figures.

[0006] In an embodiment, NeRFs may be used for synthesizing novel views of a scene or object given desired / queried camera pose parameters. NeRFs may be trained by capturing multiple images from different viewpoints (e.g., multi-view), in which an artificial neural network learns to synthesize how an object or scene would look like given an arbitrary camera pose completely based on the captured 2D images. By inverting the NeRFs in an iNeRF framework, pose estimation may be performed via an initial pose estimate, which then generates a synthesized view which is compared to the current observation. The error may then be backpropagated to update and refine the pose estimation until convergence of the pose estimation and the observed image is achieved.

[0007] In at least one embodiment, the invention contemplates applying Bundle-Adjusting Neural Radiance Fields (BARF) approach for training NeRF from an imperfect (or even unknown) camera pose, where BARF can effectively optimize the neural scene representations and, at the same time, resolve large camera pose misalignment. This may enable view synthesis and localization of image sequences from unknown camera poses and may help with having imperfect ground truth and limited data, so the data requirements don't have to be that high which may make it useful because there is no need to collect data and / or having to get an accurate camera pose to train the NeRF. It should be appreciated that the method of the invention disclosed herein allows for application specific functionality for complex parts, scenes, etc. especially in "smart" factories by using iNeRF for camera pose, scene pose and object pose estimation, as opposed to just object pose estimation. Moreover, the incorporation of BARF as the initial, or forward, NeRF instead of the trained "vanilla" NeRF to address imperfect ground truths may result in a robust method for pose estimation based on real visual images which require no CAD models and no need to consider domain shift.

[0008] It should be appreciated that the invention is applicable to a multiple branch approach (i.e., multiple NeRFs, segmentation, multiple estimates).

[0009] Referring to FIG.1 and FIG. 2, a method 100 for generating an application specific pose estimation responsive to a camera pose estimation, object pose estimation and scene pose estimation for a camera, a scene and at least one object is provided, in accordance with an embodiment. The method 100 includes generating an observed image (I obs ) of the scene and an object, as shown in operational block 102. This may be accomplished using a camera which captures an image of the object / scene from one or more angles. In one embodiment, the camera may be mounted on a robotic arm which moves around a stationary object and takes an image of the object at a predetermined inspection angle. For example, if a blade of a turbo fan is being inspected, the camera may take an image of the blade from a specific inspection angle. It should be appreciated that in another embodiment, the camera may be stationary and the object may be moving. The observed image is processed to generate a scene image and an object image, as shown in operational block 104. This may be accomplished by segmenting the observed image using a 2D segmentation approach to isolate and obtain the current scene (i.e., foreground / background) pose and the current object pose from the observed image.

[0010] The current scene pose and the object scene pose are separately fed into specialized NeRFs to generate a scene image rendering and an object image rendering, respectively, as shown in operational block 106. The scene image rendering and the object image rendering are composited to generate a rendered output (I rndr ) as shown in operational block 108. The final pose estimate is determined by comparing the observed image (I obs ) with the rendered output (I rndr ) to determine the loss metric (i.e., the difference between the observed image (I obs ) and the rendered output (I rndr )) and repeatedly backpropagating the loss metric to update the current pose estimate until the observed image (I obs ) and the rendered output (I rndr ) converge (i.e., match), as shown in operational block 110. It should be appreciated that the loss metric is a metric that corresponds to a difference between the observed image (I obs ) and the final rendered output (I rndr ). In an embodiment, the invention may include a predetermined loss metric threshold, where if the loss metric is below the predetermined loss metric threshold, then the method may include backpropagating the loss metric to update the current pose estimate and repeating the backpropagating until the loss metric is equal to or greater than the predetermined loss metric threshold.

[0011] The rendered output (I rndr ) is the final pose estimate when the rendered output (I rndr ) matches the observed image (I obs ). This final pose estimate may then be used to perform application specific functions, such as mapping a defect seen in the image space to a 3D CAD model. It should be appreciated that the method of the invention may be used with other ranges of the electromagnetic spectrum, such as infrared.

[0012] It should be appreciated that, although the invention is described hereinabove with regards to only one object, it is contemplated that in other embodiments the invention may be used for multiple objects. Moreover, although the invention is described hereinabove with regards to the camera being movable and the object being stationary (i.e., static), it is contemplated that in other embodiments the camera may be stationary and the object may be movable. Moreover, the invention uses "multiple branches" and "multiple NeRFs" to achieve better rendering of scene (i.e., foreground / background, multiple objects, different camera pose vs object pose in a scene, etc.) such that multiple estimates are obtained. The invention may be used for application specific tasks involving complex parts, scenes, etc. especially in smart factories.

[0013] The term "about" is intended to include the degree of error associated with measurement of the particular quantity based upon the equipment available at the time of filing the application.

[0014] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, element components, and / or groups thereof.

[0015] In accordance with one or more embodiments, the processing of at least a portion of the method in FIG. 2 may be implemented by a controller / processor disposed internal and / or external to a computing device. In addition, processing of at least a portion of the method in FIG. 2 may be implemented through a controller / processor operating in response to a computer program. In order to perform the prescribed functions and desired processing, as well as the computations therefore (e.g. execution control algorithm(s), the control processes prescribed herein, and the like), the controller may include, but not be limited to, a processor(s), computer(s), memory, storage, register(s), timing, interrupt(s), communication interface(s), and input / output signal interface(s), as well as combination comprising at least one of the foregoing.

[0016] Additionally, the invention may be embodied in the form of a computer or controller implemented processes. The invention may also be embodied in the form of computer program code containing instructions embodied in tangible media, such as floppy diskettes, CD-ROMs, hard drives, and / or any other computer-readable medium, wherein when the computer program code is loaded into and executed by a computer or controller, the computer or controller becomes an apparatus for practicing the invention. The invention can also be embodied in the form of computer program code, for example, whether stored in a storage medium, loaded into and / or executed by a computer or controller, or transmitted over some transmission medium, such as over electrical wiring or cabling, through fiber optics, or via electromagnetic radiation, wherein when the computer program code is loaded into and executed by a computer or a controller, the computer or controller becomes an apparatus for practicing the invention. The computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing.

[0017] A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer-readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire When implemented on a general-purpose microprocessor the computer program code segments may configure the microprocessor to create specific logic circuits.

[0018] Additionally, the processor may be part of a computing system that is configured to or adaptable to implement machine learning models which may include artificial neural networks, such as deep neural networks, convolutional neural networks, recurrent neural networks, vision transformers, encoders, decoders, or any other type of machine learning model. The machine learning models can be trained in a supervised, unsupervised, or hybrid manner. While the present disclosure has been described with reference to an exemplary embodiment or embodiments, it will be understood by those skilled in the art that various changes may be made and equivalents may be substituted for elements thereof without departing from the scope of the invention as defined by the claims. Moreover, the embodiments or parts of the embodiments may be combined in whole or in part without departing from the scope of the invention according to the claims. Therefore, it is intended that the present disclosure not be limited to the particular embodiment disclosed as the best mode contemplated for carrying out this disclosure, but that the present disclosure will include all embodiments falling within the scope of the claims.

Claims

1. A method (100) for generating an application specific pose estimation responsive to a camera pose, a scene pose and an object pose, the method comprising: generating (102) an observed image (Iobs) of an object and a scene using a camera; processing the observed image to generate a current pose estimate having a current object pose and a current scene pose, wherein the processing comprises segmenting (104) the observed image using a 2D segmentation approach to isolate and obtain the current scene pose and the current object pose from the observed image; processing the current pose estimate to: generate (106) an object render by applying the object image to a first specialized NeRF, and generate (106) a scene render by applying the scene image to a second specialized NeRF; generating (108) a final rendered output (Irndr) by compositing the object render and the scene render; generating (110) a final pose estimate, wherein generating a final pose estimate includes: comparing the observed image (Iobs) with the final rendered output (Irndr) to obtain a loss metric, backpropagating the loss metric to update the current pose estimate, and repeating the generating a final pose estimate until the final rendered output (Irndr) matches the observed image (Iobs).

2. The method of claim 1, wherein generating (102) an object image includes operating a moving camera to generate an image of a stationary object from at least one predefined angle relative to a position of the object.

3. The method of claim 1 or 2, wherein generating (102) an object image includes operating a stationary camera to generate an image of a moving object from at least one predefined angle relative to a position of the object.

4. The method of any preceding claim, wherein generating (108) a final rendered output (Irndr) includes compositing the object render and the scene render together using a predetermined compositing technique.

5. The method of any preceding claim, wherein: generating (106) an object render includes applying the object image to a BARF, and generating (106) a scene render by applying the scene image to a BARF.