Daily scene restoration engine

The scene reconstruction engine uses imaging polarimeters and octree data structures to efficiently reconstruct everyday scenes, addressing limitations in light field reconstruction and computational complexity, achieving high-resolution models comparable to human perception.

JP2026016640APending Publication Date: 2026-02-03QUIDIENT LLC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025182658
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-02-08
Filing Date
2025-10-29
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Conventional 3D imaging systems struggle to efficiently reconstruct everyday scenes, which are densely volumetric, contain occluded, shiny, and partially transparent objects, and have complex light fields, due to limitations in light field reconstruction and computational complexity.

Method used

A scene reconstruction engine that utilizes imaging polarimeters to capture polarization characteristics of light, forming a volumetric scene model with spatially sorted hierarchical data structures like octrees, enabling efficient reconstruction of everyday scenes with angular resolution comparable to the human eye.

Benefits of technology

The engine achieves accurate and efficient reconstruction of everyday scenes, overcoming limitations of existing technologies by incorporating polarization data and spatial processing techniques, resulting in models with resolution suitable for human perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026016640000001_ABST
    Figure 2026016640000001_ABST
Patent Text Reader

Abstract

To provide an everyday scene reconstruction engine.SOLUTION: The stored volumetric scene model of the real scene is generated from data defining a digital image of a light field in the real scene including different types of media. The digital images are formed by the camera from opposite facing poses, and each digital image includes image data elements defined by stored data representative of light field flux received by photosensitive detectors in the camera. The digital images are processed by a scene reconstruction engine to form a digital volumetric scene model representing the real scene. The volumetric scene model includes (i) volumetric data elements defined by stored data representing one or more media properties and (ii) solid angle data elements defined by stored data representing a bundle of light fields. Adjacent volumetric data elements form a corridor.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 321,564, filed April 12, 2016, U.S. Provisional Patent Application No. 62 / 352,379, filed June 20, 2016, U.S. Provisional Patent Application No. 62 / 371,494, filed August 5, 2016, U.S. Provisional Patent Application No. 62 / 427,603, filed November 29, 2016, U.S. Provisional Patent Application No. 62 / 430,804, filed December 6, 2016, U.S. Provisional Patent Application No. 62 / 456,397, filed February 8, 2017, and U.S. Provisional Patent Application No. 62 / 420,797, filed November 11, 2016, which are incorporated herein by reference in their entireties.

[0002] The present invention relates generally to the field of 3D imaging, and more particularly to volumetric scene reconstruction. [Background technology]

[0003] A 3D image is a digital 3D model of a real-world scene captured for various purposes, including visualization and information extraction. They are acquired by 3D imagers, variously called 3D sensors, 3D cameras, 3D scanners, VR cameras, 360° cameras, and depth cameras. They serve the 3D information needs of applications used in global sectors, including defense, security, entertainment, education, healthcare, infrastructure, manufacturing, and mobile.

[0004] Numerous methods have been developed to extract 3D information from a scene. Many involve active light sources such as lasers and have limitations such as high power consumption and limited range. A near-ideal method uses two or more images from inexpensive cameras (devices that form images by sensing a light field using detectors) to generate a detailed scene model. The term Multi-View Stereo (MVS) is used herein, but this and its variants are also known by other names, including photogrammetry, Structure-from-Motion (SfM), and Simultaneously Localization and Mapping (SLAM), among others. Many such methods are described in Furukawa's reference, "Multi-View Stereo: A Tutorial," which frames MVS within the framework of an image / geometry consistency optimization problem. Robust implementation of photometric consistency and efficient optimization algorithms has proven important for achieving successful algorithms.

[0005] To improve the robustness of scene model extraction from images, improved modeling of light transport is needed, which includes the properties of light's interaction with objects, including transmission, reflection, refraction, scattering, etc. Jarosz's paper "Efficient Monte Carlo Methods for Light Transport in Scattering Media" (2008) provides a thorough analysis of this subject.

[0006] In the simplest version of MVS, if the camera viewpoint and pose are known for two images, the location of a "landmark" 3D point in the scene can be computed if that point's projection is found in the two images (its 2D "feature" point) using some form of triangulation. (A feature is a property of an entity expressed in terms of a description and a pose. Examples of features include spots, glints, or buildings. A description can be i) used to find instances of the feature at several poses in a field (a space in which an entity can be posed), or ii) formed from descriptive properties at one pose in a field.) The surface is extracted by combining many landmarks. This works as long as the feature points are indeed correct projections of fixed landmark points in the scene and are not caused by any viewpoint-dependent artifacts (e.g., specular reflections, edge crossings). This can be extended to many images and situations where the camera viewpoint and pose are unknown. The process of resolving landmark placement and camera parameters is called bundle adjustment (BA), although there are many variations and other names used for specific applications and situations. This topic is comprehensively covered by Triggs in his paper "Bundle Adjustment - A Modern Synthesis" (2009). An important subtopic in BA is the ability to compute solutions without explicitly generating analytical derivatives, which becomes increasingly computationally difficult as the situation becomes more complex. This is introduced by Brent in his book "Algorithms for Minimization Without Derivatives".

[0007] While the two properties of light, color and intensity, are used in MVS, they have significant limitations when used in everyday scenes. These include an inability to accurately represent surfaces without texture, non-Lambertian objects, and transparent objects within a scene (an object is a medium that is expected to be juxtaposed. Examples of objects include leaves, twigs, trees, fog, clouds, and the Earth). To address this, a third property of light, polarization, has been found to extend scene reconstruction capabilities. The use of polarimetric imaging in MVS is called Shape from Polarization (SfP). Wolff, U.S. Pat. No. 5,028,138, discloses a basic SfP device and method based on specular reflection. Diffuse reflection, if present, is assumed to be unpolarized. Barbour, U.S. Pat. No. 5,890,095, discloses a polarimetric imaging sensor device and a micropolarizer array. Barbour, U.S. Pat. No. 6,810,141, discloses a general method for using SPI sensors to provide information about an object, including information about its 3D geometry. d'Angelo, German Patent No. 102004,062,461, discloses an apparatus and method for determining geometry based on shape from shading (SfS) combined with SfP. d'Angelo, German Patent No. 102006,013,318, discloses an apparatus and method for determining geometry based on SfS combined with SfP and a block-matching stereo algorithm for adding range data to sparse point sets. Morel, International Publication No. 2007,057,578, discloses an apparatus for SfP of highly reflective objects.

[0008] Koshikawa's paper, "A Model-Based Recognition of Glossy Objects Using Their Polarimetrical Properties," is generally considered the first to disclose the use of polarization information to determine the shape of dielectric glossy objects. Later, Wolff presented a basic polarization camera design in his paper, "Polarization camera for computer vision with a beam splitter." Miyazaki's paper, "Determining shapes of transparent objects from two polarization images," developed the SfP method for transparent or reflective dielectric surfaces. Atkinson's paper, "Shape from Diffuse Polarization," explains the basic physics of surface scattering and describes equations for determining shape from polarization in the cases of diffuse and specular reflection. Morel's paper, "Active Lighting Applied to Shape from Polarization," describes an SfP system for reflective metallic surfaces using integrating dome and active lighting. It explains the basic physics of surface scattering and describes equations for determining shape from polarization in the cases of diffuse and specular reflection. In his paper "3D Reconstruction by Integration of Photometric and Geometric Methods," d'Angelo describes an approach to 3D reconstruction based on sparse point clouds and dense depth maps.

[0009] MVS systems are largely based on resolving surface parameters, but improvements have been found by increasing the dimensionality of the modeling using dense methods. Newcombe describes an advancement in this in his paper "Live Dense Reconstruction with a Single Moving Camera" (2010). Wurm describes another method in his paper "OctoMap: A Probabilistic, Flexible, and Compact 3D Map Representation for Robotic Systems" (2010).

[0010] When applying MVS methods to real-world scenes, the computational complexity required can quickly become impractical for many applications, especially for mobile and low-power operation. In fields outside of MVS, such as medical image processing, where such computational problems have previously been solved, the use of octree and quadtree data structures and methods has proven effective. This is particularly the case when implemented on modest, dedicated processors. This technique promises to enable the use of a wide range of simple, inexpensive, low-power processors to be applied to computationally challenging situations. The basic octree concept was introduced by Meagher in his papers "Geometric Modeling Using Octree Encoding" and "The Octree Encoding Method for Efficient Solid Modeling." It was later extended to orthographic image generation in U.S. Patent No. 4,694,404. [Prior art documents] [Patent documents]

[0011] [Patent Document 1] U.S. Patent No. 5,028,138 [Patent Document 2] U.S. Patent No. 6,810,141 [Patent Document 3] German Patent No. 102004062461

Patent document 4

Patent document 5

Patent document 6

Patent document 7

Non-licensed literature

[0012] [Non-licensed document 1] Furukawa, reference "Multi-View Stereo: A Tutorial" [Non-licensed document 2] Jarosz, paper "Efficient Monte Carlo Methods for Light Transport in Scattering Media" (2008) [Non-licensed document 3] Triggs, paper "Bundle Adjustment - A Modern Synthesis" (2009)

Non-licensed Document 4

Non-licensed Document 5

Non-licensed Document 6

Non-licensed Document 7

[0013] The following simplified summary is intended to provide a first basic understanding of some aspects of the systems and / or methods described herein. This summary is not an extensive overview of the systems and / or methods described herein. It is not intended to identify all key / critical elements or to delineate the entire scope of such systems and / or methods. Its sole purpose is to present some concepts in a simplified form as a prelude to the more detailed description that is presented later.

[0014] In some embodiments, data defining a digital image is acquired while the camera is oriented in different directions (e.g., opposite orientations) within a workspace of a natural scene (e.g., including different types of media through which light (e.g., propagating electromagnetic energy within and / or outside the visible frequency / wavelength spectrum) is transported or reflected) and input directly into the scene reconstruction engine. In other embodiments, such digital image data has already been acquired and stored for later access by the scene reconstruction engine. In either case, the digital image includes image data elements defined by stored data representing the light field flux received by the photosensitive detectors within the camera.

[0015] An exemplary scene reconstruction engine processes such digital images to form a digital volumetric scene model representing the real scene. The scene model may include volumetric data elements defined by stored data representing one or more media characteristics. The scene model may also include solid angle data elements defined by stored data representing sensed light field flux. Adjacent volumetric data elements form corridors in the scene model, and at least one of the volumetric data elements in at least one corridor represents a partially optically transparent medium.

[0016] The volumetric scene model data is stored in a digital data memory for a desired subsequent application (e.g., providing design data for applications involving real scenes, displaying images perceptible to humans, as well as many other applications known to those skilled in the art).

[0017] In some exemplary embodiments, the acquired image data elements represent at least one predetermined polarization characteristic of the imaged real scene light field.

[0018] In some exemplary embodiments, some corridors may occupy an area between (i) the position of the camera's photosensitive detector when the associated digital image was acquired and (ii) a volumetric data element representing a media element comprising a reflective surface element.

[0019] In some exemplary embodiments, some corridors may include volumetric data elements located at the distal ends of the corridors that represent featureless reflective media surfaces when observed through unpolarized light.

[0020] In some exemplary embodiments, some corridors may include volumetric data elements located at the distal ends of the corridors that represent localization orientation gradients on featureless solid media surfaces when observed through unpolarized light (e.g., recessed "bumps" on a vehicle exterior caused by hail damage).

[0021] In some exemplary embodiments, the scene restoration engine receives a user identification of at least one scene restoration goal, and the scene restoration process continues iteratively until the identified goal is achieved with a predetermined accuracy. Such an iterative process may continue, for example, until the angular resolution of at least some volumetric data elements in the scene model is comparable to or better than the angular resolution of the average human eye (e.g., about 1 arc minute or about 0.0003 radians).

[0022] In some exemplary embodiments, digital images of at least some media in the real scene are acquired at different distances between the camera and the media of interest (e.g., close-up images of smaller media elements such as flowers, leaves on trees, etc.).

[0023] In some exemplary embodiments, the volumetric data elements are stored in digital memory in a spatially sorted hierarchical manner to facilitate scene reconstruction processing and / or subsequent application use of the reconstruction model data.

[0024] In some exemplary embodiments, the solid angle data elements are stored spatially in a solid angle octree (SAO) format to facilitate scene reconstruction processing and / or subsequent application use of the reconstruction model data.

[0025] In some exemplary embodiments, some volumetric data elements include a representation of a single-center multi-directional light field.

[0026] In some exemplary embodiments, the volumetric data elements at the distal ends of at least some of the corridors are associated with stereoscopic data elements centered at the center of the volumetric scene model.

[0027] In some exemplary embodiments, at least some of the intermediately located volumetric data elements within the corridor are associated with a multi-directional light field of the stereoscopic data elements.

[0028] In some exemplary embodiments, the scene restoration process uses a non-differential optimization method to calculate the minimum of a cost function used to identify interest points in the acquired digital images.

[0029] In some exemplary embodiments, the scene reconstruction process includes refining the digital volumetric scene model by iteratively performing the steps of: (A) comparing (i) projected digital images of the previously constructed digital volumetric scene model with (ii) corresponding generated digital images of the real scene; and (B) modifying the previously constructed digital volumetric scene model to more accurately fit the compared digital images of the real scene, thereby generating a newly constructed digital volumetric scene model.

[0030] In some exemplary embodiments, the camera may be embedded in a user-carried portable camera (e.g., in an eyeglass frame) to acquire digital images of the real scene. In such embodiments, the acquired digital image data may be transmitted to a remote data processor in which at least a portion of the scene reconstruction engine is located.

[0031] An exemplary embodiment includes at least (A) a machine apparatus in which the claimed functions are arranged, (B) performing method steps to implement an improved scene restoration process, and / or (C) a non-volatile computer program storage medium containing executable program instructions that, when executed on a compatible digital processor, implement the improved scene restoration process and / or generate an improved scene restoration engine machine.

[0032] Additional features and advantages of the exemplary embodiments are described below.

[0033] These and other features and advantages will be better and more completely understood by reference to the following detailed description of exemplary, non-limiting, illustrated embodiments taken in conjunction with the following drawings. [Brief explanation of the drawings]

[0034] [Figure 1]FIG. 1 illustrates a scene that may be scanned and modeled, according to some example embodiments. [Figure 2A] FIG. 1 is a block diagram of a 3D image processing system according to some exemplary embodiments. [Figure 2B] FIG. 1 is a block diagram of a scene reconstruction engine (a device configured to use images to reconstruct a volumetric model of a scene) hardware according to some example embodiments. [Figure 3] 1 is a flowchart of a process for capturing images of a scene and forming a volumetric model of the scene, according to some example embodiments. [Figure 4A] 1 is a geometric diagram illustrating a volume field and volume elements (voxels), according to some exemplary embodiments; [Figure 4B] FIG. 1 is a geometric diagram illustrating a solid angle field and solid angle elements (saels, pronounced "sails"), according to some example embodiments. [Figure 5] FIG. 2 is a geometric diagram illustrating a workspace-scale plan view of the kitchen shown in FIG. 1, according to some exemplary embodiments. [Figure 6] FIG. 1 is a 3D diagram illustrating a voxel-scale view of a small portion of a kitchen window, in accordance with some exemplary embodiments. [Figure 7A] 1 is a geometric diagram illustrating a plan view of a model of a kitchen scene, according to some example embodiments. [Figure 7B] 1 is a geometric diagram showing a plan view of a model of a kitchen scene, with the kitchen shown as a small dot at this scale, according to some example embodiments. [Figure 8A] 1 is a geometric diagram illustrating camera poses used to reconstruct voxels, according to some exemplary embodiments. [Figure 8B]1 is a geometric diagram illustrating various materials that may occupy a voxel, according to some exemplary embodiments. [Figure 8C] 1 is a geometric diagram illustrating various materials that may occupy a voxel, according to some exemplary embodiments. [Figure 8D] 1 is a geometric diagram illustrating various materials that may occupy a voxel, according to some exemplary embodiments. [Figure 8E] 1 is a geometric diagram illustrating various materials that may occupy a voxel, according to some exemplary embodiments. [Figure 9A] FIG. 1 is a geometric diagram illustrating a single-center unidirectional sail arrangement that may be used, for example, to represent a light field or sensor frustum, according to some exemplary embodiments. [Figure 9B] FIG. 1 is a geometric diagram illustrating a single-center multidirectional sail arrangement that may be used, for example, to represent a light field or sensor frustum, according to some exemplary embodiments. [Figure 9C] FIG. 1 is a geometric diagram illustrating a single-center omnidirectional sail arrangement that may be used, for example, to represent a light field or sensor frustum. [Figure 9D] FIG. 1 is a geometric diagram illustrating a single-center isotropic sail arrangement that may be used, for example, to represent a light field or sensor frustum, in accordance with some exemplary embodiments. [Figure 9E] FIG. 1 is a geometric diagram illustrating a planar-centered unidirectional sail arrangement that may be used, for example, to represent a light field or sensor frustum, according to some exemplary embodiments. [Figure 9F] FIG. 1 is a geometric diagram illustrating a multi-center omnidirectional sail arrangement that may be used, for example, to represent a light field or sensor frustum, according to some exemplary embodiments. [Figure 10]FIG. 1 is an isometric diagram illustrating a bidirectional light interaction function (BLIF) relating an incident (going from outside a closed boundary to inside the closed boundary), a responsive light field, an emissive light field, and an exitant (going from inside a closed boundary to outside the closed boundary) light field, according to some example embodiments. [Figure 11] FIG. 1 is a schematic diagram illustrating scene restoration engine (SRE) functionality, according to some example embodiments. [Figure 12] FIG. 1 is a schematic diagram illustrating data modeling, according to some example embodiments. [Figure 13] FIG. 1 is a schematic diagram illustrating light field physics capabilities, according to some example embodiments. [Figure 14] FIG. 2 is a functional flow diagram illustrating a software application, according to some example embodiments. [Figure 15] FIG. 1 is a functional flow diagram illustrating plan processing, according to some example embodiments. [Figure 16] FIG. 4 is a functional flow diagram illustrating a scanning process, according to some example embodiments. [Figure 17] FIG. 1 is a schematic diagram illustrating sensor control functions, according to some exemplary embodiments. [Figure 18A] FIG. 1 is a functional flow diagram illustrating scene resolution, according to some example embodiments. [Figure 18B] FIG. 2 is a functional flow diagram illustrating a process used to update a posted scene model, according to some example embodiments. [Figure 18C] FIG. 2 is a functional flow diagram illustrating a process used to restore the petals of the daffodils in the scene of FIG. 1, according to some example embodiments. [Figure 18D]FIG. 18D is a functional flow diagram illustrating the process used to directly solve one mediel of the daffodil petal reconstructed in FIG. 18C for BLIF, according to some example embodiments. [Figure 19] FIG. 1 is a schematic diagram illustrating spatial processing operations, according to some example embodiments. [Figure 20] FIG. 1 is a schematic diagram illustrating light field operations, according to some example embodiments. [Figure 21] FIG. 1 is a geometric diagram illustrating a solid angle octree (SAO) (in 2D), according to some example embodiments. [Figure 22] FIG. 1 is a geometric diagram illustrating a solid angle octree subdivision along one arc, according to some example embodiments. [Figure 23A] FIG. 1 is a geometric diagram illustrating feature correspondences of surface normal vectors. [Figure 23B] FIG. 1 is a geometric diagram illustrating the projection of landmark points onto two image feature points for registration, according to some exemplary embodiments. [Figure 24] 1 is a geometric diagram illustrating registration cost as a function of parameters, according to some exemplary embodiments; [Figure 25A] FIG. 10 is a schematic diagram illustrating generation of updated cost values ​​after a PUSH operation, according to some example embodiments. [Figure 25B] FIG. 10 is a functional flow diagram illustrating an alignment procedure, according to some exemplary embodiments. [Figure 26] FIG. 10 is a geometric diagram illustrating movement from node centers to minimum points for alignment, according to some exemplary embodiments. [Figure 27] FIG. 10 is a geometric diagram illustrating the movement of minimum points in a projection for registration, according to some example embodiments. [Figure 28A] FIG. 10 is a geometric diagram illustrating the front face of an SAO bounding cube, according to some example embodiments. [Figure 28B]FIG. 10 is a flow diagram illustrating a process for generating a position-invariant SAO, according to some exemplary embodiments. [Figure 29] FIG. 10 is a geometric diagram illustrating the projection of an SAO onto the surface of a quadtree, according to some example embodiments. [Figure 30] FIG. 1 is a geometric diagram illustrating the projection of an octree in incident SAO generation, according to some example embodiments. [Figure 31A] FIG. 2 is a geometric diagram illustrating the relationship between an outgoing SAO and an incoming SAO, according to some example embodiments. [Figure 31B] FIG. 31B is a geometric diagram illustrating the resulting incident SAO from FIG. 31A, in accordance with some exemplary embodiments. [Figure 32] FIG. 1 is a geometric diagram illustrating a bidirectional light interaction function (BLIF) SAO in 2D, according to some exemplary embodiments. [Figure 33] FIG. 1 is a conceptual diagram illustrating a hail damage assessment application, according to some illustrative embodiments. [Figure 34] FIG. 1 illustrates a flow diagram of an application process for a hail damage assessment (HDA), according to some example embodiments. [Figure 35] FIG. 1 is a flow diagram illustrating an HDA inspection process, according to some example embodiments. [Figure 36A] A geometric diagram showing an external view of the coordinate system of a display (a device that stimulates human senses to create concepts of entities such as objects and scenes), according to some exemplary embodiments. [Figure 36B] FIG. 1 is a geometric diagram illustrating an internal view of a display coordinate system, according to some exemplary embodiments. [Figure 36C] 1 is a geometric diagram illustrating an orthogonal view in a display coordinate system, according to some exemplary embodiments. [Figure 37] FIG. 10 is a geometric diagram illustrating the geometric movement of node centers from a parent node to a child node in the X direction, according to some exemplary embodiments. [Figure 38]FIG. 1 is a schematic diagram illustrating an implementation of an octree geometric transformation in the X dimension, according to some exemplary embodiments. [Figure 39] FIG. 2 is a geometric diagram illustrating a perspective projection of octree node centers onto a display screen, according to some exemplary embodiments. [Figure 40A] FIG. 1 is a geometric diagram illustrating spans and windows in a perspective projection, according to some exemplary embodiments. [Figure 40B] FIG. 10 is a geometric perspective view illustrating origin rays and centers of spans according to some exemplary embodiments. [Figure 40C] FIG. 1 is a geometric diagram showing span zones when subdividing a node. [Figure 40D] FIG. 10 is a geometric diagram showing spans and origins after subdividing a node. [Figure 41A] FIG. 10 is a schematic diagram illustrating the calculation of span and node center values ​​resulting from a PUSH operation, according to some example embodiments. [Figure 41B] This is a continuation of Figure 41A. [Figure 42] 1 is a geometric diagram illustrating a measure of scene model accuracy (a measure of deviation between elements in a scene model and corresponding elements in the real scene that the scene model represents), according to some example embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0035] The following terms are provided with non-limiting examples for ease of reference. Corridor. A passageway through a permeable medium. Frontier. The incident light field at the boundary of the scene model. Imaging polarimeter: A camera that senses polarimetric images. Light. Electromagnetic waves with frequencies including visible, infrared, and ultraviolet light. Light field: The flow of light in a scene. Medium. A volumetric region that contains some material through which light flows, or no material through which light flows. A medium can be homogeneous or heterogeneous. Examples of homogeneous media include open space, air, and water. Examples of heterogeneous media include the surface of a mirror (part air and part silvered glass), the surface of a sheet of glass (part air and part transparent glass), and a volumetric region containing a pine tree branch (part air and part organic material). Light flows through a medium due to phenomena including absorption, reflection, transmission, and scattering. Examples of partially transparent media include a pine tree branch and a sheet of glass. Workspace: The area of ​​a scene that contains the camera positions from which images of the scene are captured.

[0036] A scene reconstruction engine (SRE) is a digital device that uses images to reconstruct 3D models of real scenes. Currently available SREs are virtually incapable of reconstructing an important type of scene known as an everyday scene (where "everyday" means "everyday"). Generally, everyday scenes are i) densely volumetric, ii) contain occluded, shiny, partially transparent, and featureless objects (e.g., white walls), and iii) contain complex light fields. The majority of scenes people occupy in the course of their daily lives are mundane. One reason that today's SREs cannot effectively reconstruct everyday scenes is that they do not reconstruct light fields. To successfully reconstruct shiny, partially transparent objects such as cars and bottles, the light field in which the objects are embedded must be known.

[0037] Exemplary embodiments include a new type of SRE that provides capabilities such as, for example, efficient reconstruction of everyday scenes to reach useful levels of scene model accuracy (SMA) and their use. According to some embodiments, a scene reconstruction engine uses images to create a volumetric scene model of a real scene. The volumetric elements of the scene model represent corresponding regions of the real scene in terms of i) the type and characteristics of the media occupying the volumetric elements and ii) the type and characteristics of the light field present at the volumetric elements. In embodiments, a novel scene resolution method is used for efficiency. To ensure that reading and writing from highly detailed volumetric scene models is efficient during and after their creation, and to achieve models that can be used by many applications requiring accurate volumetric models of a scene, novel spatial processing techniques are used in embodiments. One type of camera that has significant advantages for everyday scene reconstruction is the imaging polarimeter.

[0038] A 3D imaging system includes a scene reconstruction engine and a camera. There are three broad 3D imaging technologies currently available: time-of-flight, multi-viewpoint, and light-field focal plane. Each technology is unable to effectively reconstruct everyday scenes in at least one way. Time-of-flight cameras are virtually unable to reconstruct everyday scenes at longer distances. A key deficiency of multi-viewpoint cameras is their inability to determine shape within featureless surface regions, leaving gaps in the model. Light-field focal plane cameras have very low depth resolution because the stereo baseline is very small (the width of the imaging chip). These deficiencies remain a major price-performance bottleneck, relegating these technologies to narrow, specialized applications.

[0039] Scene 101 shown in Figure 1 is an example of a real-world scene that is also everyday. Scene 101 includes a kitchen area, including region 103 containing a portion of a tree, region 105 containing a portion of a cabinet, region 107 containing a portion of a hood, region 109 containing a portion of a window, region 111 containing a portion of a mountain, region 113 containing a flower in a jar, and region 115 containing a portion of the sky. Scene 101 is an everyday scene because it contains non-Lambertian reflecting objects, including partially transparent and fully transparent objects. Scene 101 also extends through a window to the frontier.

[0040] Most of the scenes people occupy in the course of their daily lives are everyday. Yet, conventional scene reconstruction engines are unable to cost-effectively reconstruct everyday scenes to the accuracy required by most potential 3D image processing applications. Exemplary embodiments of the present invention implement a scene reconstruction engine that can efficiently and accurately model various types of scenes, including everyday ones. However, the exemplary embodiments are not limited to everyday scenes and may also be used to model real scenes.

[0041] Preferably, the scene reconstruction engine generates a scene model with, at least in some parts, an angular resolution comparable to that discernible by the average human eye. The human eye has an angular resolution of approximately 1 arcminute (0.02°, or 0.0003 radians, which corresponds to 0.3 m at a distance of 1 km). Over the past few years, image display technologies have begun to realize the ability to create projected light fields that are nearly indistinguishable from the light fields in the real scene they represent (as perceived by the naked eye). The inverse problem of reconstructing a volumetric scene model from images with sufficient resolution to achieve similar accuracy as perceived by the naked eye has not yet been effectively solved by other 3D image processing techniques. In contrast, scene models reconstructed by exemplary embodiments approach and potentially exceed the accuracy thresholds mentioned above.

[0042] 2A is a block diagram of a 3D image processing system 200 according to some exemplary embodiments. The 3D image processing system 200 may be used, for example, to scan a real scene, such as scene 101, and form a corresponding volumetric scene model. The 3D image processing system 200 includes a scene reconstruction engine (SRE) 201, a camera 203, application software 205, a data communication layer 207, and a database 209. The application software 205 includes a user interface module 217 and a job script processing module 219. The 3D image processing system 200 may be embodied in mobile and wearable devices (such as glasses).

[0043] SRE 201 contains instruction logic and / or circuitry for performing various image capture and image processing operations, as well as overall control of 3D image processing system 200 for generating volumetric scene models. SRE 201 is further described below with respect to FIG. 11.

[0044] Camera 203 may include one or more cameras configurable to capture images of a scene. In some exemplary embodiments, camera 203 is polarimetric in that it captures the polarization characteristics of light in a scene and generates an image based on the captured polarization information. Polarimetric cameras are sometimes referred to as imaging polarimeters. In some exemplary embodiments, camera 203 may include a PX 3200 polarimetric camera from Photon-X, Inc. of Orlando, Florida, USA, an Ursa polarimeter from Polaris Sensor Technologies, Inc. of Huntsville, Alabama, USA, or a PolarCam™ snapshot micropolarizer camera from 4D Technology, Inc. of Tempe, Arizona, USA. In some exemplary embodiments, camera 203 (or the platform to which it is attached) may include one or more of a position sensor (e.g., a GPS sensor), an inertial navigation system including a motion sensor (e.g., an accelerometer), and a rotation sensor (e.g., a gyroscope) that can be used to indicate the position and / or attitude of the camera. When camera 203 includes multiple cameras, the cameras may all be of the same specification, or may include cameras with different specifications and capabilities. In some embodiments, two or more cameras, which may be in the same or different positions within a scene, may be operated by image processing systems that are synchronized with each other to scan the scene.

[0045] The application software 205 comprises instruction logic for acquiring images of a scene using components of the 3D image processing system and generating a 3D scene module. The application software may include instructions for acquiring inputs related to the scanning operation and the model to be generated, instructions for generating goals and other commands for the scanning operation and model generation, and instructions for causing the SRE 201 and camera 203 to perform actions for the scanning operation and model generation. The application software 205 may also include instructions for performing processes after the scene model is generated, such as, for example, an application using the generated scene model. An exemplary flow of the application software is described with reference to FIG. 14. An exemplary application using a volumetric scene model generated according to an embodiment is described below with reference to FIG. 33.

[0046] The data communication layer 207 comprises components of the 3D imaging system for communicating with each other, one or more components of the 3D imaging system for communicating with external devices via local area network or wide area network communication, and subcomponents of components of the 3D imaging system for communicating with other subcomponents or other components of the 3D imaging system. The data communication layer 207 may include interfaces and / or protocols for a communication technology or combination of communication technologies. Exemplary communication technologies for the data communication layer include one or more communication buses (e.g., PCI, PCI Express, SATA, Firewire, USB, Infiniband, etc.), as well as network technologies such as Ethernet (IEEE 802.3) and / or wireless communication technologies (Bluetooth, WiFi (IEEE 802.11), NFC, GSM, CDMA2000, UMTS, LTE, LTE-Advanced (LTE-A)), and / or other short-, medium-, and / or long-range wireless communication technologies.

[0047] The database 209 is a data store for storing configuration parameters for the 3D image processing system, scene images captured via scans, a library of previously acquired images and volumetric scene models, and volumetric scene models currently being generated or refined. At least a portion of the database 209 may reside in the memory 215. In some embodiments, a portion of the database 209 may be distributed between the memory 215 and an external storage device (e.g., cloud storage or other remote storage accessible via the data communication layer 207). The database 209 may store some data in a manner that is efficient for writing, accessing, and retrieving, but is not limited to a particular type of database or data model. In some exemplary embodiments, the database 209 uses an octree and / or quadtree format to store volumetric scene models. The octree and quadtree formats are further described with respect to FIG. 20 and other figures. In some embodiments, the database 209 may use a first type of data format to store scanned images and a second type of data format to store polarimetric scene model information. Octrees are spatially sorted, so regions of a scene can be accessed directly. In addition, they can be accessed efficiently in a specified direction in space, such as the direction of a particular ray in a scene. They are also hierarchical, so that a coarse-to-fine algorithm can process information at a lower resolution level depending on the results at the coarse level until higher resolution information is required for the operation of the algorithm. This also makes access from secondary memory efficient; only lower resolution information is kept in primary memory until the higher resolution is actually required.

[0048] 2B is a block diagram of SRE hardware according to some example embodiments, comprising at least an input / output interface 211, at least one processor 213, and memory 215. The output interface 211 provides one or more interfaces through which the 3D image processing system 200 or its components may interact with a user and / or other devices. The input / output interface 211 may enable the 3D image processing system 200 to communicate with input devices (e.g., keyboard, touchscreen, voice command input, controller for guiding camera movement, etc.) and / or output devices such as a screen and / or additional storage devices. The output interface may enable wired and / or wireless connections to any of the input, output, or storage devices.

[0049] The processor 213 may include, for example, one or more of a single or multi-core processor, a microprocessor (e.g., a central processing unit or CPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, or a system-on-chip (SOC) (e.g., an integrated circuit that includes a CPU and other hardware components such as memory, networking interfaces, and the like). In some embodiments, each or any of these processors uses an instruction set architecture such as x86 or Advanced RISC Machine (ARM). The processor 213 may perform operations of one or more of the SRE 201, the camera 203, the application software 205, the data communication layer 207, the database 209, and the input / output interface 211. In some exemplary embodiments, the processor 213 includes at least three distributed processors such that the 3D image processing system 200 includes a first processor in the SRE 201, a second processor in the camera 203, and a third processor that controls the processing of the 3D image processing system 200.

[0050] The memory 215 includes random access memory (RAM) (such as dynamic RAM (DRAM) or static RAM (SRAM)), flash memory (e.g., based on NAND or NOR technology), hard disk, magneto-optical medium, optical medium, cache memory, registers (e.g., holding instructions), or other types of devices that perform storage of data and / or instructions (e.g., software executed on or by the processor) to non-volatile or non-volatile storage. The memory may be a volatile memory or a non-volatile computer-readable storage medium. In an exemplary embodiment, the volumetric scene model information and scanned images may be distributed by the processor 213 among different memories (e.g., volatile memory vs. non-volatile memory, cache memory vs. RAM, etc.) based on the dynamic execution needs of the 3D image processing system.

[0051] The user interface module 217 implements software operations to obtain input from and output information to a user. The interface 217 may be implemented using browser-based or other user interface generation technologies.

[0052] The job script processing module 219 receives goals and initial parameters for scanning the scene and model generation for the scene, and implements operations to generate a sequence of operations that achieve the desired goals in scanning and model generation.

[0053] While FIG. 2 illustrates one embodiment in which camera 203 is separate from SRE 201 and other components of 3D image processing system 200, it will be understood that contemplated embodiments according to the present disclosure include embodiments in which the camera is in a separate housing from the rest of the system and communicates via a network interface, embodiments in which the camera and SRE are in a common housing (in some embodiments, the camera is mounted on a circuit board with a chip that implements the operation of the SRE) and communicates with database 209 via a network interface, embodiments in which camera 203 is integrated into the same housing as components 201, 205-215, and the like.

[0054] 3 shows a flowchart of a process 300 for scanning a scene and recovering a volumetric model of the scene, according to some example embodiments. For example, the process 300 may be performed by the 3D image processing system 200 to scan the real scene 101 and form a volumetric scene model of the scene. The at least one processor 213 may execute instructions associated with the process 300, such as those provided by the application software 205, the SRE 201, the camera 203, and other components of the 3D image processing system 200.

[0055] After entering process 300, in operation 301, the 3D imaging system may present a user interface to a user for providing input related to activating and calibrating one or more cameras, initializing a scene model, and controlling subsequent operations of process 300. For example, the user interface may be generated by user interface module 217 and displayed on an output device via input / output interface 211. A user may provide input using one or more input devices coupled to the 3D imaging system via input / output interface 211.

[0056] In operation 303, one or more cameras are activated and optionally calibrated. In some embodiments, camera calibration is performed to determine response models for various properties (e.g., geometric, radiometric, and polarimetric properties) measured by the cameras. Calibration may be performed as an offline step, where calibration parameters may be stored in memory, and the cameras and / or 3D image processing system may access the stored calibration information for each camera and perform the necessary calibration, if any. In embodiments in which the cameras (or the platform to which they are attached) include position sensors (e.g., GPS sensors), motion sensors (e.g., accelerometers), and / or rotation sensors (e.g., gyroscopes), these sensors may be calibrated to, for example, determine or set the current position and / or attitude of each camera. Camera calibration is further described below with respect to sensor modeling module 1205 shown in FIG. 12.

[0057] In operation 305, a scene model may be initialized. A volumetric scene model may be represented in memory by volume elements for the scene space. Initializing the scene model before starting the scan allows the scene to be scanned in a manner that is tailored to the particular scene and therefore more efficient. Initialization may also prepare data structures and / or memory for storing the images and scene model.

[0058] Initializing the scene model may include a step in which a user specifies a description of the scene via a user interface. The scene description may be a high-level description, such as specifying that the scene includes a kitchen with a countertop, a hood, a vase with flowers, a wooden cupboard, and a glass window when preparing to image the scene of FIG. 1. The plan views shown in FIGS. 5 and 7 may be an example of a scene description provided to the 3D image processing system. The above exemplary description of the kitchen scene of scene 101 is merely an example, and the scene description provided to the image processing system may be more or less specific than the above exemplary description. The input description may be provided to the image processing system by one or more of text entry, menu selection, voice command, and the like. In some exemplary embodiments, the input description of the scene may take the form of a CAD design of the space.

[0059] The initializing step may also include selecting the format in which the volumetric scene model will be stored (e.g., octree or other format), the storage location, the file name, the camera to be used for scanning, etc.

[0060] Subsequent operations of process 300 are directed to iteratively refining the scene model by characterizing each volume element that represents a portion of the scene being modeled.

[0061] In operation 307, one or more goals for the scene model and the corresponding sequence of plans and tasks are determined. The one or more goals may be specified by a user via a user interface. These goals may be specified in any form made available by the user interface.

[0062] An exemplary goal specified by a user in connection with image processing the scene 101 may be to have a scene model in which petals are determined with a pre-specified level of accuracy (e.g., 90% accuracy of the model with respect to the model's description), i.e., to have volume elements in the volumetric scene model corresponding to the real scene occupied by petals determined as including petals (or material corresponding to petals) with 90% or greater accuracy. Other goal criteria may be lowering the accuracy with respect to one or more aspects of the scene region below a specified threshold (e.g., the thickness of the petals in the obtained model to be within 10 microns of a specified thickness) and coverage level (e.g., a certain percentage of volume elements in the scene space are resolved with respect to the media type contained therein). Another goal criterion may be iteratively refining the model until incremental improvement within the scene or portion of the scene with respect to one or more criteria falls below a threshold (e.g., repeating until at least a predetermined portion of the volume elements change their media type determination).

[0063] The image processing system, automatically or in combination with manual input from a user, generates a plan containing one or more tasks to achieve a goal. As noted above, a goal may be specified by requirements that meet one or more identified criteria. A plan is a sequence of tasks arranged to meet these goals. The plan process is described below with respect to FIG. 15. A task may be considered an action to be performed. The scanning process, which may occur when performing a task, is described below with respect to FIG. 16.

[0064] In the example of scanning a scene 101, to satisfy the goal of creating a scene model in which flower petals are determined with a pre-specified level of accuracy, the generated plan may include a first task to move the camera in a specific path through the scene space and / or pan the camera to obtain an initial scan of the scene space, a second task to move the camera in a specific path and / or pan the camera to obtain a detailed scan of the flower as specified in the goal, and a third task to test the goal criteria against the current volumetric scene model and repeat the second task until the goal is satisfied. The tasks may also specify parameters for image acquisition, image storage, and scene model calculation. Camera movement may be specified in terms of any of position, path, pose (orientation), distance traveled, speed, and / or time.

[0065] The accuracy of a scene model can be determined for a volume element in the model by comparing current determinations of the medium within the volume element and the volume element's optical properties with images captured from corresponding positions and poses. Figure 42 illustrates a measure of scene model accuracy (SMA) in accordance with some example embodiments. In the example of Figure 42, positional properties are indicated for several voxels (volume elements) within a region of adjacent voxels. Voxel centers 4201, 4211, and 4221 in the true scene model are indicated. The true scene model, as that term is used in this example, is an estimated high-accuracy model of the scene. The test scene model for which SMA is sought includes estimated coordinates of points 4205, 4215, and 4225 corresponding to voxel centers 4201, 4211, and 4221 in the true model. The distances 4203, 4213, and 4223 between corresponding points in the two models define the (inverse) SMA. Before calculating the distance, the test object model is typically adjusted in a manner that preserves internal structure while reducing global (systematic) misregistration between it and the true model. A 6-DOF rigid body coordinate transformation may be applied to the test object model based on an iterative closest-point fitting operation to the true model. The example in FIG. 42 uses spatial location as the modeled feature to generalize to features modeled by SRE. Deviation functions may be defined over heterogeneous combinations of features (e.g., positional deviation of point-like features plus radiometric deviation of light field features at corresponding voxels). Penalty functions may be used to mitigate the effect of outlier deviations when calculating the SMA. An application will always propose a method for constructing a robust deviation function.

[0066] In operation 309, the scene model is roughed in by moving and / or panning the camera around the scene space. This may be performed according to one or more of the tasks in the generated plan. In some exemplary embodiments, the camera may be moved by a user with or without prompting by the 3D image processing system. In some other exemplary embodiments, the camera movement may be controlled by the 3D image processing system, for example, the camera is mounted on a mounting or railing that allows it to be controllably moved, or in a UAV (e.g., a drone that can be controlled to operate freely within the scene space).

[0067] An important consideration in this operation is the ability of an embodiment to perform the scan in a manner that minimizes interaction between the light field and the camera in scene space. Rough-in of a scene model is intended to form an initial scene model of the scene space. Rough-in may include moving and / or panning the camera within the scene space to capture one or more images of each of the entities in the scene specified by the goals and of the space surrounding the goal entities. Those skilled in the art will understand that rough-in may involve any amount of image capture and / or model creation, and that a minimum number of images of the goal entities and minimum coverage of the scene space is sufficient to form the initial scene model, but that a greater number of images covering a larger portion of the goal entities and / or the scene space will produce a more complete and accurate initial scene model. Generally, a high-quality initial scene model resulting from rough-in reduces the amount of iterative refinement of the scene model subsequently required to obtain a model within the specified set of goals.

[0068] Operations 311-313 iteratively refine the initial scene model until all specified goal conditions are met. In operation 311, selected aspects of the scene space undergo further scanning. For example, the bidirectional light interaction function (BLIF) of a flower may be determined by acquiring images while moving on a trajectory around a small area of ​​the petals, leaves, or bottle shown in FIG. 1 . BLIF is described below with respect to FIG. 10 . The trajectory and scanning parameters may follow one of the tasks generated from the initially specified goals. For example, if the rough-in involved moving and / or panning the camera over a wide area, the refinement process may involve moving the camera within a small trajectory around a particular object of interest. Similarly, the scanning parameters may be changed between the rough-in and refinement stages; the rough-in may involve scanning with a wide FOV, while the refinement process may involve using a significantly smaller FOV (e.g., a close-up of the object of interest (OOI) view). This operation may include, for example, determining material properties by comparing measured light field parameters to a statistical model.

[0069] In each of operations 309-311, the acquired images may be stored in memory. Generation or refinement of the volumetric scene model may be performed concurrently with image acquisition. Refining the volumetric scene model involves improving the estimates of each volume element of the scene model. Further description of how the estimates of the volume elements are improved is provided below with respect to FIGS. 18A-18B. The volumetric scene model may be stored in memory in a manner that is space-efficient for storage and time-efficient for reading and writing. FIGS. 4A-4C and 9A-9E illustrate parameters used in maintaining the scene for the volumetric scene model, while FIG. 20 and related discussion describe how the volume elements and aspects thereof are stored in an efficient data structure.

[0070] In operation 313, it is determined whether the conditions of one or more specified goals are met. If it is determined that the conditions of the goals are met, process 300 ends. If it is determined that the conditions of one or more goals are not met, operations 311-313 are repeated until the conditions of all goals are met. After the current instance of scanning scene 101, the currently acquired scene model may be tested to determine whether the conditions of the specified goals are met, and if not, the scanning parameters may be adjusted and operation 311 may be continued again.

[0071] After operation 313, the process 300 ends. The scene model of the scene 101 stored in memory can be used in any application.

[0072] 4A and 4B show how the SRE models space for reconstruction purposes. Understanding these fundamental geometric entities helps understand the following description of scene model entities used in some examples.

[0073] A volume element 407 represents a finite spatial extent in 3D space. A volume element is known by the abbreviated name "voxel." A voxel is bounded by a closed surface. A voxel may have a non-degenerate shape. It need not be cubic or exhibit any particular symmetry. A voxel is characterized by its shape and the size of the volume it encloses, expressed in cubic meters or other suitable units. One or more voxels together make up a volume field 405. In an exemplary embodiment, an SRE (e.g., SRE 201) uses a volume field to represent the medium (or lack thereof) occupying an area in a scene. In some embodiments, a hierarchical spatial construct can be used to achieve increased processing throughput and / or reduced data size of the volume field. See the discussion of octrees below.

[0074] A solid angle element 403 represents an angular extent in 3D space. Solid angle elements are known by the abbreviated name "sail." A sail is bounded by a surface that is radially flat. That is, a line extending from the origin outward along the bounding surface of the sail is a straight line. Although shown as such in FIG. 4B, a sail need not have a conical shape or be symmetrical. A sail can have a shape that fits the definition of a general conical surface and is characterized by its shape and the number of angular units it subtends, usually expressed in steradians. A sail has infinite extent along the radian direction. One or more sails together make up a solid angle field 401. In an exemplary embodiment, an SRE (e.g., SRE 201) uses a solid angle field to represent the radiance of the light field in different directions at a voxel. In the same manner as for voxels above, a hierarchical spatial construct can be used to speed up and / or reduce the data size of the solid angle field. See the discussion of the solid angle octree below with reference to FIG. 21.

[0075] Figure 5 shows a plan view of the exemplary everyday scene shown in Figure 1. The scene model description input to the 3D image processing system in process 300 may include a description similar to Figure 5 (e.g., a text description, a machine-readable sketch of the plan view, etc.). In an exemplary embodiment, any one of user-specified scene information, an initial scene model generated from a rough-in (e.g., see the description of operation 309 above), and / or an iteratively refined version of the scene model (e.g., see the description of operation 311 above), or any combination thereof, may be a digital representation of the scene space depicted in Figure 5.

[0076] The scene model 500 may include several indoor and outdoor entities for the example kitchen scene shown in FIG. 1. The workspace 517 divides the scene space into two broad regions. The workspace is the scene region accessed by the camera (physical camera device) used to record (sample) the light field. It may be defined as the convex hull of such camera positions or by other suitable geometric constructs that indicate the accessed positions. Voxels inside the workspace can be observed from a full sphere of surrounding viewpoints (a "full orbit"), according to the density of actual viewpoints and according to occlusions by media within the workspace. Voxels outside the workspace can be observed from only a hemisphere (half-space) of viewpoints. This supports the conclusion that opaque closed surfaces outside the workspace (e.g., statues, basketballs) cannot be completely reconstructed using observations recorded by cameras inside the workspace.

[0077] Inside the workspace, a jar containing flowers rests on a curved countertop 511. Region 113 containing the flowers and jar has been designated as OOI for detailed reconstruction. Several other entities have been placed outside the workspace. Region 103 containing part of a tree is placed outside the windowed wall of the kitchen. Windows 505 and 506 are within the wall. Region 109 containing part of one of the door panes is placed on windowed door 503. Another region 105 contains part of a cupboard. Near stove hood 507 is region 107 containing part of the hood.

[0078] Three overall stages of scene reconstruction (in this example) are shown in accordance with some exemplary embodiments. First, an SRE (e.g., SRE 201) “roughs in” the entire scene within the entire vicinity of the OOI by approximately panning (513) the camera in multiple directions (which may also include moving along a path) to observe and reconstruct the light field. In this example, a human user performs the panning as guided and prompted by the SRE (e.g., SRE 201 as guided by application software 205). Voxels 522 terminate in corridors 519 and 521 originating from the two camera viewpoints from the scene rough-in pan 513. Next, the SRE performs an initial BLIF reconstruction of media (e.g., petals, leaves, stems, bottles) present within the OOI by guiding a short camera trajectory arc 515 of the OOI. Finally, the SRE achieves a high-resolution reconstruction of the OOI by guiding a detailed scan of the OOI. A detailed scan involves one or more full or partial orbits 509 of the OOI and requires as many observations as necessary to meet the accuracy and / or resolution goals of the reconstruction workflow. A more complete description is provided below with reference to the SRE operation (FIGS. 14-18B).

[0079] In some exemplary embodiments, a scene model (including an everyday scene model) exclusively includes medials and radiels. Medials are voxels that contain the medium that interacts with the light field as described by BLIF. Radiels are radiometric components of the light field and may be represented by sails paired with radiometric power values ​​(radiance, flux, or other suitable quantities). Medials may be primitive or composite. A primitive medial is one whose BLIF and radiant light field are modeled with high consistency by one of the BLIF models available to SRE within a given workflow (e.g., in some applications, less than 5% RMS relative deviation between the radiance values ​​of the observed exit light field (formed digital image) and the radiance values ​​of the exit light field (projected digital image) predicted by the model). A composite medial is one whose BLIF is not modeled in this way. After observation, composite medial, or regions of composite medial, are also known as "observed but unexplained" or "observed but unexplained" regions. Composite media can become primitive media as more observations and / or updated media models become available to the SRE. Primitive media can generally be represented with greater parsimony (in the Occam's razor sense) than composite media. Primitive media more narrowly (implying "usefully") answer the question "What type of media occupies this voxel in space?"

[0080] As can be seen in FIG. 1, the kitchen window in region 109 has open space on either side (in front of and behind the glass). FIG. 6 shows a representation of a portion of region 109 (shown in FIGS. 1 and 5) including the kitchen window glass from a volumetric scene model according to some example embodiments. FIG. 6 illustrates how voxels are used to accurately represent media in scene space. A model 601 of a portion of glass region 109 consists of a bulk glass medium 603, a glass-against-air surfel 605, and a bulk air medium 607. A surfel is a medium that represents a planar boundary between regions of different types of bulk medium, typically with different refractive indices, that causes reflection when light is incident on the boundary. A reflective surface (reflective medium) contains one or more surfels. Each surfel is partially occupied by glass (609) and partially occupied by air (613). Glass and air meet at a planar boundary 611. A typical SRE scenario involves a parsimonious model of the glass's BLIF. Air-to-glass surfels are represented as type "simple surfels," with appropriately detailed BLIFs describing the refractive index and other relevant properties. The bulk glass medium inside the window glass is represented as type "simple bulk," with physical properties similar to those of the simple glass surfels.

[0081] Model 601 may be thought of as a "corridor of voxels" or "corridor of media" representing multiple voxels (or media) in the direction of the field of view of a camera capturing a corresponding image. The corridor may extend horizontally from the camera through glass, for example, in region 109. The voxels and / or media within the corridor may or may not be of uniform size.

[0082] In the example scenes of Figures 1 and 5, in an example embodiment, different media regions are restored less or more parsimoniously as informed by the restoration goals and the recorded observations. A more significant share of computing resources is focused on restoring the OOI as primitive media compared to the rest of the scene, which may be represented as composite media, each with an exit light field used in the restoration calculations performed on the OOI.

[0083] FIG. 7A is a geometric diagram showing a plan view of the model of the kitchen scene shown in FIG. 1. The scene model 500 is bounded by a scene model boundary 704. The scene model boundary is appropriately sized to accommodate the OOI and other entities whose outgoing light field significantly affects the OOI's incoming light field. Three corridors 709A, 709B, and 709C extend a short distance from the kitchen windows 503, 505, and 506 to the scene model boundary 704. Where the corridors terminate at the scene model boundary, the incoming light field defines three frontiers 703A, 703B, and 703C. Each of the three frontiers is a "surface light field" consisting of radii pointing roughly inward along the frontier's corridor. The kitchen's interior corridor 709D covers the entire interior volume of the kitchen. The area 713 beyond the kitchen's opaque wall 701 is not observed by the cameras in the exemplary workspace. The intermediate space 715 region extends from the workspace 517 to the scene model boundary 704 (the intermediate space is the space between the workspace and the scene model boundary).

[0084] FIG. 7B is a geometric diagram showing a plan view of another model of the kitchen scene shown in FIG. 1. Compared to FIG. 7A, the scene model boundary 705 is significantly farther from the workspace. In some exemplary scenarios, as the camera moves around to capture the kitchen scene and / or as the reconstruction process progresses, the scene model naturally expands in spatial extent from the size shown in FIG. 7A to the size shown in FIG. 7B. SRE's tendency toward parsimony (Occam's razor) in representing the scene region being reconstructed tends to "push" the scene model boundary outward to a distance where disparity is not meaningfully observable (i.e., medial multi-view reconstruction is not robustly achievable). As reconstruction progresses, the scene model boundary also tends to become centered on the workspace. The scene model 500 in FIG. 7A can be considered an earlier stage of reconstruction (of the kitchen scene), while the scene model 501 in FIG. 7B can be considered a later stage of reconstruction.

[0085] Three corridors 719A, 719B, and 719C extend outward from the workspace to the scene model boundary, which is now at a considerable distance. Similar to the frontiers in the narrower scene model of Figure 7A, each of frontiers 717A, 717B, and 717C defines a surface light field representing light incident on a respective portion of the scene model boundary 704. For example, the surface light field of frontier 717B generally has a "smaller disparity" than the surface light field of frontier 703B, the corresponding frontier from the earlier stages of the reconstruction shown in Figure 7A. That is, a given subregion of frontier 717B has a radius that is oriented in a narrower directional span (toward kitchen window 503) than a subregion of frontier 703B. At the limit, when the scene model boundary is extremely far from the workspace (e.g., in advanced stages of reconstructing a wide, open space such as outdoors), each subregion of the frontier contains a single, very narrow radius pointing radially inward toward the center of the workspace. Sky-containing region 115 shown in FIG. 1 is represented within frontier 717A. Intermediate space 711 is large compared to intermediate space 715 in FIG. 7A. Mountain 707 is located within the intermediate space. Depending on the media types in the real scene, the atmospheric conditions (e.g., haze) in the observation corridor, the available media models, and the operating settings of the SRE, some degree of reconstruction of the media comprising the mountain is possible.

[0086] FIG. 8A is a geometric diagram illustrating the poses of one or more cameras used to reconstruct media. In the example of FIG. 8A, one or more cameras image voxel 801 from multiple viewpoints. Each viewpoint 803, 805, 807, 809, 811, and 813 records the exit radiance value (and polarimetric properties) for a particular radius of the light field emanating from voxel 801. The reconstruction process can use images recorded at significantly different distances, as shown for poses 803 and 807, with respect to voxel of interest 801. The model of light transport (including BLIF interactions) used in the reconstruction takes into account the relative orientation and differences in subtended solid angles between the voxel of interest and the imaging viewpoints.

[0087] Figures 8B, 8C, 8D, and 8E are geometric diagrams showing various possible media types for voxel 801 in an example scenario. These are typical media involved in scene modeling in SRE. Glass 815 media is a simple bulk glass, as in Figure 6. Wood 817 media represents the heterogeneous media that make up tree branches in their natural configuration, including air between the solid media (leaves, wood) of the tree. Wood media can be composite media because BLIF's "low-dimensional" parametric model does not fit the observed light field emanating from the wood media. Low-dimensional parametric models use reasonably few mathematical quantities to represent the salient properties of an entity. Wood 819 media is surfels, representing the wood surface relative to air. Wood media can be simple if the physical wood medium is sufficiently homogeneous in material composition so that low-dimensional parametric BLIF can model its light field interaction behavior. In the example of Figure 8E, the BLIF of a metal 821 surfel may be represented by a scalar value for the refractive index, a second scalar value for the extinction coefficient, and a third scalar value for the anisotropic "grain" of the surface in the case of a matte metal. A determination of whether the light field associated with a voxel indicates a particular type of media may be made by comparing the light field associated with the voxel to predetermined statistical models of the BLIF for various media types. The media modeling module 1207 maintains statistical models of the BLIF for various media types.

[0088] 9A-9F show exemplary sail arrangements that may be present in a scene. FIG. 9A shows a single-center unidirectional sail arrangement 901. A single radial sail in a light field is an example of this sail arrangement. FIG. 9B shows a single-center multidirectional arrangement 903 of sails. A collection of multiple radial sails in a light field is an example of this sail arrangement, where the sails share a common origin voxel. FIG. 9C shows a single-center omnidirectional arrangement 905 of sails. A collection of multiple radial sails in a light field is an example of this sail arrangement, where the sails share a common origin voxel and together completely cover (fill) a sphere of direction. When each sail is paired with a radiance value to generate a radial, a single-center omnidirectional sail arrangement is called a "point light field." FIG. 9D shows a single-center omnidirectional isotropic sail arrangement 907. A single radial sail in an isotropic light field is an example of this sail arrangement. Because the light field is isotropic in voxels (equal radiance in all directions), a point light field can be represented by a single coarse-grained radial covering the entire sphere of directions. In some SRE embodiments, this may be realized as the root (coarsest) node of a solid angle octree. Further details regarding solid angle octrees are provided below with reference to Figure 21.

[0089] FIG. 9E shows a planar-centered unidirectional arrangement of sails 911. A collection of sails subtended by pixels in the idealized focal plane of a camera (one sail per pixel) is an example of this sail arrangement type. Each pixel conveys a radiance value for the radius it subtends in the scene. Note that the planar-centered unidirectional sail arrangement 911, when arranged on a 2D manifold of voxels and paired with a radiance value, is a subtype of the more general (non-planar) multi-centered multi-directional type of sail arrangement, also called a "surface light field." FIG. 9F shows a sail arrangement 913 with multiple volumetric centers and omnidirectional sails. A collection of point light fields (defined above) is an example of this sail arrangement type. In some embodiments, a well-arranged collection of such point light fields yields a useful representation of the light field within an elongated region of scene space.

[0090] Figure 10 is an isometric view showing BLIFs relating incident, emitted, and outgoing light fields. Figure 10 shows a model that can be used to represent interactions occurring in a single medium, where the medium consists of voxels 1003 and associated BLIFs 1005. Radii of the incident light field 1001 enter the medium. The BLIFs operate on the incident light field to generate a response light field 1011 that exits the medium. The total outgoing light field 1007 is a combination of the response light field and the (optional) emitted light field 1009. The emitted light field is emitted by the medium regardless of stimulation by incident light.

[0091] 11 is a block diagram of an SRE, such as SRE 201, illustrating some of the operational modules according to some exemplary embodiments. Operational modules 1101-1115 include instruction logic for performing certain functions in scanning a scene and / or scene reconstruction and may be implemented using software, firmware, hardware, or any combination thereof. Each of modules 1101-1115 may communicate with others of modules 1101-1115 or with other components of the 3D image processing system via a data communication layer, such as data communication layer 207 of 3D image processing system 200 described above.

[0092] The SRE command processing module 1101 receives commands from a calling environment, which may include a user interface or other component of a 3D image processing system. These commands may be implemented as compiled software function calls, as interpreted script directives, or in any other suitable form.

[0093] The plan processing module 1103 creates and executes plans toward an iterative scene reconstruction goal. The scene reconstruction goal may be provided to the plan processing module 1103 by application software controlling the SRE. The plan processing module 1103 may control the scan processing module 1105, the scene resolution module 1107, and the scene display module 1109 according to one or more tasks defined for each plan to achieve the goal. More details regarding the plan processing module's decomposition of a goal into a sequence of subcommands are provided below with reference to FIG. 15. In some embodiments, the application software may bypass the plan processing function and interface directly with modules 1105-1109. The plan processing module 1103 may perform sensor calibration before scanning of the scene begins.

[0094] The scan processing module 1105 acquires the sensed data necessary to drive one or more sensors (e.g., a camera or multiple cameras operating in unison) to achieve the scene reconstruction goal. This may include dynamic guidance regarding the sensor's pose and / or other operating parameters, which may be communicated or informed by a human user's actions to the sensor control module 1111 via a user interface provided by the application software controlling the process. The sensor control module 1111 manages the individual sensors to acquire data and directs the sensed data to the appropriate modules that consume it. In response to the sensed data, the sensor control module 1111 may dynamically adjust geometric, radiometric, and polarimetric degrees of freedom as needed to successfully complete the ongoing scan.

[0095] The scene resolution module 1107 estimates values ​​of one or more physical properties of the hypothesized scene model. The estimated values ​​maximize the consistency between the hypothesized scene model and observations of the corresponding real scene (or between the hypothesized model and one or more other scene models). Scene resolution, including details regarding consistency calculation and updating of modeled property values, is further described with respect to Figures 18A-18B below.

[0096] The spatial processing module 1113 operates on a hierarchically subdivided representation of scene entities. In some embodiments, some of the operations of the spatial processing 1113 are selectively performed with improved efficiency using an array of parallel computing elements, specialized processors, and the like. One example of such improved efficiency is the transport of light field radiance values ​​between media in a scene. In one exemplary embodiment, the incident light field generation module 2003 (shown in FIG. 20) processes each incident radial using a small group of FPGA cores. When processing a large number of media and / or incident radials, FPGA-based exemplary embodiments can perform light transport calculations for thousands of radials in parallel. This is in contrast to conventional CPU-based embodiments, which may process at most tens of incident radials simultaneously. GPU-based embodiments can operate thousands in parallel, but at a significantly higher cost in terms of power consumption. Similar discussion of efficiency when operating on a large number of incident and / or exit radials applies to the incident / exit light field processing module 2007 (shown in FIG. 20). Spatial processing operations, including solid angle octrees, are further described below with reference to FIG. 19 and subsequent figures.

[0097] The scene display module 1109 prepares a visual representation of the scene for human viewing. Such a representation may be realistic in nature, analytical in nature, or a combination of the two. The SRE provides a scene display 1109 function that generates synthetic images of the scene in two broad modes: realistic and analytical. The realistic representation of the scene synthesizes a "first person" image as it would appear to a real camera if immersed in the scene as represented by the current state of the reconstructed volumetric scene model at a specified viewpoint. The light energy received by a pixel of the virtual camera is calculated by reconstructing the light field of the scene model at that pixel. This is accomplished by solving for (and integrating) the radii of the light field at the synthesized pixel using the scene solution module 1107. Synthesis may incorporate modeled characteristics of the camera as described with reference to the sensor modeling module 1205 above. Pseudocolor coloring may be used for the synthesized pixel values, as long as the pseudocolor assigned to the pixel is based on the reconstructed radiometric energy at that pixel. An analytical representation of a scene represents a portion of the scene and synthesizes an image that is not a real image (as explained above).

[0098] The light field physics processing module 1115 operates based on measuring the light field and modeling aspects of the light field within a scene model. This module is described below with respect to FIG. 13.

[0099] The data modeling module 1117 operates to model various aspects including the scene to be scanned and the sensor to be used. The data modeling module 1117 is described with respect to FIG.

[0100] In addition to the above modules 1103-1117, the SRE may also include other modules. These may include, but are not limited to, interfacing with a database, allocating local and remote computing resources (e.g., load balancing), and handling operational errors gracefully. For database access, the SRE may use specialized data structures to efficiently read and write large amounts of spatial data. In some embodiments, octrees, quadtrees, and / or solid angle octrees may be used for storing the vast number of medials and radii present at various resolutions in the scene model. The SRE may store typically much smaller amounts of non-spatial data, such as sensor parameters, operational settings, and the like, using standard database records and techniques. The SRE may provide analytical queries (e.g., "What total volume is represented by a contiguous group of voxels in a given BLIF XXX?").

[0101] 12 is a block diagram illustrating a data modeling module 1201 of an SRE, such as, for example, data modeling module 1117 of SRE 201, according to some example embodiments. Data modeling module 1201 may include a scene modeling module 1203, a sensor modeling module 1205, a media modeling module 1207, an observation modeling module 1209, a (feature) kernel modeling module 1211, a resolution rule module 1213, a rule merging module 1215, and an operation setting module 1217.

[0102] Scene modeling module 1203 includes operations to generate an initial model of the scene (e.g., see the description of operation 307 in FIG. 3) and to refine aspects of the initial scene model in conjunction with other data modeling modules. This initial scene model is described above with respect to FIGS. 5 and 7. The voxels and other elements from which the scene is modeled are described above with respect to FIGS. 4, 6, 8B-8E, and 9A-9F. Updating / refinement of the scene model (referred to as the "assumed scene model") is described with respect to FIGS. 18A-18B.

[0103] The sensor modeling module 1205 represents the characteristics of sensors (e.g., cameras 203) used in the scene reconstruction process of the measured scene. Sensors may be divided into two broad categories: cameras that sense the light field and sensors that sense other characteristics of the scene. Each camera is modeled by one or more of its geometric, radiometric, and polarimetric properties. The geometric properties may dictate how a given scene space is mapped to one or more spatially indexed light-sensitive elements (e.g., pixel photosites) of the camera. The radiometric properties may dictate how strongly a radiance value of a particular wavelength of light excites a pixel when incident. The polarimetric properties may dictate the relationship between the polarization properties of the radii (e.g., the elliptical polarization state as represented by the Stokes vector) and the excitation intensity at the pixel when incident. These three classes of properties together define the forward mapping from the sensed radii for each pixel to the digital values ​​output by the camera.

[0104] By appropriately inverting this mapping, the sensor modeling module 1205 enables a corresponding inverse mapping from digital pixel values ​​to physically meaningful properties of the radiation. Such properties include the radiance, spectral band (wavelength), and polarization state of the observed radiation. The polarization state of light is characterized in some embodiments by a polarization ellipsoid (a four-element Stokes vector). In some embodiments, a simplified polarimeter architecture models a reduced set of polarimetric properties. Sensing the linear component rather than the circular component of the full polarization state is one example of this. The inability to sense the circular polarization component can limit the ability to accurately reconstruct organic media, such as plant leaves, which tend to induce significant circular polarization in the reflected light.

[0105] The above geometric, radiometric, and polarimetric properties of the camera are represented by a camera response model. The sensor modeling module 1205 may determine these response models. For some cameras, some of the response models may be parameterized with a small number of degrees of freedom (or even a single degree of freedom). Conversely, some of the response models may be parameterized with a large number of degrees of freedom. The dimensionality (number of degrees of freedom) of the components of the camera response model may depend on how much uncertainty a given reconstruction goal can tolerate in the various light field properties measured by the camera. For example, a polarimeter using a filter mask of “micropolarization” elements on the camera pixels may require separate 4×4 Mueller correction matrices (millions of real scalars) for each pixel to measure the polarimetric properties of the light field with the uncertainty required by the example goal. In contrast, a polarimeter using a rotating polarizing filter may require only a single global 4×4 Mueller correction matrix (16 real scalars) that is sufficient for all pixels. In another example, SRE corrects for camera lens distortion using an idealized "pinhole camera" model. For a given lens, the correction involves anywhere from five to fourteen or more real-valued correction parameters. In some cases, a parametric model may not be available or practical. In such cases, a more literal representation of the response model, such as a lookup table, may be employed.

[0106] The above response model can be flexibly applied at different stages of the restoration flow: for example, there can be speed advantages in correcting camera lens distortion on-the-fly for each individual pixel rather than as a pre-processing step on the whole image.

[0107] If not supplied by the vendor or other external source, the camera response model is discovered by performing one or more calibration procedures. Geometric calibration is ubiquitous in the field of computer vision and is performed, for example, by imaging a chessboard or other optical target of known geometric shape. Radiometric calibration can be performed by imaging a target with known spectral radiance. Polarimetric calibration can be performed by imaging a target with known polarization states. Low uncertainty in all three response models is desirable for reconstruction performance, which depends on the ability of the SRE to predict light field observations based on the light field of an assumed scene model. If one or more of the response models has high uncertainty, the SRE may have weaker predictive ability for the observations recorded by that sensor.

[0108] Exemplary embodiments enable accurate 6-DOF localization of a camera with respect to a scene it contains, where the scene itself (when stationary) may serve as a stable radiance target for finding one component of the radiometric response, and the mapping from incident radiance to incident irradiance varies across the focal plane depending on the camera's optical configuration (e.g., lens type and configuration). In an exemplary scenario, the 6-DOF camera pose is resolved with high accuracy when the scene resolution module 1107 frees the pose parameters to vary with the medial and radial parameters of the scene being solved. In the present invention, robust camera localization is possible when the observed light field exhibits a consistent, gentle gradient with respect to pose changes (many existing localization methods require a relatively steep gradient).

[0109] Non-camera sensors may also be calibrated in a suitable manner. For example, position, motion, and / or rotation sensors may be calibrated and / or determined so that subsequent movement can be tracked relative to an initial value. A time-of-flight range sensor, for example, may record observations of a scene area to determine an initial pose for a camera observing the same area. The scene resolution module 1107 may use the initial pose estimate to initialize and then refine a model of the scene, including the pose of the time-of-flight sensor and the pose of the camera at various viewpoints. In another example, an inertial navigation system rigidly attached to a camera records an estimate of its pose while the camera observes the scene. When the camera subsequently observes the scene from a different viewpoint, the pose estimate of the inertial navigation system at the new viewpoint may be used to initialize and refine the pose estimate of the camera at the new viewpoint.

[0110] The media modeling module 1207 is concerned with participating (e.g., light-interacting) media types that are characterized primarily, but not exclusively, by their BLIFs 1005. Media modeling manages a library of media that exist in a database. Media modeling also maintains a hierarchy of media types. Both "wood" and "plastic" media types fall into the "dielectric" supertype category, while "copper" falls into the "metal" supertype category. During scene reconstruction, media may be decomposed into one or more of these media types, with likelihoods assigned to each type.

[0111] The observation modeling module 1209 relates to sensed data observations. Observations recorded by cameras and other sensors are represented. Calibration and response models are represented when they relate to a particular observation rather than the sensor itself. An example of this is the distortion of a camera lens, which changes over the course of an image scan due to dynamic focus, aperture, and zoom control. An observation from a camera at a particular viewpoint includes radiance integrated at pixels. The observation model for such an observation includes pixel radiance values, an estimated pose of the camera at the time of the observation, temporal reference information (e.g., image timestamp), calibration values ​​and / or response models specific to the observation (e.g., zoom-dependent lens distortion), and calibration values ​​and / or response models that are observation-independent.

[0112] The chain of calibration and / or response models at different levels of observation locality forms a nested sequence. The nesting order may be embodied using database record references, two-way or one-way memory pointers, or other suitable mechanisms. The nesting order enables traceability of information from various levels of reconstructed scene models back to the original source observations in a manner that allows for a "forensic analysis" of the data flow that produced a given reconstruction result. This information also enables re-invocation of the reconstruction process at various stages using alternative goals, settings, observations, previous models, and the like.

[0113] Kernel modeling module 1211 relates to patterns and / or functions used to detect and / or characterize (extract scene feature signatures from) scene features from recorded observations. The SIFT function, ubiquitous in computer vision, is an exemplary kernel function in the area of ​​feature detection that may be used in exemplary embodiments. Kernel modeling manages a library of kernel functions that reside in a database and are available for feature detection in a given reconstruction operation. Further details regarding scene features are provided below when describing the detection and use of scene features in operation 1821 with reference to FIG. 18B.

[0114] The resolution rules module 1213 is responsible for the order in which the scene resolution module 1107 evaluates (for consistency) the hypothesized values ​​of modeled properties of scene entities. For example, the hypothesized media types shown in FIGS. 8B-8E may be evaluated in parallel or in some prioritized sequential order. Within the evaluation of each hypothesized media type, various numerical ranges for the property's value may be evaluated in parallel or sequentially as well. An example of this would be the different ranges (bins) for the angular degrees of freedom of the normal vector of metallic surfels 821. The desired sequence of sequential and / or parallel evaluation of hypotheses in the previous example is represented by a resolution rules data construct. The resolution rules module 1213 manages a library of resolution rules (contained in a database). The hypothesis evaluation order is determined in part by the hierarchy of modeled media types maintained by the media modeling module 1207. Further details regarding the evaluation of model hypotheses are provided below when describing the scene resolution process 1800 with reference to FIG. 18A.

[0115] The merge rules module 1215 is concerned with merging finer spatial elements to form coarser spatial elements. Media and radii are typical examples of such spatial elements. The merge rules module 1215 manages a library of merge rules (contained in a database). Merge rules have two main aspects. First, they dictate when a merge operation should be performed. In the case of media, merging may be indicated for media whose values ​​for positional, orientational, radiometric, and / or polarimetric properties fall within a certain mutual tolerance across several media. The radiometric and polarimetric properties involved in the merge decision may be properties of the media's BLIF, response light field, emission light field, (total) outgoing light field, and / or incoming light field. In the case of light fields, merging may be indicated for radii whose values ​​for positional, orientational, radiometric, and / or polarimetric properties fall within a certain mutual tolerance across several radii. Second, the merge rules dictate the resulting type of coarser spatial element that is formed when finer spatial elements merge. In the case of media, the hierarchy of modeled media types maintained by the media modeling module 1207 determines, in part, the resulting coarser media. The previous description of the merge rules module 1215 remains valid if the media and light fields in the scene are parameterized using spherical harmonics (higher harmonics merge to form lower harmonics) or any other suitable spatially-based system of functions.

[0116] The operational configuration module 1217 operates to allocate local and remote CPU processing cycles, GPU computing cores, FPGA elements, and / or dedicated computing hardware. The SRE may rely on module 1217 to allocate some types of computations to specific processors, allocate computations with load balancing considerations, etc. In one example, if scene solving 1107 is repeatedly hampered by a lack of updated incident light field information in some scene regions of interest, the operational configuration module 1217 may allocate additional FPGA cores to the outgoing / incoming light field processing module 2005 (shown in FIG. 20 ). In another example, if network bandwidth between the SRE and the cloud suddenly decreases, the operational configuration module 1217 may begin caching cloud database entities in local memory to eliminate long delays in fetching cloud entities on demand, at the expense of using a larger share of local memory.

[0117] FIG. 13 is a block diagram illustrating a light field physics processing module 1301 of an SRE, such as the light field physics processing module 1115 of the SRE 201, according to some demonstrative embodiments. Module 1301 includes operations that model the interaction between a light field and a medium in a scene, as represented by BLIF. The light field physics processing module includes a microfacet (Fresnel) reflectance module 1303, a microfacet integration module 1305, a volumetric scattering module 1307, an emission and absorption module 1309, a polarization modeling module 1311, and a spectral modeling module 1313. An adaptive sampling module 1315 is used to make the modeling problem tractable by focusing available computing resources on the rays with the greatest impact on the reconstruction operation. In some embodiments, the light field physics module uses spatial processing (such as provided by the spatial processing module 1113), including optional dedicated hardware, to achieve light field operations with lower power consumption and / or improved processing throughput. Further details regarding this are provided below with reference to FIG. 20.

[0118] Light field physics models the interactions between a medium and the light that enters (incident) and leaves (emits) it. Such interactions can be complex in real scenes when all or many known phenomena are involved. In an exemplary embodiment, the light field physics module represents these interactions using a simplified "light transport" model. As noted above, FIG. 10 illustrates the model used to represent interactions occurring in a single medium. The emitted light field is emitted by the medium independently of stimulation by incident light. Energy conservation dictates that the total energy of the response light field is less than the total energy of the incident light field.

[0119] In the light transport model, each response radiant (radiance) exiting a medium is a weighted combination of the radiants (radiances) entering that medium or a collection of contributing media (in the case of "light hopping," such as subsurface scattering, where light exits a medium in response to entering a different medium). This combination is usually, but not always, linear. The weights may be a function of the wavelength and polarization state of the incident radiant, which may change over time. As noted above, a radiation term may also be added to account for light emitted by the medium when not stimulated by the incident light. The light transport interaction used in the exemplary embodiment may be represented by the following equation:

[0120]

number

[0121] where: x is a voxel (position element). x' is the voxel that contributes to the outgoing radiance at x. X' are all voxels that contribute to the outgoing radiance at x. ω is the sail of the outgoing radiance. ω' is the sail of incident radiance. x → ω and x ← ω are sails defined by voxel x and sail ω. L(x→ω) is the radiance of the outgoing radial at the sail x→ω. L(x'←ω') is the radiance of the incident radial at the sail x'←ω'. L e (x→ω) is the radiative intensity of the outgoing radiator at the sail x→ω. f l (x→ω, x'←ω') is the light-hopping BLIF that relates the incoming radial at x'←ω' to the outgoing radial at x→ω. dω' is the solid angle subtended by the sail ω'. dx' is the surface area represented by voxel x'.

[0122]

number

[0123] is the global sphere (4π steradians) of the incident sail.

[0124] The previous and following light transport equations do not explicitly show dependencies on wavelength, polarization state, and time, and those skilled in the art will understand that these equations can be extended to model these dependencies.

[0125] As mentioned above, some media exhibit the phenomenon of "light bouncing," where the response light field of a medium depends not only (or even at all) on the medium's incident light field, but also on the light fields entering one or more other media. Such bouncing gives rise to important scene characteristics of some types of media. For example, human skin exhibits significant subsurface scattering, a subclass of general light bouncing. When light bouncing is not modeled, the incident contribution area X' shrinks to a single voxel x, the outer integral disappears, and

[0126]

number

[0127] It will look like this. where: f l (x,ω←ω') is a BLIF that relates the incident radius at x←ω' to the outgoing radius at x→ω, with no optical jumps.

[0128] When the medial is assumed to be of type surfel, the BLIF becomes a conventional Bidirectional Reflectance Distribution Function (BRDF),

[0129]

number

[0130] This becomes: where: f r (x,ω←ω') is the BLIF relating the incident radial at x←ω' to the exit radial at x→ω. n is the surface normal vector at surfel x. (n·ω') is the cosine perspective coefficient that balances the reciprocal coefficient present in the traditional BRDF definition.

[0131]

number

[0132] is a continuous hemisphere (2π steradians) of the incident sail centered on the surfel normal vector.

[0133] When light is transported through a space that is modeled as perforated, radiance is conserved within the sail of propagation (along the path). That is, at a given voxel, the radiance of the incident radiant is equal to the radiance of the corresponding outgoing radiant in the non-perforated medium at the end of the interaction before entering the voxel in question. The medium in perforated space has an identity function, BLIF, where the outgoing light field is equal to the incident light field (and the emitted light field has zero radiance at all sails). For scene models consisting of (non-perforated) medium regions in perforated space, radiance conservation is used to transport light along paths that intersect only the perforated medium. An embodiment using other radiometric units, such as radiant flux, can be formed with equivalent conservation laws.

[0134] The microfacet reflection module 1303 and the microfacet integration module 1305 together model the light field at the boundary between different types of media as dictated by changes in refractive index. This covers the common case of surfels in a scene. Microfacet reflection models the reflectance component of all scattering interactions at each small microfacet, including macroscopic surfels. Each microfacet is modeled as an optically smooth mirror. Microfacets may be opaque or (non-negligibly) transparent. Microfacet reflection uses the well-known Fresnel equation to model the ratio of reflected radiance to incident radiance at various surfaces of interest (as dictated by adaptive sampling 1315) at a voxel of interest. The polarization states of the incident light field (e.g., 1001 in FIG. 10) and the outgoing light field (e.g., 1007 in FIG. 10) are modeled as described below with reference to the polarization modeling module 1311.

[0135] The microfacet integration module 1305 models the total scattering interaction on the statistical distribution of microfacets present in a macroscopically observable surfel. For the reflected component, this involves summing the outgoing radiance over all the microfacets that make up the macroscopic surfel. The camera pixels record such macroscopic radiance.

[0136] The volumetric scattering module 1207 models scattering interactions occurring in transmissive media within a scene, including the generally anisotropic scattering of sails in such media. Volumetric scattering is realized in terms of a scattering phase function or other suitable formula.

[0137] The Polarization Modeling module 1311 models the changing polarization state of light as it propagates through a transmission medium and interacts with surfaces. The Stokes vector formalism is used to represent the radial polarization state of a light field. Stokes vectors are expressed in a specified geometric frame of reference. Polarization readings recorded by different polarimeters must be reconciled by transforming their Stokes vectors into a common frame of reference. This is accomplished by Mueller matrix multiplication, which represents the change in coordinate system. Polarization modeling performs this transformation as needed when observations at multiple pixels and / or viewpoints are compared in reconstruction.

[0138] The polarization modeling module 1311 also models the relationship between the polarization states of the incident and exiting radii at various sails after reflection. This relationship is expressed in the polarimetric Fresnel equations that govern the reflection of s-polarized and p-polarized light at dielectric and metallic surfaces. For light incident on a dielectric or metallic surface in some given medium, the polarimetric Fresnel equations dictate how the reflected and transmitted (refracted) radiances relate to the radiance of the incident light. In conjunction with polarimetric observations of a scene, the polarimetric Fresnel equations enable accurate reconstruction of reflective surfaces that appear featureless when observed with a non-polarimetric camera.

[0139] The Emission and Absorption module 1309 models the emission and absorption of light in a medium. Emission in a medium is represented by a sampled and / or parametric form of the emitted light field 1009. Absorption at a given wavelength is represented by an extinction coefficient or similar quantity, which indicates the attenuation per unit distance of light as it travels through the medium in question.

[0140] The spectral modeling module 1313 models the transport of light of different wavelengths (in different wavebands). The total light field contains light in one or more wavebands. The spectral modeling module 1313 subdivides the light field into several wavebands as needed for the reconstruction operation, given the spectral characteristics of the camera observing the scene. The wavelength dependency on the interaction of the light field with the medium type is expressed in the BLIF of the medium.

[0141] The adaptive sampling module 1315 determines the optimal angular resolution to use to model the radii at various directions for each medium of interest in the scene model. This determination is generally based on the SMA goal for the medium (characteristics), its assumed BLIF, and uncertainty estimates for the radii entering and exiting the medium (characteristics). In a preferred embodiment, the adaptive sampling module 1315 determines the appropriate subdivision (tree level) of the solid angle octree representing the incident and exiting light fields associated with the medium. For example, the BLIF of a shiny (highly reflective) surface has a "specular lobe" that is strictly aligned with the specular direction relative to the incident radii. When the scene solver 1107 begins the process of calculating the consistency between the observed radii and the corresponding modeled radii predicted by the assumed BLIF, the adaptive sampling module 1315 determines that the most significant medium incident radii are those in the opposite specular direction with respect to the assumed surface normal vector. This information can then be used to focus available computing resources on determining modeled radiance values ​​for the determined incident radii. In a preferred embodiment, determining these radiance values ​​is accomplished via a light field operations module 1923 (shown in FIG. 19).

[0142] 14 illustrates an application software process 1400 for using an SRE in a 3D image processing system according to some embodiments. For example, the process 1400 may be performed by the application software 205 described above with respect to the SRE 201.

[0143] After entering process 1400, in operation 1401, the application software may generate and / or provide a script of SRE commands. As discussed above with reference to FIG. 11 , such commands may be implemented in many forms. The command script may be assembled by the application software as a result of user interaction via a user interface, such as user interface 209. For example, as described with respect to process 300 above, user input may be obtained regarding scene information, one or more goals for the restoration, and scan and / or restoration configuration. The application software may configure the goals and operating parameters of the restoration job based on the obtained user input and / or other aspects, such as execution history or stored libraries and data. The command script may also be drawn from another source, such as existing command scripts stored in a database.

[0144] In operation 1403, the application software invokes the SRE on a prepared command script, which initiates a series of actions managed by the SRE, including, but not limited to, plan, scan, solve, and display. If human input and / or action is required, the SRE may notify the application software to prompt the user appropriately. This may be achieved using a callback or another suitable mechanism. An example of such a prompted user action occurs when a user moves a handheld camera to a new viewpoint and records more image data required by the 3D image processing system (e.g., scene resolution module 1107) toward satisfying some restoration goal. An example of prompted user input occurs when the scene resolution function is fed a previous scene model with insufficient information to resolve an ambiguity in the type of media occupying a particular voxel of interest. In this case, the application software may prompt the user to select an alternative media type.

[0145] In operation 1405, it is determined whether the SRE completed the command script without returning an error. If the SRE reached the end of the command script without returning an error, in operation 1407, the application software applies the restoration results in some manner useful in the application domain (e.g., application-dependent use of the restored volumetric scene model). For example, in assessing hail damage to an automobile, the surface restoration of the automobile's hood may be stored in a database for virtual inspection by a human inspector or for detection of hail dents by a machine learning algorithm. In another exemplary application, scanning a scene, such as a kitchen, at regular intervals and scene restoration may be used to detect the condition of perishable goods scattered throughout the space so that new orders can be initiated. For a more detailed description of restoration jobs, see the description of exemplary process 1840 with reference to FIG. 18C below. Process 1400 may end after operation 1407.

[0146] If the SRE returns an error in operation 1405, an optional error handling function (e.g., in the application software) may attempt to compensate by adjusting goals and / or operational settings in operation 1409. As an example of modifying operational settings, if a given SMA goal condition is not met when restoring OOI in a scene, operation 1409 may increase the total processing time and / or number of FPGA cores estimated for the script job. Process 1400 then proceeds to reinvoking the SRE (operation 1403) on the command script after such adjustments. If the error handling function cannot automatically compensate for the returned error in operation 1409, the user may be prompted to edit the command script in operation 1411. This editing may be accomplished through user interaction similar to that employed when initially assembling the script. After the command script has been edited, process 1400 may proceed to operation 1401.

[0147] FIG. 15 is a flowchart of a planning process 1500 according to some example embodiments. Process 1500 may be performed by plan processing module 1103 of SRE 201. Process 1500 parses and executes plan commands toward a reconstruction goal. Plan commands may include plan goals and action settings. In one example, the plan commands initiate reconstruction (including image processing) of the BLIF and voxel-wise medial geometry of the flower and glass bottle of FIG. 1 to within a specified upper uncertainty limit.

[0148] After entering the process 1500, in operation 1501, the plan goal is obtained.

[0149] Plan settings are obtained in operation 1503. The plan settings include operational settings such as those described with reference to operational settings module 1217 above.

[0150] In operation 1505, database information is retrieved from the connected database 209. The database information may include a plan template, a previous scene model, and a camera model.

[0151] In operation 1507, computing resource information is obtained regarding the availability and performance characteristics of local and remote CPUs, GPU computing cores, FPGA elements, data stores (e.g., core memory or disk), and / or specialized computing hardware (e.g., ASICs that perform specific functions to accelerate incoming / outgoing light field processing 2005).

[0152] In operation 1509, a sequence of low-level sub-commands is generated, including scan processing operations (e.g., scan processing module 1105), scene resolution operations (e.g., scene resolution module 1107), scene display (e.g., scene display module 1109), and / or sub-plan operations. Plan goals, operational settings, accessible databases, and available computing resources inform the process of decomposing into sub-commands. Overall plan goal satisfaction may include satisfaction of sub-command goals and overall tests of recovery validity and / or uncertainty estimates at the plan level. In some embodiments, overall plan goal satisfaction may be specified as satisfaction of a predetermined subset of sub-command goals. For further details regarding decomposing plan goals into sub-command sub-goals, see the description of exemplary process 1840 with reference to FIG. 18C below.

[0153] Pre-image processing motion checks and necessary sensor calibrations come before the actual image processing of the relevant scene entities. Scan processing operations are directed to obtaining input data for motion checks and sensor calibrations. Scene solving operations involve computing response models resulting from several calibrations.

[0154] Examples of sensor-oriented operational checks performed at this level are verifying that relevant sensors (e.g., camera 203) are powered on, that they pass basic control and data acquisition tests, and that they have valid current calibrations. In an exemplary solution-oriented operational check, the plan processing operation verifies that inputs to the scene resolution operation are scheduled to be validly present before the resolution operation is invoked. Incomplete or invalid calibrations are scheduled to occur before scene sensing and / or resolution actions that require them. Calibrations that involve physical adjustments to sensors (hard calibrations) must be performed rigorously before sensing operations that rely on them. Calibrations that discover sensor response models (soft calibrations) must be performed rigorously before scene resolution operations that rely on them, but may occur after sensing of scene entities whose sensed data feeds into the scene resolution operation.

[0155] Through a connected database, the plan processing operation can access one or more libraries of subcommand sequence templates. These include generic subcommand sequences that achieve some common or otherwise useful goal. For example, a 360-degree reconstruction of the flower and glass bottle of FIGS. 1 and 5 can be achieved by the following approximate sequence template: Scanning process 1105 roughens the scene light field by guiding an "outward pan" scan 513 at a few positions near the object of interest. Scanning process 1105 then guides a "left-to-right short-baseline" scan 515 of a small region of each different media type that makes up the object of interest (flower, leaf, stem, bottle). Scene solving 1107 then performs a BLIF reconstruction of each media type by maximizing radiometric and polarimetric consistency metrics given the roughened-in model of the scene light field. (Scanning process 1105 performs one or more additional BLIF scans if more observations are needed to reach a specified uncertainty upper bound.) The scanning process then guides a high-resolution orbital (or partial orbital) scan 509 of the object of interest. Scene solving 1107 then recovers the detailed geometry of the object of interest (a BLIF pose-related property) by maximizing a consistency metric similar to that used in BLIF discovery (but maximized over the geometric domain rather than the domain of pose-independent BLIF properties).

[0156] In general, some lightweight scene resolution operations 1107, usually involving spatially localized scene entities, may be performed within the scope of a single scan command. More extensive scene resolution operations 1107 are typically performed outside the scope of a single scan. These more extensive scene resolution operations 1107 typically involve a broader spatiotemporal distribution of scene data, larger amounts of scene data, particularly tighter constraints on reconstruction uncertainty, tuning of existing or new models, etc.

[0157] Following generation of the sub-command sequence, in operation 1511, process 1500 iteratively processes each subsequent group of sub-commands until a plan goal is reached (or an error condition occurs, time and / or resource reservations are exhausted, etc.). After or during each iteration, if the plan goal is reached (in operation 1513), process 1500 outputs the results sought in the plan goal in operation 1515 and stops iterating. If the plan goal is reached in operation 1513, the plan processing command queue is updated in operation 1517. This update in operation 1517 may involve adjusting the goal and / or behavior settings of the sub-commands and / or the plan command itself. New sub-commands may also be introduced, existing sub-commands may be removed, and the order of the sub-commands may be changed by update operation 1517.

[0158] FIG. 16 is a flowchart of a scan processing process 1600 according to some example embodiments. Process 1600 may be executed by plan processing module 1105 of SRE 201. Scan processing process 1600 drives one or more sensors in scanning a scene, for example, as described above with reference to FIG. 11. A scan includes sensed data acquisition within the scene. A scan may also include an optional scene resolution operation (1107) and / or scene display operation (1109). A scan accomplishes some relatively atomic sensing and / or processing goal, as described in the section above regarding the sequencing of plan subcommands with reference to FIG. 15. In particular, scan processing process 1600 manages the acquisition of observations necessary for operational checks and sensor calibration. See the description of example process 1840 with reference to FIG. 18C for examples of operational ranges of individual scans.

[0159] As with plan processing commands, scan processing commands (scan commands) include scan goals along with operational settings. The scan processing process 1600 generates a sequence of sensing operations and, optionally, scene resolution operations (1107) and / or scene display operations (1109). The scan goals, operational settings, and accessible databases inform the generation of this sequence.

[0160] After entering process 1600, scan goals are obtained in operation 1601. For examples of scan goals derived from a plan, see the description of exemplary process 1840 with reference to Figure 18C below.

[0161] Scan settings are obtained in operation 1603. The scan settings include operational settings such as those described with reference to operational settings module 1217 above.

[0162] Database information is obtained in operation 1605. The database information may include a scan template, a previous scene model, and a camera model.

[0163] In operation 1607, a subcommand sequencing function generates a sequence of subcommands as described above. The subcommand sequencing function is largely responsible for the sensing operations, which are described extensively below with reference to FIG. 17. A lightweight scene resolution operation 1107 is also within the scope of the subcommand sequencing function. An example of such a resolution operation 1107 is seen in the feedback-guided acquisition of images for BLIF discovery at step 311 in the functional flow presented in FIG. 3. Scene resolution 1107 is invoked one or more times as a new group of sensed images becomes available. The derivative of BLIF model consistency versus camera pose is used to guide camera motion in a manner that reduces the uncertainty of the recovered model. After scene resolution 1107 reports satisfaction of a specified consistency criterion, feedback is provided to terminate the incremental sensing operation.

[0164] Although slight refinement of existing sensor response models can be achieved through the Scene Solve 1107 subcommand, the full initialization of response models is outside the scope of the Scan Process 1105 and must be handled at a higher process level (e.g., Plan Process 1103).

[0165] In operation 1609, the scan processing module 1105 executes the next group of queued sub-commands. In operation 1611, satisfaction of the scan goals is evaluated 1611, typically with respect to SMA restoration goals for some portion of the imaged scene. If the goal conditions are not met, operation 1615 updates the sub-command queue in a manner that will help reach the scan goals. For an example of update 1615, see the description of exemplary process 1840 with reference to Figure 18C below. If the scan goal conditions are successfully met, the scan results are output in operation 1613.

[0166] 17 is a block diagram of a sensor control module 1700, according to some exemplary embodiments. Sensor control module 1700 may correspond, in some exemplary embodiments, to sensor control module 1111 of SRE 201 described above. In some embodiments, sensor control module 1700 may be used by a scan processing module, such as scan process 1105, to record observations of a scene.

[0167] The acquisition control module 1701 manages the length and number of camera exposures used to sample the scene light field. Multiple exposures are recorded and averaged as needed to mitigate thermal and other time-varying noise in the camera's photosites and readout electronics. Multiple exposures are stacked for synthetic HDR image processing when required by the radiometric dynamic range of the light field. The exposure scheme is dynamically adjusted to account for the flicker period of artificial light sources present in the scene. The acquisition control module 1701 time-synchronizes camera exposures to different polarization filter states when time-multiplexed polarimetry is used. Similar synchronization schemes are used with other optical modalities when time-multiplexed. Examples of this are exposures with multiple aperture widths or multiple spectral (color) filter wavebands. The acquisition control module 1701 also manages the temporal synchronization between different sensors in the 3D image processing system 207. For example, a camera mounted on a UAV can be triggered to expose at the same instant as a tripod-mounted camera observing the same OOI from a different perspective. The two viewpoints may then be input together into a scene resolution operation (eg, performed by the scene resolution module 1107) that recovers the characteristics of the OOI at that moment.

[0168] The analog control module 1703 manages the various gains, offsets, and other controllable settings of the analog sensor electronics.

[0169] The binarization control module 1705 manages the binarization bit depth, digital offset, and other controllable settings of the analog / digital quantization electronics.

[0170] The optics control module 1707 manages adjustable aspects of the camera lens, when available on a given lens. These include zoom (focal length), aperture, and focus. The optics control module 1707 adjusts the zoom setting to achieve the appropriate balance between the size of the FOV and the angular resolution of the light field samples. For example, a relatively wide FOV may be used when roughing in the scene, and then the optics control module 1707 may significantly narrow the FOV for high-resolution image processing of the OOI. The lens aperture is adjusted as needed (e.g., when extracting a relatively weak polarization signal) to balance focal depth of field against recording sufficient radiance. When electromechanical control is not available, the optics control module 1707 may be fully or partially implemented (via a software application user interface) as guidance to a human user.

[0171] The polarimetry control module 1711 manages the scheme used to record the polarization of the light field. When a time-multiplexed scheme is used, the polarimetry control manages the sequence and timing of the polarization filter states. The polarimetry control also manages the number of exposures, multiple exposures to reduce noise, and different nesting orders that are possible when interleaving the polarization sampling states.

[0172] The motion control module 1715 manages the controllable degrees of freedom of the sensor's orientation in space. In one example, the motion control commands an electronic pan / tilt unit to the orientation required for comprehensive sampling of the light field within a region of space. In another example, a UAV is oriented to place a car on a track for hail damage imaging. As with the optics control module 1707, this function can be fully or partially realized as guidance to a human user.

[0173] The data transport control module 1709 manages settings related to the transport of sensed scene data over the data communication layer in the 3D image processing system 200. Data transport deals with transport failure policies (e.g., retransmission of dropped image frames), relative priorities between different sensors, data chunk sizes, etc.

[0174] The proprioceptive sensor control module 1713 interfaces with an inertial navigation system, inertial measurement unit, or other such sensor that provides information about position and / or orientation and / or motion within a scene. The proprioceptive sensor control synchronizes proprioceptive sensor sampling as needed with sampling performed by cameras and other associated sensors.

[0175] 18A is a flowchart of a scene resolution process 1800, according to some example embodiments. Process 1800 may be performed by scene resolution module 1107 of SRE 201. Scene resolution process 1800 may create and / or refine a recovered scene model as directed by a scene resolution command (solve command). Scene resolution includes the major steps of initializing an assumed scene model in operation 1811 and then iteratively updating the model in operation 1819 until the goal specified in the solve command is reached in operation 1815. After the goal is reached, the resulting scene model is output in operation 1817.

[0176] After entering process 1800, one or more goals for the scene are obtained in operation 1801. Each goal may be referred to as a solution goal.

[0177] In operation 1803, process 1800 obtains solution settings. The solution settings include operational settings such as those described above with reference to operational settings module 1217. The solution settings may also include information regarding the processing capabilities of the underlying mathematical optimizer (e.g., the maximum number of allowed iterations of the optimizer) and the degree of spatial domain context to be considered in the solution operation (e.g., weak expected gradients in the properties of media regions that are assumed to be homogeneous).

[0178] A database is accessed in operation 1805. Process 1800 may load relevant scene observations from the appropriate database in operation 1807 and load a previous scene model from the appropriate database in operation 1809. The loaded observations are compared to predicted results resulting from the assumed model in operation 1813. The loaded previous model is used to initialize the assumed scene model in operation 1811.

[0179] Before engaging in iterative updates to the hypothesized scene model (in operation 1819), process 1800 initializes the hypothesized scene model to a starting state (configuration) in operation 1811. This initialization depends on the retrieved solution goal and one or more (a priori) scene models retrieved from the relevant database (in operation 1809). In some restoration scenarios, multiple distinct hypotheses are feasible in the initialization phase. In that case, the update flow (not necessarily its first iteration) explores multiple hypotheses. See the discussion regarding the creation and elimination of alternative possible scene models (in operation 1833) with reference to Figure 18B below.

[0180] The iterative updating in operation 1819 adjusts the hypothesized scene model to maximize consistency with observations of the real scene. Scene model updating in operation 1819 is further described below with reference to FIG. 18B . In the case of alignment between existing scene models, process 1800 instead calculates 1813 the consistency between the models to be adjusted. In an example of such model alignment, a model of a car hood that has been damaged by hail is aligned against a model of the same hood before the damage occurred. The consistency calculation in operation 1813 is based on the deviation between the intrinsic medial properties (e.g., BRDF, normal vector) of the hood model in this example alignment, rather than the deviation between the emitted light fields of the two models.

[0181] Time can be naturally included in the scene model. Thus, the temporal dynamics of an image processing scene can be recovered. This is possible when a time reference (e.g., a timestamp) is provided for the observations fed into the solving operation. In one exemplary scenario, a car drives through a surrounding scene. With access to time-stamped images from a camera observing the car (and optionally, with access to a model of the car's motion), the car's physical properties can be recovered. In another example, the deforming surface of a human face is recovered at multiple time points while changing expression from neutral to smiling. Scene solving 1107 can generally recover scene models where the scene media configuration (including BLIF) changes over time, provided it has observations from a sufficiently large number of spatial and temporal observation viewpoints per spatiotemporal region to be recovered (e.g., voxels at multiple instantaneous times). Recovery can also be performed under a changing light field when a model (or sampling) of the light field's dynamics is available.

[0182] At each scene-solving iteration, a metric indicating the degree of consistency between the assumed model and the observations and / or other scene models being matched during calibration is computed in operation 1813. The consistency metric may include a heterogeneous combination of model parameters. For example, the surface normal vector direction, refractive index, and spherical harmonic representation of the local light field at the assumed surfels may be input together into a function that computes the consistency metric. The consistency metric may also include multiple modalities (types) of sensed data observations. Polarimetric radiance (Stokes vector) images, sensor pose estimates from an inertial navigation system, and surface extent estimates from a time-of-flight sensor may be input together into the consistency function.

[0183] In a typical everyday reconstruction scenario, model consistency is calculated for each individual voxel. This is achieved by combining per-observation measures of deviation between a voxel's predicted emitted light field 1007 (e.g., in a projected digital image) and the corresponding observation of the actual light field (e.g., in a generated digital image) across multiple observations of the voxel. The consistency calculation in operation 1813 may use any suitable method for combining the per-observation deviations, including but not limited to, the sum of the squares of the individual deviations.

[0184] In the example of FIG. 8A , one or more cameras image voxel 801 from multiple viewpoints. Each viewpoint 803, 805, 807, 809, 811, and 813 records the exit radiance value (and polarimetric properties) for a particular radius of the light field emanating from voxel 801. A medium is assumed to occupy voxel 801, and scene solver 1107 must calculate model consistency against one or more BLIF 1005 assumptions (the light field model is held constant in this example). For each viewpoint, a given BLIF 1005 assumption predicts (models) the exit radiance values ​​expected to be observed by each viewpoint by operating on the incident light field 1001. Multi-view consistency between the assumed BLIF 1005 assumptions and the camera observations as described above is then calculated. The evaluated BLIF 1005 assumptions may fall into two or more distinct classes (e.g., wood, glass, or metal).

[0185] At this algorithmic level, the scene resolution module 1107 operation cannot explicitly deal with the geometric properties of the voxels 801. All physical information input to the consistency calculation 1813 is contained within the BLIF of the voxels 801 and the surrounding light field. Classification as surfels, bulk media, etc. does not affect the fundamental mechanism of computing the radiometric (and polarimetric) consistency of the assumed BLIF.

[0186] Figure 18B is a flowchart of a hypothesized scene update process 1819, according to some exemplary embodiments. In some exemplary embodiments, the process 1819 may be performed by the update assumed scene model operation 1819 described with respect to Figure 18A. The hypothesized scene update process 1819 represents details of the update 1819 operation performed on the hypothesized scene model at each scene solution iteration. The scene update 1819 includes one or more of the internal functions shown in Figure 18B. The internal functions may be performed in any suitable order.

[0187] Process 1819 detects and uses observed features of the scene in operation 1821. Such features are typically sparsely represented in observations of the scene (i.e., in the case of image processing, features are significantly fewer than pixels). Features are detectable in a single observation (e.g., an image from a single viewpoint). Feature detection and characterization are expected to have significantly lower computational complexity than fully physics-based scene reconstruction. Features may have unique signatures that are resilient to viewpoint changes, useful for inferring the structure of the scene, particularly the sensor viewpoint pose at the time the observation was recorded. In some examples, localized (point-like) features, strict with respect to their position within the field of total radiance (emanating from the field of the scene medium), are used for 6-DOF viewpoint localization.

[0188] When polarimetric image observations are input to the scene resolution module 1107, two additional types of features become available. The first are point-like features within the field of polarimetric radiance. In many exemplary scenarios, these point-like features of polarimetric radiance measurably increase the total number of features detected compared to point-like features of total radiance alone. In some examples, point-like features of polarimetric radiance may result from gradients formed from adjacent surfel normal vectors and may be used as localized feature descriptors to label corresponding features on the two polarimetric images. The second type of feature that becomes available in the polarimetric images is planar features within the field of polarimetric radiance. Where point-like features are said to be positional features (they convey information about relative position), planar features may be said to be orientational features (they convey information about relative orientation). In a prime example, a user of a 3D imaging system performs a polarimetric scan of an empty wall in a typical room. The wall completely fills the field of view of the imaging polarimeter. Even if no position feature is detected, the polarimetric radiance signature of the wall itself is a strong feature of orientation. In this example, the polarimetric feature of orientation is used to estimate the orientation of the polarimeter with respect to the wall. The estimated orientation is fed into the general scene model adjustment module 1823 operation by itself or in conjunction with position and / or other orientation information. Both of the above polarimetry-enabled feature types can be detected on featureless reflective surfaces when observed with a non-polarimetric camera.

[0189] In operation 1823, process 1819 may adjust the scene model in a way that enhances some metric indicative of goal satisfaction. In a common example, least-squares fit between the assumed model and the observations is used as the sole metric. The adjustment may be guided by the derivative of the satisfaction metric (as a function over the model solution space). The adjustment may proceed by derivative-free stochastic and / or heuristic methods, such as pattern search, random search, and / or genetic algorithms. In some cases, machine learning algorithms may guide the adjustment. Derivative-free methods may be particularly beneficial in real-world scenarios with sampled and / or noisy observations (the observed data exhibits jaggedness and / or cliffs). For an example of derivative-free adjustment via hierarchical parameter search, see the description of process 1880 with reference to FIG. 18D below.

[0190] A suitable model optimizer may be used to implement the above tuning scheme. In some embodiments, spatial processing operations (e.g., those described with respect to spatial processing module 1113) may be used to quickly explore the solution space under a suitable optimization framework. In addition to tuning the assumed scene model itself, subgoals (of the solution goal) and solution settings may also be adjusted.

[0191] Uncertainty estimates for the various degrees of freedom of the scene model are updated for each scene region as appropriate in operation 1825. The observation status of a scene region is updated as appropriate in operation 1827. The observation status of a region indicates whether light field information from that region has been recorded by the camera involved in the reconstruction. A positive observation status necessarily indicates that a direct line of sight observation (through a given medium) has occurred. A positive observation status indicates that a non-negligible amount of radiance observed by the camera can be traced back to the region in question via a series of BLIF interactions within a region already reconstructed with high consistency. Topological coverage information for the reconstruction and / or observation coverage of the scene region is updated in operation 1829.

[0192] Process 1819 may, in operation 1831, divide and / or merge the hypothesized scene model into multiple sub-scenes. This is typically done to focus available computing resources on areas of high interest (high-interest areas are divided into their own sub-scenes). The division may be in space and / or time. Conversely, two or more such sub-scenes are merged into a unified scene (a super-scene of the constituent sub-scenes). A bundle adjustment optimizer or some other suitable optimization method may be used to adjust the sub-scenes based on their mutual consistency.

[0193] Process 1819 may generate and / or eliminate alternative possible scene models in operation 1833. Alternative models may be explored sequentially and / or in parallel. At suitable junctures in the iterative solution process, alternative models that become highly inconsistent with related observations and / or other models (when adjusted) may be eliminated. In the example of FIGS. 8A-8E, four bottom views show hypothesized media to which voxel 801 may be solved. These hypotheses define distinct BLIF alternative models. The four BLIFs differ in type (structure). This distinguishes between BLIF hypotheses that are the same type (e.g., "metal surfels") but have different values ​​for parameters specific to that type (e.g., two hypothesized metal surfels that differ only in normal vector direction). As noted above, the four distinct hypotheses may be explored sequentially and / or in parallel by different mathematical optimization routines in the solve 1107 function.

[0194] The above scene resolution operation can be expressed in a concise mathematical form. Using the symbols introduced above with reference to the light field physics 1301 function, a general scene resolution operation in a single medium is:

[0195]

number

[0196] and where:

[0197]

number

[0198] is the total number of sails that contribute to the response light field sail x → ω

[0199]

number

[0200] This is BLIF, which has light bounce and is applied to the image.

[0201]

number

[0202] Sail

[0203]

number

[0204] is the radiance of each radii at

[0205]

number

[0206] is BLIF f l The outgoing radial and incoming light fields in the sail x→ω predicted by

[0207]

number

[0208] is the radiance of the L observed is the radiance recorded by a camera observing a single voxel x or multiple voxels X. error(L predicted -L observed ) is a function that includes a robustification and regularization mechanism that generates an inverse consistency measure between the predicted and observed radii of the scene light field. An uncertainty-based weighting factor is applied to the difference (deviation, residual) between the predicted and observed radiances.

[0209] When applied to volumetric regions (multiple media), the solve operation:

[0210]

number

[0211] is expanded as follows. where

[0212]

number

[0213] is BLIF f l The outgoing radial and incoming light fields for all sails x→ω predicted by

[0214]

number

[0215] is the radiance of the X is the observed area of ​​the medi-el.

[0216] In the useful case where the incident light field is estimated with high confidence, the solve operation can be constrained to solve for BLIF, but keeping the incident light field constant (illustrated for a single medium, with hops).

[0217]

number

[0218] Conversely, when the BLIF is estimated with high confidence, the solve operation can be constrained to solve for the incident light field, but keep the BLIF constant (illustrated for a single medium, with jumps).

[0219]

number

[0220] The sub-operations involved in computing the individual error contributions and determining how to perturb the model at each argmin iteration are complex and are explained in the text before the equations for scene solving 1107.

[0221] FIG. 18C is a flowchart of a goal-driven SRE job, according to some example embodiments, for the kitchen scene restoration scenario presented in FIGS. 1, 5, and 7. FIG. 18C complements the general flowchart of FIG. 3 by providing a detailed description of the operations and decision points that the SRE follows in the restoration job. The example job goals in FIG. 18C are also narrowly specified to define numerical, quantitative goal criteria in the following description. Descriptions of common "boilerplate" operations, such as motion checks and standard camera calibrations, are provided in previous sections and will not be repeated here.

[0222] The restoration process 1840 begins in this exemplary scenario, beginning with operation 1841, in which the application software 205 creates a script for a job (forming a script of SRE commands) whose goal is to restore the petals of a daffodil flower to a desired spatial resolution and level of scene model accuracy (SMA). The desired spatial resolution may be specified in terms of the angular resolution for one or more of the viewpoints at which the images were recorded. In this example, the goal SMA is specified in terms of the mean square error (MSE) of the light field radii emanating from those medial to which the restored model is assumed to be of type "daffodil petal." A polarimeter capable of characterizing the complete polarization ellipse is used in this example. The polarization ellipse is parameterized as the Stokes vector [S0, S1, S2, S3], and the MSE is

[0223]

number

[0224] It takes the form (see equations [1] to [7]). where: s predicted is the Stokes vector of the radiance of the assumed outgoing radial at x→ω, and BLIF f l and the incident light field

[0225]

number

[0226] is assumed. The model of the incident light field may include polarimetric properties. "Light bouncing" is not modeled in this example scenario. s observed is the Stokes vector of radiance recorded by a polarimeter observing the corresponding medium in a real scene. The error is a function that provides an inverse consistency measure between the predicted and observed radials. In the polarimetric case of this example, this function can be realized as the squared vector norm (sum of squares of components) of the deviation (difference) between the predicted and observed Stokes vectors.

[0227] The desired SMA in this example scenario is a polarimetric radiance RMSE (square root of MSE) of 5% or less calculated across the radii emanating from the petal media of a daffodil in the assumed scene model relative to the radiance of the corresponding observed radii, expressed in watts per steradian per square meter. Equation [8] above yields the MSE for a single assumed media. The MSE for a collection of media, such as the assumed arrangement of media making up the petals, is calculated by summing the predicted versus observed polarimetric radiance deviations across the media in the collection.

[0228] Also, in operation 1841, the application software 205 specifies a database whose contents include a polarimetric camera model of the polarimeter used in image processing, an a priori model of a typical kitchen in a single-family home during daylight hours, and a prior model of daffodil petals. The prior model, including an approximate BLIF of a typical daffodil petal, is adjusted in rough-in operations 1847 and 1851 in a manner that maximizes its initial match to the real scene. An exemplary adjustment operation inputs observed radiance values ​​for the sphere of exit radii as initial values ​​for the medial in the prior model. The observed radiance values ​​are measured by the polarimeter performing the rough-in image processing in the real scene.

[0229] In operation 1843, the plan processing module 1103 generates a sequence of scan and solve commands toward the 5% RMSE goal specified in the job script. Given the daffodil petal OOI recovery goal, the plan processing 1103 retrieves a suitable plan template from the connected database. The template, in this example, is the "General Quotidian OOI Plan." The template specifies a sequence of sub-operations 1847, 1849, 1851, 1853, 1855, and 1857. Based on the overall job goal, the template also specifies sub-goals for some of the operations. These sub-goals are detailed in the following paragraphs and determine the conditional logic operations 1859, 1865, 1867, 1869, and 1873, as well as the associated iterative adjustment operations 1861, 1863, 1871, and 1875 that are performed along some of the conditional flow paths. The SMA targets in the sub-goals are customized according to the 5% RMSE goal of the overall job.

[0230] In operation 1845, plan process 1103 begins an iterative cycle of sub-steps. In operation 1847, scan process 1105 guides an initial "outward pan" sequence of rough-in images of the scene (e.g., camera pose in item 513 of FIG. 5 ) and then inputs initial values ​​for the exit point light field of the scene medial as described above with reference to the previous model initialization. The sub-goal of operation 1847 in this example is to estimate (reconstruct) the incident light field radius at 3° angular resolution for all medial within the hypothesized region occupied by the daffodil petals. Scan process 1105 determines the 3° goal for incident radius resolution based on a mathematical model (e.g., simulated reconstructions) and / or historical reconstruction runs with (estimated) BLIFs of daffodil petals contained in previous models in a database.

[0231] In operation 1849, the scan process 1105 guides a cyclic sequence of BLIF rough-in image processing of daffodil petals (e.g., camera pose of item 515 in FIG. 5). The BLIF rough-in image processing is guided toward regions of the petals that are assumed to be homogeneous in media and / or spatial configuration (shape). In the exemplary case of petals, homogeneity in this regard may require that the imaged "petal patches" have approximately constant color and thickness, and that the patches have negligibly small or approximately constant curvature. The scan process 1105 then instructs the scene solver 1107 to estimate BLIF parameters, as conveyed in Equation [6], to be applied to the media of the petal patches (1851). A sub-SMA goal of 3% relative RMSE is set for the BLIF solving operation. Similar to the 3° radial goal in the previous paragraph, the 3% BLIF recovery goal is determined based on a mathematical model and / or historical execution data. If the scene solve 1107 fails 1865 to meet the 3% goal, the scan process 1105 updates 1863 its internal command queue to image the petal patch from additional viewpoints and / or with increased image resolution (e.g., zoom in to reduce the field of view). If the 3% goal is not met and the scan process 1105 has exhausted some image processing and / or processing power, the plan process 1103 updates 1861 its internal command queue to guide additional "outward pan" image processing sequences and / or more accurately rough in the scene by recovering the incident light field (at the petal patch) with a finer angular resolution in the radius. If the 3% goal is met, control proceeds to operation 1853.

[0232] In operation 1853, the scan process 1105 guides the sequence of detailed image processing of the daffodil petals (e.g., camera pose 509 in FIG. 5 ). The scan process 1105 then instructs the scene solver 1107 to refine 1855 the roughened-in reconstruction of the petal medial. A subgoal of 8% relative RMSE is set for the refinement operation. The 8% goal applies to all petal medial, not just the homogeneous patches (which are more easily reconstructed) that were assigned the 3% goal in the previous description of operation 1851. The 8% goal is determined based on mathematical models and / or historical data on reconstruction runs. The SRE predicts that an 8% RMSE goal when solving for the petal medial, while holding the BLIF (BRDF portion of the BLIF) and light field parameters constant, will reliably result in a 5% final RMSE when the BLIF and / or light field parameters are allowed to “float” in the final whole-scene refinement operation 1857. If the scene solve 1107 fails 1869 to meet the 8% goal, the scan process 1105 updates 1871 its internal command queue to image petal patches from additional viewpoints and / or at increased resolution. If the 8% goal is not met and the scan process 1105 has exhausted some image processing and / or processing power, control proceeds to operation 1861 as described above for BLIF rough-in operations 1849 and 1851. If the 8% goal is met, control proceeds to operation 1857.

[0233] In operation 1857, scene solve 1107 performs a full refinement of the scene model (e.g., large-scale bundle adjustment in some embodiments). Solve operation 1857 typically involves more degrees of freedom than rough-in solve operations 1851 and 1855. BLIF parameters (the BRDF portion of BLIF) and pose parameters (e.g., petal surfel normal vectors) are, for example, allowed to vary simultaneously. In some scenarios, portions of the scene light field, parameterized using standard radial and / or other basis functions (e.g., spherical harmonics), are also allowed to vary in final refinement operation 1857. If operation 1857 fails (1859) to meet the control job's SMA goal of 5% RMSE on petal medial, then the plan process updates (1875) its internal command queue to acquire additional petal images and / or updates (1861) its internal command queue as described above with reference to operations 1849 and 1851. If the 5% goal condition is met, the process 1840 completes normally and the application software uses the reconstructed model of the daffodil petals for any suitable purpose.

[0234] 18D is a flowchart, according to some example embodiments, of the operations involved in solving in detail for the daffodil petal medial as described with reference to operation 1855 in the description of 18C above. A hierarchical derivative-free solving method is used in this example scenario. Process 1880 provides a BLIF solution for a single medial (and / or its entrance and / or exit radials) at some desired spatial resolution. In a preferred embodiment, instances of process 1880 are run in parallel for many hypothesized medials (voxels) in the daffodil petal OOI region.

[0235] In operation 1881, a hierarchical BLIF refinement problem is defined. The assumed medial BLIF parameters are initialized to the values ​​found in BLIF rough-in solution operation 1851. Also initialized is a traversable tree structure (e.g., a binary tree for each BLIF parameter) representing a hierarchical subdivision of the numerical range of each BLIF parameter. In operation 1883, a minimum starting level (depth) of this tree is set, thereby preventing premature failure of the tree traversal due to overly coarse quantization of the parameter range. In operation 1885, parallel traversals of the tree begin at each of the minimum-depth starting nodes determined in operation 1883. Each branch of the tree is traversed independently of other branches. Massive parallelism can be achieved in embodiments using, for example, many simple FPGA computing elements to perform the traversal and node processing.

[0236] In operation 1887, the consistency of the BLIF assumptions is evaluated at each tree node. Parallel traversals evaluate the consistency of each node independently of other nodes. The consistency evaluation proceeds according to equation [8]. The consistency evaluation in operation 1887 requires that certain robustness criteria be met (1893). In this example, one such criterion is that the daffodil petal observations produced by the scanning process 1105 provide sufficient coverage (e.g., 3° ​​angular resolution in a 45° cone centered on the assumed BLIF's primary specular lobe). Another exemplary criterion is that the modeled scene radii entering (incident) the medial satisfy similar resolution and coverage requirements. Both the robustness criteria for the observed radii and the modeled incident radii are highly dependent on the assumed BLIF (e.g., the geometry of the specular lobe). When a criterion is not met (1893), operation 1891 requests additional necessary image processing viewpoints from the scanning process 1105 module. When so indicated, operation 1889 requests the availability of additional incident radial information from the scene modeling 1203 module. These requests may take the form of adding entries to a priority queue. Requests for additional incident radial information generally start a chain of light transport requests that are serviced by the light field operations module 1923.

[0237] After the robustness criteria are met (1893), if the consistency criteria are not met (1895), the tree traversal stops (1897) and the BLIF hypothesis for the node is discarded (1897). If the consistency criteria are met (1895), the BLIF hypothesis for the node is added to a list of valid candidate BLIFs. After the tree is exhaustively traversed, the candidate with the greatest consistency (lowest modeling error) is output in operation 1899 as the most likely BLIF for the media. If the candidate list is empty (none of the tree nodes met the consistency criteria), the voxel fails to meet the hypothesis that it is of media type "daffodil petal."

[0238] 19 is a block diagram illustrating a spatial processing module 1113 of an SRE, according to some example embodiments. The spatial processing module 1113 may include a set operation module 1903, a geometry module 1905, a generation module 1907, an image generation module 1909, a filtering module 1911, a surface extraction module 1913, a morphological operation module 1915, a connection module 1917, a mass property module 1919, a registration module 1921, and a light field operation module 1923. The operations of the spatial processing modules 1903, 1905, 1907, 1911, 1913, 1915, 1917, 1919 are generally known, and those skilled in the art will understand that they may be implemented in many ways, including software and hardware.

[0239] In some exemplary embodiments, media in a scene is represented using an octree. A basic implementation is described in U.S. Pat. No. 4,694,404 (see, e.g., Figures 1a, 1b, 1c, and 2), which is incorporated herein by reference. Each node in the octree can have any number of associated property values. Storage methods are presented in U.S. Pat. No. 4,694,404 (including, e.g., Figure 17). It will be understood by those skilled in the art that equivalent representations and storage formats may be used.

[0240] Nodes, including processing operations, may be added to the bottom of the octree to increase resolution, and to the top to increase the size of the scene being represented. Separate octrees may be combined using the UNION Boolean set operation to represent larger data sets. Thus, a collection of spatial information that can be treated as a group (e.g., 3D information from one scan or a collection of scans) can be represented in a single octree and processed as part of the combined set of octrees. Changes may be applied independently to individual octrees within a larger collection of octrees (e.g., fine-tuning alignment as processing progresses). Those skilled in the art will understand that they may be combined in multiple ways.

[0241] Octrees can also be used as masks to remove some portion of another octree or set of octrees using the INTERSECT or SUBTRACT Boolean set operations or other operations. Half-space octrees (the entire space on one side of a plane) can be geometrically defined and generated as needed to remove half of a sphere, for example, the hemisphere on the negative side of the surface normal vector at a point on the surface. The original data set is not modified.

[0242] Images are represented in the exemplary embodiment in one of two ways. They may be a conventional array of pixels as described in U.S. Pat. No. 4,694,404 (including, for example, Figures 3a and 3b) or a quadtree. Those skilled in the art will recognize that there are alternative equivalent representation and storage methods. A quadtree that includes all nodes down to a particular level (but not the nodes below it) is also called a pyramid.

[0243] Exemplary embodiments represent light emanating from a light source, its passage through space, light incident on a medium, its interaction with a medium, including light reflected from a surface, and light emanating from a surface or medium. A "solid angle octree," or SAO, is used for this purpose in some exemplary embodiments. In its basic form, the SAO uses an octree representing a hollow sphere to model directions radiating outward from the center of the sphere, which is also the center of the SAO's octree region. The SAO represents light entering or exiting that point. This may then represent light entering or exiting a surrounding region. This center may coincide with a point in another dataset (e.g., a point in a volumetric region of space, such as the center of an octree node), but their coordinate systems are not necessarily aligned.

[0244] A solid angle is represented by a node in the octree that intersects with the surface of a sphere centered at its center point. Here, intersection can be defined in multiple ways. Without loss of generality, the sphere is considered to have unit radius. This is illustrated in Figure 21, which is shown in 2D. Point 2100 is the center, vector 2101 is the X-axis of the SAO, and 2102 is the Y-axis (the Z-axis is not shown). Circle 2103 is the 2D shape that corresponds to the unit sphere. The root node of the SAO is a square 2104 (a cube in 3D). The SAO domain is divided into four quadrants (octants in 3D), namely node a and its three sibling nodes in the figure (or seven for a total of eight in 3D). Point 2105 is the node center at this level. This is then subdivided into child nodes at the next level (nodes b, c, and f are examples). Point 2106 is the node center at this level of subdivision. The subdivision continues to any level necessary as long as the nodes intersect the circle (sphere in 3D). Nodes a through e (and other nodes) intersect the circle (sphere in 3D) and are P (parent) nodes. Non-intersecting nodes such as f are E (empty) nodes.

[0245] Nodes that intersect with the circle through which the ray passes both it and the central point (or in some formulations, the area around that point) remain P nodes and are given a set of properties that characterize that ray. Nodes that are not intersected by the ray (or are not of interest for some other reason) are set to E. SAOs can be constructed bottom-up from sampled information (e.g., images), or generated top-down from some mathematical model, or constructed on demand (e.g., from a projected quadtree).

[0246] Although there may be terminal F (full) nodes, operations in this invention typically operate on P nodes, creating and, if necessary, deleting them. Lower-level nodes contain solid angles represented at higher resolution. Often, some measure of the difference in property values ​​contained in child nodes is stored in the parent node (e.g., minimum and maximum, average, variance). In this way, the immediate needs of the operational algorithm can be determined on the fly when a child node needs to be accessed and processed. In some cases, lower-level nodes in octrees, SAOs, and quadtrees are generated on the fly during processing. This can be used to support adaptive processing, where future operations depend on the results of previous operations when performing some task. In some cases, the step of accessing or generating new information, such as a lower level of such a structure for higher resolution, will take time or be performed by a different process, perhaps remotely, such as in the cloud. In such cases, the operational process may generate a request message for new information to be accessed or generated for a later processing operation while the current operation uses the available information. Such requests may have priority. Such requests are then analyzed and, on a priority basis, they are communicated to the appropriate processing units.

[0247] In an exemplary embodiment, individual rays may be represented. For example, a node that intersects with a unit sphere may contain a property with high precision points or a list of points on the sphere. These may be used to determine children that contain ray intersections when lower-level nodes are created, at whatever resolution is required, possibly on the fly.

[0248] Starting with an octant of the SAO, the status of a node (that intersects with the unit sphere) can be determined by calculating the distance (d) from the origin to the near and far corners of the node. There are eight distinct pairs of near / far corners within the eight octants of this region. The value d2 = dx 2 +dy 2 +dz 2 can be used because a sphere has unit radius. In one formula, a node intersects with the unit sphere if the value for the near corner is <1 and for the far corner is ≥1. If it does not intersect with the sphere, it is an E-node. Another method is to calculate the d2 value for the center of the node and then compare it to the minimum and maximum radius2 values. The distance increment for a PUSH operation is a power of 2, so dx 2 , dy 2 , and dz 2 The value can be calculated efficiently with shift and add operations (described below with reference to Figure 25A).

[0249] In a preferred embodiment, a more complex formula is used to select P nodes for better node distribution (e.g., more equal surface area on the unit sphere). Node density can also be lowered to reduce or prevent overlap. For example, they can be constructed so that there is only one node at a certain level of solid angle. In addition, in some methods, just four child nodes may be used, which simplifies the distribution of lighting by dividing by four. Depending on the node arrangement, gaps may be allowed between non-E nodes, including lighting samples and perhaps other characteristics. SAOs can be pre-generated and stored as templates. Those skilled in the art will appreciate that numerous other methods can be employed. (FIG. 22 shows a 2D example of an SAO with rays.)

[0250] Octrees and quadtrees are combined using the Boolean set operations UNION, INTERSECTION, DIFFERENCE (or SUBTRACTION), and NEGATION in set operation processing module 1903. The handling of property information is specified in such operations by calling functions.

[0251] In the geometric processing module 1905, geometric transformations are performed on the octrees and quadtrees (eg, rotation, translation, scaling, skewing).

[0252] Octrees and quadtrees are generated from other forms of geometric representations in a generation processing module 1907. This implements a mechanism for determining the status of nodes (E, P, and so on) in the original 3D representation. This can be done all at once to generate an octree, or incrementally as needed.

[0253] The image generation processing module 1909 computes the image of the octree as either a traditional array of pixels or a quadtree. Figure 36A shows the geometry of image generation in 2D. The display coordinate system is the X axis 3601 and the Z axis 3603, which intersect at the origin 3605. The Y axis is not shown. The external display screen 3607 is on the X axis (in 3D, the plane formed by the X and Y axes). The observer 3609 is located on the Z axis with the viewpoint 3611. The size of the external display screen 3607 in X is length 3613, measured in pixels, but may be of any size.

[0254] 36B shows an internal view of a geometry of size 3615 that is scaled to the next highest power of two in size in X of external display screen 3613 to define display screen 3617, used internally. For example, if external display screen 3607 has a width 3613 of 1000 pixels, then display screen 3617 has a size 3615 of 1024. Those skilled in the art will appreciate that alternative methods may be used to represent this geometry.

[0255] Figure 36C shows the orthogonal projection of octree node center 3618 from viewpoint 3611 onto display screen 3617. Node center 3618 has an x-value 3621 and a z-value 3619, which projects to point 3627 on display screen 3617 with a value on the X-axis of x' 3625. Because projection 3623 is orthogonal, the value of x' 3625 is equal to x 3621.

[0256] Figure 37 shows the geometry of octree node centers in 2D and how geometric transformations are performed, for example, to calculate node center x value 3709 from the node center of its parent node, node center 3707. This is shown for general rotation in the XY plane, where axis 3721 is the X axis and axis 3723 is the Y axis. This mechanism is used for general 3D rotations for image generation, where node center x and y coordinates are projected onto a display screen and the x value can be used as a measure of depth from the viewer in the direction of the screen. In addition to display, this is used for general geometric transformations in 3D. A parent node 3701, centered at node center point 3707, is typically subdivided into its four children (eight in 3D) as a result of a PUSH operation to one of the children. The coordinate system of the octree containing node 3701 is the IJK coordinate system, distinguishing it from an XYZ coordinate system. The K axis is not shown. The I direction is the I axis 3703, and the J direction is the J axis 3705. A move from node center 3707 to the node center of its child is a combination of a move by two vectors (three in 3D) in the direction of an axis in the octree coordinate system. In this case, the vectors are i vector 3711 in the I direction and j vector 3713 in the J direction. In 3D, a third vector, k, would be used in the K direction but is not shown here. A move to one of the four children (eight in 3D) involves adding or subtracting each of the vectors. In 3D, the selection of addition or subtraction is determined by three bits of the three-bit child number. For example, a 1 bit can indicate addition to the associated i, j, or k vector, and a 0 bit can indicate subtraction. As shown, the move is in the positive direction for both i and j, moving to the child node whose node center is at 3709.

[0257] The geometric operation shown is a rotation of node centers in the octree IJ coordinate system into a rotated XY coordinate system with an X axis 3721 and a Y axis 3723. To accomplish this, the distance from the parent node center to the child node center in the X and Y coordinate system (and Z in 3D) is pre-computed for vectors i and j (and k in 3D). For the case shown, the translation in the X direction for the i vector 3711 is distance x i 3725. For the j vector, this distance is the x j distance 3727. In this case, x i is the positive distance and x j is in the negative direction. To calculate the x value of the child node center 3709, these values, or x i and x j , are added to the x value of the parent node center 3707. Similarly, x translations to the centers of other child nodes will be various combinations of additions and subtractions of x i and x j (and x k in 3D). Similar values ​​are calculated for the distance in the Y (and Z in 3D) direction from the parent node center, which gives a general 3D rotation. Operations such as the geometric operations of translation and scaling can be implemented in a similar manner by one skilled in the art.

[0258] Importantly, the length of the i vector 3711 used to move in X from the parent node center 3707 to the child node center 3709 is exactly twice the length of the vector in the I direction used to move to the child at the next level of subdivision, such as node center 3729. This is true for all levels of subdivision, and is true for the j vector 3713 (and in 3D, the k vector) as well as Y and Z. The difference values ​​xi 3725, xj 2727, and other values ​​can thus be calculated once for a region of the octree for a particular geometric operation and then divided by 2 for each additional subdivision (e.g., PUSH). Then, in a POP, the center can be restored by multiplying the difference by 2 and reversing the addition or subtraction when returning to the parent node. These values ​​can then be put into, for example, a shift register and shifted right or left as needed. These operations can also be achieved in other ways, such as precomputing the shifted values; precomputing eight different sums for the eight children, using a single adder for the dimensions, and using a stack to restore the values ​​after the POP.

[0259] This is illustrated in Figure 38, where registers are used to calculate the x value of a node center as a result of octree node subdivision. A similar implementation is used for y and z. The starting x position of the node center is loaded into x register 3801. This can be a region of the octree or the starting node. The xi value is loaded into shift register xi 3803. Similarly, xj is loaded into shift register xj 3805, and in 3D, xk is loaded into shift register xk 3807. After moving to a child node, such as with a PUSH, the value in x register 3801 is modified by adding or subtracting the values ​​in the three shift registers with adder 3809. The combination of additions or subtractions to the three registers corresponds, as appropriate, to the three bits of the child number. The new value is loaded back into x register 3801 and is available as the converted output x value 3811. The addition and subtraction operations can be undone to return to the parent value, such as with a POP operation. Of course, sufficient space must be allocated to the right of each shift register to prevent loss of precision. The child node sequence used in the subdivide operation must of course be preserved. Alternatively, the x values ​​can be saved on a stack, or another mechanism can be used.

[0260] Extending this to perspective projection is shown in Figure 39. Now, node center 3618 is projected along projection vector 3901 onto display screen 3617 to a point on the display screen with X value x' 3903. The value of x' will not, in general, be equal to x value 3621. It is a function of the distance from viewpoint 3611 to the display screen in the Z direction, distance d 3907, and the distance from the viewpoint to node center 3618 in the Z direction, distance z 3905. The value of x' can be calculated using a similar triangle as follows: x' / d=x / z or x'=xd / z [Equation 9]

[0261] This requires a general-purpose division operation, which requires significant hardware and clock cycles compared to simpler mathematical operations such as integer shifts and adds. It is therefore desirable to recast the perspective projection operation into a form that can be implemented with simple arithmetic operations.

[0262] As shown in FIG. 40A, we use "spans" to perform perspective projections. A window 4001 is defined on the display screen 3617. It extends from the X value of the bottom of window 4003 to the X value of the top of window 4005. As noted above, the size in pixels of the display screen 3617 is a power of two. Window 4001 begins as the entire screen and therefore begins as a power of two. It is then subdivided into half-sized sub-windows as needed to enclose the node projection. Thus, all windows have sizes that are powers of two. While window 4001 is shown in 2D, in 3D it is a 2D window of the display screen 3617, and the window size in each dimension, X and Y, is a power of two. For display purposes, the window size may be the same in X and Y. In other uses, these may be maintained independently.

[0263] Ray 4021 is from viewpoint 3611 to the top of window 4005. In 3D, this is a plane extending in the Y direction that forms the top edge of the window in the XY plane. Ray 4023 is from the viewpoint to the bottom of window 4003. The line segment in the X direction that intersects node center 3618 and lies between the bottom of the intersection of span 4011 with ray 4023 and the top of the intersection of span 4013 with ray 4021 is span 4009.

[0264] If the span is translated in the z direction by a specified amount, the span 4009 will increase or decrease by an amount that depends only on the gradients of the top-projected ray 4021 and bottom-projected ray 4023, and not on the z position of the node center 3618 of node 4007. In other words, the size of the span will change by the same amount for the same step in z no matter where it is in Z. Thus, if the step is taken from a parent node, the change in span will be the same regardless of where the node is located. If a child node is then subdivided, the change in span will be half the change from parent to child.

[0265] Figure 40B expands on the concept with a window center ray 4025 from viewpoint 3611 to the center of window 4001, with window center 4019 dividing window 4001 into two equal parts. Similarly, the point dividing the span into two equal parts is the center of span 4022. This is used as a reference point to determine the X value of the center of node 3618, which is offset from the center of span 4022 by node center offset 4030. Thus, knowing the location of the point center of span 4022 and node center offset 4030 yields the location in X of the node center.

[0266] As discussed above regarding the size in X of a span when translated in the z direction, the center of span 4022 can similarly be determined by a step in Z, regardless of its position in Z. Combined with the changes in X, Y, and Z when moving from the center of the parent node to the center of the child as shown in FIG. 37, the change in size in X of the span can be calculated for each subdivision. Similarly, the center of span 4022 can be recalculated for each octree subdivision by adding the difference, regardless of its position in Z. Knowing the center of span 4022 and the center shift in X after subdivision, the new position of the child's center, relative to the central ray 4025, can be obtained using only an addition or subtraction operation. Similarly, the intersection of the bounding box of node 4015 with the span can also be calculated by addition, since it is a fixed distance in X from the node center 3618.

[0267] Therefore, as shown above, the offset added or subtracted to move from the parent center point to the child center point is divided by 2 after each subdivision. Thus, the position of the center of the span 4022, the center of the node 3618, and the top and bottom X positions of the bounding box can be calculated for PUSH and POP by shift and add operations. The two measures used as usual in the calculations below are 1 / 4 of the size of the span, i.e., QSPAN 4029, and 1 / 4 of the window, i.e., QWIN 4027.

[0268] The perspective projection process works by traversing the octree with PUSH and POP operations to simultaneously determine the location of the center of the current node on the span with respect to the center of span 4022. Along with the boundaries of the bounding box on the span, the window is refined and the center ray is moved up or down as needed to keep the node within the window. The process of refining a span is illustrated in Figure 40C. An origin ray 4025 intersects the span with span 4009 at the center of span 4022. The span passes through node center 3618 and intersects the top of projection ray 4021 at the top of the span and the bottom of projection ray 4023 at the bottom. If the span needs to be refined after a child is PUSHed, a new origin ray 4025 and center of span 4022 may be required. There are three options for the new origin ray 4025. First, it can remain the same. Second, it can be raised to the middle of the top half of the original span. Or third, it can be lowered to the center of the lower half of the original span.

[0269] To determine the new span and new center, zones are defined. The zone where node center 3618 is placed after PUSH determines the movement of the origin ray. As shown, ZONE1 4035 is centered at the current span center. Then there are two upper zones and two lower zones. Upward ZONE2 4033 is the next step in the positive X direction. Upward ZONE3 4031 is above it. Similar zones, ZONE2 4037 and ZONE3 4039, are defined in the negative X direction.

[0270] The new node center location is known from the newly calculated X node offset, which is distance NODE_A 4030. Two comparisons are performed to determine the zone. If the new node value NODE_A is positive, then the movement, if any, is in the positive X direction and the value of QSPAN is subtracted from it. If NODE A is negative, then the movement, if any, is in the negative X direction and the value of QSPAN is subtracted from it. This result is called NODE_B. Of course, in 3D there are two sets of values, one in the X direction and one in the Y direction.

[0271] To determine the zone, only the sign of NODE_B is needed. There is no need to actually perform an addition or subtraction; a magnitude comparison is sufficient. In this embodiment, this is calculated because NODE_B will be the next NODE value in some cases.

[0272] The second comparison is between NODE_A and 1 / 8 of the span (QSPAN / 2). It is compared to QSPAN / 2 if NODE_A is positive, and to -QSPAN / 2 if it is negative. The results are: NODE_A NODE_B NODE_A:QSPAN / 2 Resulting Zone ≧0 ≧0 (optional) ZONE=3 ≧0 <0 NODE_A≧QSPAN_N / 2 ZONE=2 ≧0 <0 NODE_A <QSPAN_N / 2 ZONE=1 <0 <0 (optional) ZONE=3 <0 ≧0 NODE _A≧-QSPAN_N / 2 ZONE=1 <0 >0 NODE A<-QSPAN_N / 2 ZONE=2

[0273] A node center exceeding the span results in a zone 3 situation. The results of span subdivision are illustrated in Figure 40D. The origin can be raised 4057 to a new top origin 4051 by subtracting QSPAN_N from NODE_A, or lowered 4058 to a new bottom origin 4055 by adding QSPAN_N to NODE_A. Regardless, the span can be divided into half-sized spans, resulting in one of three new spans: a top span after division 4061, an unshifted span after division 4063 to the new unshifted center 4053, or a bottom span after division 4065. The following actions are performed based on the zone: CASE SHIFT DIVIDE ZONE = 1 NO YES ZONE = 2 YES YES ZONE = 3 YES NO

[0274] Thus, in the zone=3 situation, a centering operation is performed without splitting the upper and lower half spans. For zone=2, splitting is performed and the resulting spans and windows are recentered. For zone=0, the spans (or windows) are split, but no centering is required.

[0275] There are a few additional factors that control the subdivision and centering process. First, 1 / 4 of the distance of the bounding box edge in X from the node center is maintained. This is called QNODE. This is a single number (in X) because the node center is always the center of the bounding box. Separate positive and negative offsets are not needed. This is compared to a fraction of QSPAN_N (usually 1 / 2 or 1 / 4). If QNODE is larger than this, the bounding box is considered already large in terms of span (in that dimension). No span subdivision is performed (NO_DIVIDE situation). For display window subdivision, no centering is performed if the window moves beyond the display screen (NO_SHIFT situation).

[0276] Sometimes additional splitting and / or centering operations are required for PUSH. This is called a repeat cycle. It is triggered after a zone 3 situation (centering only) or when the span is large with respect to the bounding box (QSPAN_N / 8>QNODE). Another comparison inhibits a repeat cycle if the window is too small (less than one pixel with respect to the display window).

[0277] The subdivision process may continue until the window reaches a certain level within the pixel or quadtree in the case of image generation. At this point, the property values ​​in the octree node are used to determine the action to be taken (e.g., write a value to the pixel). Alternatively, when a terminal octree node is first encountered, its property values ​​can be used to write a value to the window. This may involve writing to the pixel that constitutes the window or node at the appropriate level within the quadtree. Alternatively, a "full node push," or FNP, may be initiated, in which the terminal node is used to create a new child at the next level down, inheriting appropriate properties from its parent. In this way, the subdivision process continues to some level of window subdivision, perhaps with properties modified as new octree nodes are generated in the FNP. Due to the geometry involved in perspective projection, sometimes subdivision of the octree does not cause a subdivision of the window, or the window may need to be divided multiple times. The appropriate subdivision for that operation (e.g., PUSH) is suspended.

[0278] Figure 41 is a schematic diagram of an implementation for one dimension (X or X with Y shown). The values ​​contained in the registers are shown inside each register to the left. The name may be followed by one or more items in parentheses. One item in the parentheses may be I, J, or K, indicating the dimension of the octree coordinate system for that value. A value of J, for example, indicates that it holds a value for steps in the J direction from parent to child. Another may have "xy" in parentheses, indicating that there are actually two such registers, one for the X dimension and one for the Y dimension. The parenthesized item is then followed by a number in parentheses, indicating the number of such registers: 2 indicates one for the X dimension and one for the Y dimension. A value of 1 indicates that the register can be shared by both the X and Y portions of the 3D system.

[0279] Many registers have one or two sections separated to the right. These registers are shift registers, and the space is used to store the least significant bit when the value in the register is shifted right. This is so the value does not lose precision after performing a PUSH operation. This section may contain a "lev" that indicates one bit position for each octree level that could be encountered (e.g., the maximum level to push) or a "wlev" for the number of window levels that can be encountered (e.g., the number of window subdivisions or the maximum number of octree subdivisions that reach a pixel). Some registers require room for both. The wlev value is determined by the size of the display screen in that dimension. A 1024-pixel screen, for example, requires 10 extra bits to prevent precision loss.

[0280] The oblique projection process begins by transforming the region of the octree (I, J, and K) into the display coordinate system (X, Y, and Z). This is the node center 3618. A value is stored in the NODE register 4103 for X (and Y in a second register). The Z value uses the implementation shown in Figure 38 and is not shown here. The node centers of the octree region, or vectors from the starting node to the children in the octree coordinate system, are i, j, and k. The differences in X, Y, and Z from the step at i, j, and k are calculated for the first step and placed in DIF registers 4111, 4113, and 4115 according to the method outlined in Figure 37. Separately, these three registers and adder 4123 can be used to update the NODE register 4103 after a PUSH in the orthogonal projection, as shown in Figure 38. An additional set of adders 4117, 4119, and 4121 corrects for oblique projection.

[0281] An initial span value 4009 is calculated at the node center for the starting node. One-quarter of this value, qspan 4029, is placed in QSPAN ("quarter span") register 4125. The span length differences for the initial step of the i, j, and k vectors are calculated from the starting node to its children. One-quarter of these difference values ​​are placed in QS_DIF ("quarter span difference") registers 4131, 4133, and 4135. As in Figure 38, the QS_DIF registers are used in conjunction with adder 4137 to update the QSPAN registers on PUSH. The QS_DIF values ​​are also added or subtracted from the associated DIF values ​​using adders 4117, 4119, and 4121 to account for changes in the location where the span intersects the center ray 4025 due to octree PUSH moves in the i, j, and k directions.

[0282] When a window is subdivided and the new window is half the size of the original window, the central ray 4025 may remain the same, and the upper and lower half windows are reduced in size by half. Alternatively, the central ray can be moved to the center of the upper half or the center of the lower half. This is taken into account by selector 4140. If the origin central ray does not change, the existing central ray, NODE_A, is retained. If shifted, the old central ray 4025 is moved by adding a new QSPAN value, QSPAN_N, via adder 4139, which forms NODE_B.

[0283] The configuration continues in Figure 41B, two of the hardware configurations 4150. The node center offset 4030 from the center of span 4022, value NODE, is calculated and loaded into register 4151. The difference values ​​moving from the parent node center to the children in Z are calculated and loaded into shift registers 4153, 4155, and 4157. These are used in conjunction with adder 4159 to calculate the next value of NODE for Z, NODE_N.

[0284] The QNODE value is calculated and placed in shift register QNODE 4161. This is a fixed value that is divided by 2 (shifted one place to the right) on every PUSH. It is compared to 1 / 2 of QSPAN_N using a 1-bit shifter 4163 and comparator 4171 to determine if no division occurs (NO_DIVIDE signal). It is also compared to 1 / 8 of QSPAN_N using a 3-bit shifter 4165 and a comparator 4173 to determine if the division should be repeated (REPEAT signal). Shifter 4167 is used to divide QSPAN_N by 2, and comparator 4175 is used to compare it to NODE_A. As shown, if NODE_A is >0, the positive value of QSPAN_N is used; otherwise, its sign is inverted. The output sets the Zone=2 signal.

[0285] For the display situation, this is fairly simple since there is no movement in the Z direction. The on-screen location of the window center is initialized in WIN register 4187 (one in X and one in Y). One-quarter of the window value, QWIN, is loaded into shift register 4191. This is used in conjunction with adder 4189 to maintain the window center. The screen diameter (in X and Y) is loaded into shift register 4181. This is compared to the window center by comparator 4183 to prevent the center from shifting beyond the edge of the screen, which occurs when the node's projection is off the screen. The MULT2 value is a power of two used to maintain sub-pixel precision in register 4187. A similar value, MULT, also a power of two, is used to preserve precision in the span geometry registers. Since the screen diameter in register 4181 is in pixels, the WIN value is appropriately divided by shifter 4185 to scale appropriately for comparison. Comparator 4193 is used to compare the current window size with 1 pixel, thus preventing the window subdivision from being smaller than 1 pixel or smaller than the appropriate size. Since QWIN is scaled, this must be scaled with MULT2.

[0286] The numbers next to the adders indicate the actions to be performed in different situations. These are: Adder Operation 1 PUSH: Add if the child bits (i, j, and k) are 1, otherwise subtract POP: Opposite of PUSH Otherwise (not PUSH or POP, window-only subdivision): No operation Adder Operation 2 If NODE_A>0, subtract; otherwise, add Adder Operation 3 UP: + DOWN: - Otherwise (not UP or DOWN): No operation Adder Operation 4 The central ray 4025 is shifted to the upper half and added The central ray 4025 is shifted to the lower half and subtracted Adder Operation 5 PUSH: Subtract if child bits (i, j, and k) are 1, add otherwise POP: Opposite of PUSH Otherwise (not a push or pop): add 0 to (i, j, and k)

[0287] The shift register should be shifted as follows: Registers 4111, 4113, 4115 If PUSH, shift right by 1 at the end of the cycle (to the lev bit) In the case of POP, shift left by 1 at the beginning of the cycle (from the lev bit) Register 4125 In the case of windowing, shift right by 1 at the end of the cycle (to wlev bits) In case of window merge, shift left by 1 at the beginning of the cycle (from wlev bit) Registers 4131, 4133, and 4135 (which can shift two bits per cycle) If PUSH, shift right by 1 at the end of the cycle (to the wlev bit) In the case of windowing, shift right by 1 at the end of the cycle (to wlev bits) In the case of POP, shift left by 1 at the beginning of the cycle (from the lev bit) In case of window merge, shift left by 1 at the beginning of the cycle (from wlev)

[0288] In summary, an orthogonal projection can be implemented with three registers and adders for each of the three dimensions. The geometric calculations for a perspective projection can be implemented using eight registers and six adders per dimension. In use, this performs a perspective transformation on two of the three dimensions (X and Y). The total registers for X and Y are actually 12 registers because the QSPAN and three QS_DIF registers do not need to be duplicated. Three registers and one adder are used for the third dimension (Z). The total for a full 3D implementation is therefore 15 registers and 12 adders. Using this method, the geometric calculations for an octree PUSH or POP can be performed within one clock cycle.

[0289] To generate images of multiple octrees with different placements and orientations, these can be geometrically converted into a single octree for display, or the z-buffer concept can be extended to a quadtree-z or qz-buffer, with z values ​​contained within each quadtree node to eliminate hidden portions when a front-to-back traversal cannot be forced. Those skilled in the art can devise variations of this display method.

[0290] For display, this method uses spans that move with the center of the octree node during PUSH and POP operations. A display screen can be considered to have a fixed span in that it does not move in Z. This makes using a pixel array or quadtree with a fixed window size a convenient method for display. In other projection uses, such as those described below, a fixed display may not be necessary. In some cases, multiple octrees, including SAOs, may be tracked simultaneously and independently subdivided as needed to keep the current node within span limits. In such cases, multiple spans are tracked, such as one for each octree. In some implementations, multiple spans may be tracked for an octree, such as multiple children or descendants, to improve the speed of operations.

[0291] Image processing operations on quadtrees and equivalent operations on octrees are performed in filtering processing module 1911. Neighbor search can be implemented with a sequence of POPs or PUSHs, or multiple paths can be followed in the traversal, making neighbor information continuously available without backtracking.

[0292] The Surface Extraction Processing module 1913 extracts a set of surface elements, such as triangles, from the octree model for use when required. Many methods are available for doing this, including the "marching cubes" algorithm.

[0293] The morphological operation processing module 1915 performs morphological operations (eg, dilation and erosion) on the octrees and quadtrees.

[0294] The connectivity processing module 1917 is used to identify octree and quadtree nodes that spatially touch under some set of conditions (e.g., an intersection of properties). This can be specified to include touching points (quadtree or quadtree corners) along edges (octree or quadtree) or on faces (octree). This typically involves marking nodes in some fashion starting from a "seed" node, then traversing to all connected neighbors and marking them. In other cases, all nodes are examined and all connected components are divided into disjoint sets.

[0295] The mass properties processing module 1919 calculates the mass properties (volume, mass, center of mass, surface area, moment of inertia, etc.) of the dataset.

[0296] The registration processing module 1921 is used to refine the placement of 3D points in a scene as estimated from 2D positions found in multiple images. This is described below in conjunction with Figure 23B. The light field operations module 1923 is further described below with respect to Figure 20.

[0297] Light entering a voxel containing a medium interacts with the medium. In the present invention, discrete directions are used, and the transport equation [2] becomes a summation rather than an integral, and the equation

[0298]

number

[0299] It is described by:

[0300] 20 is a block diagram illustrating the light field operations module 1923 of the SRE, which is part of the spatial processing module 1113 of the SRE 201. The light field operations module 1923 may include a position-invariant light field generation module 2001, an incident light field generation module 2003, an outgoing / incident light field processing module 2005, and an incoming / outgoing light field processing module 2007.

[0301] The light field operations module 1923 performs operations on light in the form of SAOs, computing the interaction of light with media from any source, represented as an octree. Octree nodes may contain media that transmit, reflect, scatter, or otherwise modify light according to properties stored in or associated with the node. An SAO represents light entering or exiting from some region of volumetric space at a point within that space. Light entering a scene from outside the scene is represented as a position-invariant SAO. A single position-invariant SAO is valid for any position within the associated workspace. A scene may have multiple sub-workspaces, each containing its own position-invariant SAO, but only a single workspace is described. To compute an all-point light field SAO, it needs to be supplemented with light from within the scene. For a specified point in the workspace, an incident light field is generated and then used to effectively overwrite a copy of the position-invariant SAO.

[0302] Incident light at a medium that interacts with light causes a responsive light field (SAO) that is added to the emitted light. The BLIF of the medium is used to generate a responsive SAO for the medium based on the incident SAO.

[0303] The position-invariant light field generation module 2001 generates a position-invariant SAO of incident light obtained from a workspace, for example, from an outward-facing image. By definition, an object represented by a position-invariant SAO exhibits no parallax when viewed from anywhere in the associated scene workspace. Therefore, a single position-invariant SAO is applicable for all positions in the workspace. This is further described with respect to FIG. 28A.

[0304] Light may also enter the scene, which exhibits parallax. To represent this, a surface light field may be used. In one embodiment, this light is represented by a surface light field consisting of outgoing SAOs on the scene boundary or on some other suitable surface. These are typically created as needed to represent external light.

[0305] The incident light field generation module 2003 calculates a light field SAO for any point in the scene based on the media in the scene. As used herein, a "medium" in media is something in space that emits or interacts with light. This may be determined in advance (e.g., a terrain model, surfels, CAD model) or "discovered" during processing. Additionally, it may be refined as processing operations continue (new media, higher resolution, higher quality, etc.).

[0306] As described, the position-invariant SAO is a single SAO that is a representation of the incident light at any point in the workspace. However, for point light field SAOs, this must be corrected when the space within the frontier is not completely clear. The assumption that anything in the scene inside the frontier has no parallax from inside the workspace is no longer valid. To accommodate this, an incident SAO is generated for a specified point in the workspace and then combined with the position-invariant SAO. Such an SAO can be generated outside the workspace, but the position-invariant SAO is no longer valid.

[0307] For concentric SAOs (same center point), the concept of SAO layers is introduced. SAOs are arranged in layers with decreasing lighting priority as they move away from the center. Thus, for a given point in the workspace, the highest priority SAO is generated for media and objects within the frontier, the incident SAO, and not including the frontier. The position-invariant SAO has a lower priority over the incident SAO. The two are merged to form a composite SAO, which is a point light field SAO for the arrangement. The incident SAO effectively overwrites the position-invariant SAO. It is not necessary for a new SAO to be created or for the position-invariant SAO to be modified. They can be merged with the preferred incident node in a UNION (used if a node with the property exists in both).

[0308] Position-invariant SAOs may themselves consist of layers: for example, a lower priority layer could model the sun, and something closer to the sun that does not exhibit parallax could be represented as a higher priority SAO.

[0309] A medium is represented as one or more octree media nodes. Behavior is determined by the properties of the medium, represented by a particular octree (common to all nodes) and by specific properties contained within the node. When incident light strikes the medium, the resulting response light from the BLIF interaction is modeled. In addition to changing direction, octree nodes can attenuate the light intensity, change its color, etc. (if the SAO properties are appropriately modified).

[0310] The incident light field generation module 2003 implements a procedure to generate an incident SAO for a given point within the workspace, which effectively overwrites the common copy of the position-invariant SAO when a medium is found to block the view of the frontier from the given point.

[0311] The algorithm used in the incident light field generation module 2003 is cast as a form of intersection operation between a solid angle of interest centered at a specified point and media within the region. This minimizes the need to access media outside the region needed to generate the SAO. Using the hierarchical nature of the octree, higher resolution regions of interest are only accessed when a potential intersection is indicated at a lower level of resolution.

[0312] Light from one direction should only come from the nearest spatial region of the medium in that direction. Unnecessary computations and database accesses occur when media behind the nearest region is processed. A directed traversal of the octree (based on spatial sorting) is used. By searching in an octree traversal sequence from front to back, in this case outward from a specified point, the traversal in a particular direction stops when the first node with opaque medium is encountered for a particular octree.

[0313] The outgoing / incoming light field processing module 2005 operates to generate an incoming light field SAO for a point or represented region in space, and typically the contributions of multiple outgoing SAOs in a scene must be accumulated. The operation is illustrated in 2D in Figures 31A and 31B.

[0314] The function of the Incident / Emitted Light Feed Processing module 2007 is to calculate the emitted light from a location on a surface, which is the sum of the internally generated light and the incident light reflected or refracted by the surface.

[0315] SAOs can be used to represent incident light, emitted light, and response light. In another use, SAOs are used to represent BLIF (or other related property or set of properties) for a medium. BLIF coefficients are stored as BLIF SAO node properties. These are typically defined as a set of four-dimensional weights (two angles of incidence and two angles of emergence). SAOs represent two angles. Therefore, the two remaining angles may be represented as a set of properties at each node, resulting in a single SAO. Alternatively, they may be represented by multiple SAOs. The actual weights may be stored or generated on the fly during processing operations. As one skilled in the art will appreciate, other methods may also be used.

[0316] The spherical BLIF SAO is rotated appropriately for each outgoing direction of interest using the geometry processing module 1905 (as described in FIG. 38). The coefficients are, for example, multiplied by the corresponding values ​​in the incoming light SAO and summed. Light generated within the medium is also added to the outgoing light. This is further described with respect to FIG. 32.

[0317] Figure 22 shows a 2D example of an SAO with rays. The node centered at point 2209 is bounded by rays 2211 and 2212, which are the angular limits of a solid angle (in the 2D plane). At the next lower level (higher resolution) in the SAO, the child node centered at 2210 represents the solid angle from rays 2213 to 2214.

[0318] The surface area of ​​the unit sphere represented by the SAO node can be expressed in a variety of ways, depending on the operational needs and computational limitations. For example, a property can be attached to the SAO node to indicate some area measure (e.g., the angle of a cone around a ray passing through a point) or a direction vector for use in the operation. Those skilled in the art will understand that this can be done in a variety of ways.

[0319] As mentioned above, the registration processor 1921 is used to refine the location of 3D points in a scene as estimated from 2D locations found in multiple images. Here, the 3D points in the scene are called "landmarks," and their 2D projections on the images are called "features." In addition, the camera's location and gaze direction must be refined as the images are captured. At the start of this process, multiple 3D landmark points are identified, roughly aligned to the estimated 3D points in the scene, and labeled with unique identifiers. Associated 2D features are located and labeled (minimum 2) in the images in which they appear. Approximate estimates of the camera parameters (location and gaze direction) are also known.

[0320] These estimates can be calculated in many ways. Essentially, image features that correspond to the projections of the same landmarks in 3D need to be found. This process involves finding the features and determining whether they correspond. With such correspondence pairs and approximate knowledge of the camera pose, one skilled in the art can triangulate the rays emanating from these features to obtain a rough estimate of the landmark's 3D location.

[0321] The detected features must be discriminated and properly localized, and they must be able to be redetected when corresponding landmarks appear in other images. It is assumed that the surface normal vectors for all points in the 2D image, such as 2349 in Figure 23A, have been calculated.

[0322] For every point, we compute the local scattering matrix for each normal vector as a 3x3 matrix. Σ i=1..9 (N i N i T ) [Formula 10A] Here, each N iis the normal vector (Nx, Ny, Nz) at points i=1...9, shown as 2350-2358 in 2349 of Figure 23A. We define these as feature points, points where the determinant of this matrix is ​​maximum. Such points correspond to points of maximum curvature. The determinant is invariant to surface rotation, and the same point can therefore be detected even if the surface is rotated. Descriptors can be estimated on larger neighborhoods, such as 5x5 or 7x7, proportional to the resolution of the grid where the normals are estimated from the Stokes vectors.

[0323] To find matches between features, we compute a descriptor that represents the local distribution of surface normals within the point's 3D neighborhood. Consider the point detected at location 2350 in Figure 23A. The surface 2349 shown is unknown, but the normal can be obtained from the Stokes vector value. Assume the immediate neighborhood consists of eight neighbors (2351-2358) of point 2350, each with a normal as shown in the figure. The descriptor is a spherical histogram 2359, with bins separated by uniformly spaced latitude and longitude lines. Normals with similar orientations are grouped into the same bin, as shown in the figure: normals at 2354, 2355, and 2356 are grouped in bin 2361, normals at 2357 in bin 2363, normals at 2350 in bin 2362, normals at 2358 in bin 2364, and normals at 2351-2352 in bin 2365. The histogram bin values ​​are 3 for bin 2361, 1 for bins 2363, 2362, and 2364, and 2 for bin 2365. The similarity between features is then expressed as the similarity between spherical histograms. Note that these spherical histograms are independent of color, or lack thereof, and therefore depend only on the geometric description of the surface onto which they are projected. We cannot directly take the absolute or squared difference between two spherical histograms of normals because the two views differ in orientation and one histogram is a rotated version of the histogram for the same point. Instead, we use spherical harmonics to compute invariants. If the spherical histogram is a function f(latitude, longitude) whose spherical harmonic coefficients are F(l,m), where l = 0..L-1 are latitudes and m = -(L-1)..(L-1), longitude frequencies, then the magnitude of the vector [F(l,-L+1).....F(l,0).....G(l,L-1)] is invariant to 3D rotations, and such invariants can be computed for all l = 0..L-1. We compute the sum of the squared differences between them, which is a measure of dissimilarity between the two descriptors. We can select the corresponding point in the second view that has the smallest dissimilarity to the considered point in the first view.Those skilled in the art can apply matching algorithms from the theory of algorithms (such as the Hungarian algorithm) given the (dis)similarity measures defined above. From all pairs of corresponding features, as described above, we can determine a rough estimate of the landmark location.

[0324] In general, there are n landmarks and m images. In Figure 23B, a 2D representation of the 3D registration situation, landmark point 2323 is contained within region 2310. Images 2320 and 2321 are on the edges (planes in 3D) of regions 2311 and 2312, respectively, of the associated camera positions. The two camera viewpoints are 2324 and 2325. The location in image 2320 of the projection of the initial location of landmark 2323 onto viewpoint 2324 along line 2326 is point 2329. However, the location where the landmark was detected in image 2320 is 2328. In image 2321, the projected feature location on line 2327 is 2331, but the detected location is 2330.

[0325] The detected feature points 2328 and 2330 are at fixed locations in the image. The landmark 2323, the camera locations 2324 and 2325, and the orientation of the camera regions 2311 and 2312 are initial guesses. The goal of registration is to adjust to minimize some defined cost function. There are many ways to measure the cost. The L2 norm cost is used here. It is the sum of the squared 2D distances in the image between the adjusted feature locations and the detected locations. This is, of course, only for images in which the landmarks appear as features. In Figure 23B, these two distances are 2332 and 2333.

[0326] Figure 24 is a plot of a cost function. The Y-axis 2441 is cost, and the X-axis 2440 represents adjustable parameters (landmark placement, camera placement, camera view direction, etc.). Every set of parameters on the X-axis has a cost on curve 2442. The goal is to find the set of parameters that minimizes the cost function, point 2444, given a starting point, such as point 2443. This typically involves solving a set of nonlinear equations using iterative methods. The curve shows a single minimum cost, a global minimum, although in general there can be multiple local minima. The goal of this procedure is to find a single minimum. Many methods are known for searching for other minima and ultimately identifying the global minimum.

[0327] The equations that determine the image intersection points as a function of variables are typically collected in a matrix that is applied to a parameter vector. To minimize the sum of squared distances, the derivative of the cost (the sum of squared 2D image distances) is typically set to a minimum value that indicates zero. A matrix method is then used to determine a solution that is then used to estimate the next set of parameters to use on the way to the minimum. This process becomes computationally intensive as the number of parameters (landmarks, features, and cameras) grows, and takes a relatively long period of time. The method used here does not directly calculate the derivatives.

[0328] The registration processing module 1921 operates by iteratively projecting landmark points onto an image and then adjusting parameters to minimize a cost function in the next iteration. It generates a feature location of the landmark points at their current location in the image, represented as a quadtree, using the oblique projection image generation method of the image generation module 1909 as described above. Such an image is generated with a series of octree and quadtree PUSH operations that refine the projected location.

[0329] The cost to be minimized is Σd² over all images, where d is the distance in each image from the detected feature point in the image to the associated projected feature point, measured to some resolution in pixels. Each is the sum of the x² and y² components of the distance in the image (e.g., d² = dx² + dy²).

[0330] The calculations are shown in FIG. 25A for the X dimension. Although not shown, a similar set of calculations is performed for the y values, and the squared values ​​are summed. The process for the X dimension begins by performing a projection of the landmarks at their initial estimated locations onto the quadtree. For each feature in the image, the distance in X from the detected location to the calculated location is stored in register 2507 as the value dx. Those skilled in the art will appreciate that a register as used herein can take any form of data storage. This value may be of any precision and is not necessarily a power of two. The initial value is also squared and stored in register 2508.

[0331] The edge value in shift register 2509 is the edge distance of the initial quadtree node in X and Y, which moves the center of the quadtree node to the center of the child node. The edge distance is added or subtracted from the x and y placement values ​​of the parent to move the center to one of the four children. The purpose of the calculation in Figure 25A is to calculate new values ​​for dx and dx2 after a quadtree PUSH. As an example, if a PUSH is made to a child in the positive x and y directions, it calculates new values ​​d' and d'2 as follows: dx'=dx+edge [Equation 11] dx'2=(dx+edge)2=dx2+2*dx*edge+edge2 [Formula 12]

[0332] Because the quadtree is constructed by regular subdivision by two, the edges of nodes at any level (at X and Y) can be made a power of two. Thus, the value for an edge can be maintained by placing the value 1 in the appropriate bit position in shift register 2509 and shifting it one bit position to the right with PUSH. Similarly, the edge2 value in shift register 2510 is a 1 in the appropriate bit position that is then shifted two places to the right with PUSH. Both are shifted left with POP (one place for register 2509, two places for register 2510). As shown, edge is added to dx by adder 2512 to produce a new value with PUSH.

[0333] To calculate the new dx2 value, 2*dx*edge is needed. As shown, a shifter 2511 is used to calculate this. A shifter can be used for this multiplication because edge is a power of 2. An additional left shift is used to account for the 2x multiplication. As shown, this is added to the old dx2 and edge2 values ​​by adder 2513 to calculate the new dx2 value with a PUSH.

[0334] Note that the values ​​in shift registers 2509 and 2510 are the same for the y value; they do not need to be duplicated. The adders for summing the dx2 and dy2 values ​​are not shown. Depending on the particular implementation, the calculation of the new d2 value at PUSH may be accomplished in a single clock cycle. Those skilled in the art will understand that this calculation may be accomplished in a variety of ways.

[0335] The entire process iteratively calculates a new set of parameters that drives the cost toward a minimum, stopping when the cost change falls below the minimum. There are many methods that can be used to determine each new set of parameters. An exemplary embodiment uses the well-known procedure, Powell's Method. The basic idea is to select one of the parameters and vary only it to find the minimum cost. Then, another parameter is selected and varied to move the cost to another, lower minimum. This is done for each parameter. Finally, for example, the parameter with the largest change is fixed before a parameter deviation vector is generated and used to determine the next set of parameters.

[0336] When the x, y, and z parameters of a particular landmark point (in the scene's coordinate system) are varied, the contribution to Σd2 is only the summed d2 value for that landmark, independent of other landmarks. Thus, coordinates can be modified individually and minimized in isolation. If the computation is distributed across multiple processors, many iterations of such changes can be computed quickly, perhaps in a single clock cycle.

[0337] This procedure is the registration process 2515 in Figure 25B. The initial estimated landmarks are projected onto the quadtree to determine the starting feature placement in operation 2517. The differences from the detected feature placement are squared and summed over an initial cost value. The landmarks are then independently transformed with the sum of the squared differences being minimized in turn for each parameter that moves the landmark. This is performed in operation 2519.

[0338] Parameter modifications are then performed on the camera parameters, in this case with all landmark points represented in the associated image. This is performed in operation 2521. The results of the parameter modifications are calculated in operation 2523. This process continues until a minimization goal is achieved in operation 2527, or until some other threshold is reached (e.g., a maximum number of iterations). The results of operation 2523 are then used to calculate the next parameter vector to apply in update operation 2525, if not finished. When finished, operation 2529 outputs the refined landmark and camera parameters.

[0339] The method for moving a coordinate value (x, y, or z) within a landmark represented by an octree is to step in only that dimension with each PUSH. This involves changing from using the center of the node as parameterization to using the smallest corner in the preferred method.

[0340] This is shown in Figure 26 in 2D. Coordinate axis 2601 is the X axis and axis 2602 is the Y axis. Node 2603 has its center at point 2604. Instead of using it as two parameters (x and y, or x, y, and z in 3D), point 2603 is used. Then, to move the parameters in the next iteration, subdivision is performed on the child node with its center at 2606, but the new placement for the parameters is 2607. If, in the next iteration, the movement is positive again, the next child will have its center at 2608, but the new set of parameters will change to point 2609. Thus, only the x parameter is changed, but the y value is fixed (along with all parameters other than x).

[0341] The use of minimum placement rather than node centers is illustrated in 2D in Figure 27. The quadtree plane is 2760. The two quadtree nodes that are projected spans are 2761 and 2762. Ray 2765 is the ray above the window, and ray 2764 is the ray below the window. Ray 2763 is the origin ray. Node 2766 is centered at point 2767. The original node x span distance value is 2762. The node placement for the measurement is then moved to the minimum node point (in x and y, or x, y, and z in 3D), which is point 2763. The new node x distance value is 2767.

[0342] As the octree representing the landmarks is subdivided, the step is changed by 1 / 2 with each PUSH. If the new Σd2 for this landmark (and only this landmark) for the image containing the feature of interest is smaller than the current value, the current set of parameters is moved to it. If it is higher, the same incremental change is used, but in the opposite direction. Depending on the implementation, a move to a neighboring element node may be necessary. If the cost for both is higher, the parameter set remains unchanged (selecting the child with the same value) and another iteration continues with a smaller incremental value. The process continues until the cost change falls below some threshold. The process may then continue with another dimension for the landmark.

[0343] While the above minimization can be performed for multiple landmark points simultaneously, the step of modifying parameters related to camera placement and orientation requires that Σd values ​​be calculated for all landmarks appearing in the image. However, this can be performed independently and simultaneously for multiple cameras.

[0344] An alternative method of calculating the cost (e.g., directly calculating d2) is the use of dilation. The original image, along with the detected feature points, is iteratively dilated from those points, and a cost (e.g., d2) is attached to each dilated pixel up to some distance. This is done once for each image. Features that are too close to another feature can, for example, be added to a list of features or placed in another copy of the image. The image is then converted into a quadtree with reduced features (parent features generated from the properties of their child nodes). A minimum value is used at higher levels, allowing for later abandonment of projective traversals if the minimum cost exceeds the current minimum, eliminating unnecessary PUSH operations. For higher accuracy, the quadtree is calculated at a higher resolution than the original image if the original detected features were identified with sub-pixel accuracy during the solution.

[0345] This concept can also be extended to include the use of landmarks other than points. For example, a planar curve can form a landmark. This can be translated, rotated, and then projected onto the associated image plane. The cost is the sum of the values ​​in the dilated quadtree that the projected voxel projects onto. This can be extended to other types of landmarks by those skilled in the art.

[0346] Figure 28A illustrates a position-invariant SAO sphere 2801 that is exactly bounded by a unit cube 2803 with front-facing faces 2805, 2807, and 2809. The three rear-facing faces are not shown. The center of the SAO and bounding cube is point 2811, which is also the axis of the coordinate system. The axes are X-axis 2821, Y-axis 2819, and Z-axis 2823. The axes exit the cube at the centers of the three forward-facing axes indicated by cross 2815 on the X-axis, cross 2813 on the Y-axis, and cross 2817 on the Z-axis. Each sphere is divided into six regions that project exactly from the center point onto the faces of the cube. Each face is represented by a quadtree.

[0347] The procedure for constructing a position-invariant SAO is outlined in Figure 28B, which shows generation step 2861. The SAO and its six quadtrees are initialized in some orientation within the workspace in step 2863. An image of the frontier taken from within the workspace (or frontier information obtained in some other way) is appropriately projected onto one or more faces of the bounding cube and written to the quadtrees in step 2865. For this purpose, the quadtrees are treated as being at a sufficient distance from the workspace that no parallax occurs. Because the quadtrees have variable resolution, the distance used in the projection is not important.

[0348] The properties at the lower levels of the quadtree are then appropriately attributed to the higher levels (e.g., average values) in step 2867. For example, interpolation and filtering may be performed as part of this to generate interpolated nodes in the SAO. The position-invariant SAO is then initialized by projecting onto the cube of the quadtree in step 2869. In step 2871, properties from the original image are written to the appropriate properties in the position-invariant SAO node. The position-invariant SAO is then output to the scene graph in step 2873. In certain situations, all six quadtrees may not be necessary.

[0349] The use of a quadtree in this context is a convenient way to assemble various images into frontier information for interpolation, filtering, etc. Those skilled in the art will be able to devise alternative methods for accomplishing frontier generation, including those in which quadtree utilization is not used. Node characteristics will be processed and written directly from the image containing the frontier information.

[0350] To do this, the perspective projection method of image generation module 1909 is performed from the center of the SAO to the faces of the bounding cube. This is illustrated in Figure 29 in 2D. Circle 2981 (sphere in 3D) is the SAO sphere. Square 2980 is the 2D representation of the cube. Point 2979 is the center of both. Ray 2982 is the +Y boundary of the quadtree in the +X face of the square (cube in 3D). The two quadtree nodes are node 2983 and node 2985.

[0351] The octree structure is traversed and projected onto the surface. Instead of property values ​​in the SAO nodes being used to generate display values ​​as is done in image generation, the inverse operation is performed (quadtree nodes are written into octree nodes). The projection proceeds in back-to-front order with respect to the origin, with light within the quadtree first being transferred to the outer nodes of the SAO (and reduced or eliminated within the quadtree nodes). This is the opposite of the normal front-to-back order used for display.

[0352] In a simple implementation, when the lowest level node in the quadtree is reached (assuming the lowest level is currently defined or an F node has been encountered), the property values ​​(e.g., Stokes S0, S1, S2) in the quadtree (the node onto which the center of the SAO node is projected) are copied into that node. This continues until all nodes in the solid angle octree have been visited, for that face (possibly with the use of a mask to truncate subtree traversals).

[0353] This process can be performed after all frontier images have been accumulated and the quadtree properties reduced, or with a new image in the frontier. If a new value projects onto an SAO node that has already been written, it may replace the old value or be combined (e.g., averaged), or some figure of merit may be used to make the selection (e.g., move the camera closer to the frontier for one image). Similar rules are used when initially writing the frontier images into the quadtree.

[0354] The use of a quadtree for each face helps account for the range of projected pixel sizes from various camera configurations. The original pixels from the image are reduced (e.g., averaged, filtered, interpolated) to generate higher-level (lower resolution) nodes in the quadtree. A measure of the variance of the child values ​​can also be calculated and stored as a property with the quadtree node.

[0355] In advanced implementations, the reduction and processing operations required to generate a quadtree from an image can be performed simultaneously with the image input, and therefore may add negligible processing time.

[0356] During generation, if the highest level of resolution required in the SAO (for the current operation) has been reached and the lowest level of the quadtree (pixel) has been reached, the value from the current quadtree node is written into the node. If a higher quality value is desired, projection placement within the quadtree may be used to examine a larger region of the quadtree and compute a property or set of properties that includes contributions from neighboring quadtree nodes.

[0357] On the other hand, if the lowest level of the quadtree is reached before the lowest level of the octree, the subdivision is stopped and the quadtree node values ​​(derived from the original pixels from the frontier image) are written to the lowest node level of the SAO.

[0358] Generally, it is desirable to perform the projection using a relatively low level of resolution in the solid angle octree, but to store any information necessary or desired for later use in picking up operations that generate or access lower level SAO nodes, if necessary, during the execution of the operation or query.

[0359] The multi-resolution nature of SAO is exploited. For example, a measure of the spatial rate of change of illumination may be attached to the nodes (e.g., illumination gradient). Thus, if the illumination changes quickly (e.g., in angle), lower levels of the SAO octree / quadtree can be accessed or created to represent higher angular resolution.

[0360] The incident SAO generation operation is similar to the frontier SAO, except that the image generation operation of the image generation module 1909 uses the center of the incident SAO as the viewpoint and six frontier quadtrees are generated in front-to-back order. As shown in FIG. 30, an initially empty incident SAO is projected onto the quadtree 3090 from its center at point 3091. One quadtree node 3093 at a given level is shown. At the next level of subdivision, level n+1, nodes 3098 and 3099 are shown. The octree node at the current stage of the traversal is node 3092. The two bounding projections are ray 3095 and ray 3096. The span is measured from the central ray 3097. As mentioned above, a quadtree image may be a composite of multiple projected octrees with different origins and orientations using a qz buffer. The same projection of the quadtree onto the new SAO is then performed and transferred to the frontier SAO in a similar manner to generate the frontier SAO.

[0361] Since the direction from the center of the SAO to the sample points on the SAO sphere is fixed, a vector in the opposite direction (from the sphere to the center point) can also be a property at each light field SAO node that contains the sample placement, which can then be used in calculating the incident illumination to be used.

[0362] The operation of the outgoing / incoming light field processing module 2005 is shown in 2D in Figures 31A and 31B. The outgoing SAO has a center at point 3101 and a unit circle 3103 (a sphere in 3D). A particular outgoing SAO node 3105 contains property values ​​that describe light emerging in that direction from center point 3101. Using the same procedure as generating the incoming SAO, nodes in one or more octrees representing the media in the scene are traversed in front-to-back order from point 3101. In this case, the first opaque node encountered is node 3117. The projected boundary rays are ray 3107 and ray 3108. The task is to transfer light from the outgoing SAO, whose center is at point 3101, to the incoming SAO associated with node 3117.

[0363] The incident SAO for node 3117 is the SAO centered at 3115. It may already exist for node 3117 or may be created when it first encounters incident illumination. As shown, the node with an incident SAO centered at 3115 has a representation circle (sphere in 3D) at 3113.

[0364] The orientation of the coordinate system of the media octree(s) is independent of the orientation of the coordinate system of the outgoing SDAO, which is centered at 3101, as is conventional for the image generation processing module 1909. In this implementation, the coordinate systems of all SAOs are aligned. Thus, the coordinate system of the incoming SAO associated with node 3117 is aligned with the coordinate system of the outgoing SAO, which is centered at point 3101. The node in the incoming SAO used to represent illumination from point 3101 is node 3109, which is centered at point 3111.

[0365] The calculation to identify node 3109 proceeds by maintaining the arrangement of nodes in the incoming SAO, whose center is moved with the media octree, point 3115 in the figure. The incoming SAO nodes are traversed in front-to-back order from the center of the outgoing SAO, in this case point 3101, thus finding the node on the correct side. A mask may be used to remove nodes in the incoming SAO that are obscured from view and cannot receive illumination.

[0366] Traversal of an incoming SAO node is complicated by the fact that its coordinate system is generally not aligned with the coordinates of the octree to which it is attached. Therefore, movement from a node in the incoming SAO must not only consider movement relative to its own center, but also movement of the node center in the media octree. This is achieved by using the same offsets used by the media octree nodes for PUSH added to the normal offset for the SAO itself. Although the two sets of offsets will be the same at a particular level as those used by the media octree and outgoing SAO, the calculation must accumulate the sum of both offsets independently, since the order of PUSH operations and offset calculations will generally differ from either the outgoing SAO or the media octree.

[0367] When the last level for the particular operation in progress is reached, the appropriate lighting property information from the exit node, in this case node 3105, is transferred to the entrance node, in this case node 3109. The transferred lighting is properly taken into account by modifying the associated properties at the exit node, in this case node 3105. Characterization of the appropriate lighting transfer can be performed in a number of ways, such as by a projected rectangular area determined by the width and height of the projected node in the X and Y dimensions.

[0368] As the subdivision continues, Figure 31B illustrates the result, where the outgoing SAO node 3121 projects illumination along a narrow solid angle 3123 onto the incoming SAO node 3125 on a circle 3113 (sphere in 3D), representing illumination on the media node centered at 3115.

[0369] A 2D example of a BLIF SAO is shown in Figure 32. It is defined by SAO center 3215 and circle 3205 (sphere in 3D). The surface normal vector for which the weights are defined is vector 3207 (the Y axis in this case). In the case shown, outgoing rays are calculated for direction 3211. Incoming rays at the surface location are represented by incoming ray SAOs located at the same center point. In this case, rays from four directions 3201, 3203, 3209, and 3213 are shown. The SAO nodes, including their weights, are located on circle 3205 (sphere in 3D). The outgoing ray in direction 3211 is the sum of the weights for the incoming direction multiplied by the rays from that direction. In the illustration, this is the four incoming directions and four associated sets of weights.

[0370] The BLIF SAO surface normal direction generally does not correspond to the surface normal vector in certain situations. Therefore, the BLIF is appropriately rotated around its center using the geometry module 1905 to align the normal direction with the local surface normal vector and the exit direction to align with the BLIF SAO's exit vector. Overlapping nodes in the two SAOs are multiplied and summed to determine the light characteristics of the exit direction. This needs to be performed for all exit directions of interest for a particular situation. An SAO mask can be used to exclude selected directions (e.g., into the medium) from processing.

[0371] This operation is typically performed in a hierarchical manner: for example, when there is a large deviation in the BLIF coefficients within certain angular ranges, the calculations will be performed at lower levels of the tree only within those ranges, while other regions of the orientation space are more efficiently calculated at reduced levels of resolution.

[0372] According to an exemplary embodiment, the 3D image processing system 200 can be used in vehicle hail damage assessment (HDA). Hail causes significant damage to vehicles each year. For example, approximately 5,000 hailstorms occur each year in the United States, damaging approximately one million vehicles. Damage assessment currently requires visual inspection by trained inspectors who must be quickly dispatched to the affected area. The present invention is an automated assessment system that can automatically, quickly, consistently, and reliably detect and characterize vehicle hail damage.

[0373] Most hail dents are shallow, often only a fraction of a millimeter deep. Additionally, there are often subtle imperfections around the dent itself, all of which makes manual assessment of damage and repair costs difficult. Due to uncertainty and the varying levels of effort required to repair individual dents, the estimated cost of repairing a vehicle can vary significantly, depending on the number of dents, their placement on the vehicle, and their size and depth. Variations in manual assessments lead to significant variations in estimated costs for the same vehicle and damage. Reducing this variance is a key goal of automating the process.

[0374] Characterizing hail damage is difficult using traditional machine vision and laser techniques. A key problem is the glossiness of vehicle surfaces, which has little Lambertian reflectivity. They often exhibit mirror-like properties, which, combined with the shallowness of numerous hail pits, make accurate characterization difficult using these methods. However, some properties, such as glossy surfaces, allow a trained observer to decipher the resulting light patterns when moving the field of view to assess the damage and ultimately estimate repair costs. Some exemplary embodiments of the present invention provide a system capable of characterizing and then interpreting the light reflected from a vehicle's surface from multiple viewpoints. Light incident on and exiting the vehicle surface is analyzed. The method uses polarimetric image processing, the physics of light transport, and mathematical solvers to generate a scene model that represents both the media field (geometric / optical) and the light field in 3D.

[0375] In some exemplary embodiments of the invention, polarimetric cameras acquire images of vehicle panels and parts from multiple viewpoints. This can be achieved in a number of ways. In one version, a stationary vehicle is imaged by a set of polarimetric cameras mounted on a moving gantry that pauses at multiple locations. Alternatively, a UAV such as a quadcopter, perhaps tethered, can be used to carry the cameras. This is within a custom structure to protect the vehicle and equipment from the weather in addition to controlling exterior lighting. Inside the structure, non-specialized (e.g., non-polarized) lighting is used. For example, standard floodlights within a tent could be used. An illustration of a possible implementation is shown in FIG. 33. A vehicle, such as vehicle 3309, is driven into enclosure 3303 for inspection and stops. In this implementation, a quadcopter equipped with cameras and a fixed gantry 3301 with fixed cameras are used. Lights 3311 are mounted low to the ground to illuminate the sides of the vehicle. A control light unit 3307 commands the driver to enter the enclosure, when to stop, and when the scanning process is complete and the vehicle should exit. The camera is carried by a quadcopter 3305.

[0376] The system can be used to observe and model light fields in natural environments, eliminating the need for an enclosure in some situations.

[0377] After imaging the vehicle, the HDA system builds a model of the vehicle surface. This starts from image processing or other methods (including existing 3D models of the vehicle type, if available) and segments the vehicle into parts (hood, roof, doors, handles, etc.). This is later used to determine repair procedures and costs for each detected hail dent on each part or panel. For example, in many cases, dents can be repaired by skillfully tapping the dent from underneath using a specialized tool. Knowledge of the underside structure of a vehicle panel can be used to determine access to specific dents and improve repair cost estimates.

[0378] The models built from the images for each panel are combined with knowledge of the vehicle's year, model, shape, paint properties, etc. to form a dataset for analysis or input into machine learning systems for hail damage detection and characterization.

[0379] Using polarization images alone, some surface information can be obtained regarding hail damage, other damage, debris, etc. Such images can also be used to calculate accurate surface orientation vectors for each pixel, providing additional information. As mentioned above, BLIF gives the outgoing light from a surface point for a particular exit direction as a sum of weights applied to all the incoming light striking that point.

[0380] For vehicle panels, the intensity of reflected light for a particular exit direction is affected by the orientation of the surface at the point of reflection. However, surface orientation generally cannot be determined using intensity observations alone. However, using polarimetry, orientation can be resolved (up to a small set of ambiguous alternatives). Given incident light intensities from various directions, BLIF specifies the expected polarization for the exit direction as a function of surface orientation. For non-Lambertian reflecting surfaces, this typically provides enough information to uniquely distinguish the exit direction. In some BLIF formulations, the expected polarization of the incident light is also incorporated.

[0381] Thus, given the intensity (and polarization, if known) of incident light on a surface configuration, an estimated polarization of the outgoing light for all possible directions can be calculated. By comparing the Stokes vector values ​​actually seen in the image to the possible values, a surface normal for the surface at the observed configuration can be selected. The incident light field SAO and BLIF SAO can be used to efficiently calculate the polarization of the outgoing light to compare with the perceived polarization at a pixel or group of pixels. The closest match indicates the possible surface normal direction.

[0382] If the front normals vary relatively smoothly within a region, they can be combined to form a 3D surface estimated from a single image. Abruptly varying normals are typically caused by some form of edge, which indicates the boundary where the surface terminates.

[0383] In addition to some estimation of the illumination hitting the surface from various directions, this method requires an estimation of the BLIFs of the surface. Sometimes, a priori knowledge about the illumination and object can be used to initialize the BLIF estimation process. For a vehicle, the BLIFs can be known based on knowledge about the vehicle, or a known set of BLIFs can be tested individually to find the best fit.

[0384] Although single-image capabilities are useful, they have limitations: absolute positioning in 3D (or distance from the camera) cannot be determined directly, and there may be some ambiguity in the recovered surface orientation. Also, viewpoint is important: surface regions nearly perpendicular to the camera axis suffer from low SNR (signal-to-noise ratio), which can lead to loss of surface orientation and voids (holes) in the recovered surface.

[0385] The solution is to generalize the use of polarization to multiple images. Given scene polarization observed from two or more viewpoints, the scene model can be cast into a mathematical formula that can be solved as a nonlinear least-squares problem. This is used to estimate the scene light field and surface BLIFs. The position and shape of the surface of interest can then be estimated. The scene model, consisting of surface elements (surfels) and the light field, can be gradually solved (in a least-squares sense) to ultimately produce a model that best matches the observations. Then, depending on the camera viewpoint, the absolute position and orientation of each surfel can be determined voxel-by-voxel.

[0386] The HDA system process 3401 is shown in Figure 34. In polarimetric image acquisition operation 3403, a set of polarimetric images of the vehicle is acquired from multiple viewpoints within an enclosed environment with controlled lighting. In operation 3405, a scene reconstruction engine (SRE) processes the set of images to determine the camera placement and pose within each image. Relevant external information (e.g., from an inertial measurement unit) is used, and known photogrammetry methods may also be used. The SRE estimates the scene light field to the extent necessary to determine the incident light on the surface of interest. Models of the exit light from the enclosure and lighting are acquired during system setup or in later operations and are used in this process.

[0387] BLIF features are estimated for surfaces by SRE, which is based on observing polarized light from multiple viewpoints that view a surface region that is assumed to be relatively flat (negligibly small curvature) or have a nearly constant curvature. Knowing the estimated surface illumination and BLIF, the outgoing polarization is estimated for all relevant directions. For each scene voxel, SRE estimates the assumed surfel presence (voxel occupied or vacant), subvoxel surfel locations, and surfel orientations that best explain all observations of polarized light emanating from the voxel. For example, voxel and surfel properties are evaluated based on the radiometric match between polarized light rays emanating from each considered voxel. Each resulting surfel approximately represents the area of ​​the surface that projects onto a single pixel of the nearest camera, typically located a few feet away from the farthest vehicle surface point. For each surfel, depth, normal vector, and BLIF (locally differentiated) data are calculated.

[0388] Each relevant area of ​​the vehicle surface is examined to identify potential anomalies for further analysis. This information is used by an anomaly detector module in operation 3407 to detect anomalies and further by a panel segmentation module in operation 3411 to separate the vehicle into panels for later use in planning repairs.

[0389] Those skilled in the art will appreciate that many methods can be used to leverage this information to detect and characterize hail damage. In a preferred embodiment of the present invention, machine learning and / or artificial intelligence are used to improve the recognition of hail damage in the presence of other damage and debris. In operation 3409, the HDA preprocessing module, sensed normal vectors, and 3D reconstruction are used to create a dataset for use in machine learning. This could be, for example, a composite image from a starting point directly above the potential dent, a "hail view" image, supplemented with additional information.

[0390] The process begins by determining a mathematical model of the undamaged vehicle panel being examined. This can be derived from a 3D configuration of surfels (as detected in the acquired image) that do not appear to be part of damage or debris (e.g., normals change only slowly). If a nominal model of the surface is available based on the vehicle model, it can be employed here. Next, for each surfel, the deviation of the measured or calculated value from the expected value is calculated. In the case of normals, this is the deviation of the sensed normal from the expected normal at that location on the mathematical model of the surface (e.g., dot product). For sensed vectors where the deviation is significant, the direction of the vector in the local surface plane, e.g., with respect to the vehicle's coordinate system, is calculated. This gives the magnitude and direction of the deviation of the normal vector. A large deviation pointing inward from the circle indicates, for example, a dent in the wall. Other information, such as the difference in the surfel's reflective properties from those expected and the surfel's spatial deviation from the mathematical surface (e.g., depth), is also attached to each surfel. This may be encoded, for example, into an image format that combines depth, normals, and reflectance.

[0391] The next operation is to normalize and "flatten" the data. This means resizing the surfels to a uniform size and converting them into a flat array similar to the image pixel array. A local "up" direction may also be added. The predicted impact direction calculated for a dent is then compared to the assumed directions of neighboring dents and hail movement. This then forms the initial data set to be encoded into a format that can be fed to an artificial intelligence (MI) and / or database module to recognize patterns from dents caused by hail. The shapes, highlighted with associated information, form the pattern that is ultimately recognized as a hail dent by the artificial intelligence and / or database module in operation 3413.

[0392] The results are input into a repair planner module in operation 3415, where the dent information is analyzed and a repair plan is generated that is later used to estimate costs. The results of the panel segmentation module from operation 3411 are now used to analyze panel anomalies. The resulting plan and related information are then sent by an output processing module in operation 3417 to relevant organizations, such as insurance companies and repair shops.

[0393] The HDA inspection process 3501 is used to inspect vehicles for hail damage and is shown in Figure 35. It begins with detecting a vehicle being driven into an inspection station in operation 3503. In the illustrated embodiment, the vehicle stops for inspection, but those skilled in the art will know that the vehicle can also be inspected while moving, either under its own power or by some external means. An image is then acquired in operation 3505.

[0394] The SRE uses various parameters to guide operations in the reconstruction of the scene. Operation 3507 obtains settings and goal parameters related to the HDA inspection. The reconstruction is then performed in operation 3509. The results are then analyzed by an artificial intelligence and / or database module in operation 3511. Areas where anomalies are detected are examined. These are classified as caused by hail or something else, such as non-hail damage, debris such as tar, etc. Those identified as hail dents are further characterized (e.g., size, placement, depth).

[0395] The 3D model of the vehicle, including identified dents and other anomalies, is distributed to the systems of the relevant parties, possibly at geographically distributed locations, in operation 3513. Such information is used for different purposes in operation 3515. For example, an insurance company may use an in-house method to calculate the amount to be paid for repairs. Inspectors may access the 3D model to examine the situation in detail, possibly including original or synthetic images of the vehicle, vehicle panels, or individual dents. The 3D model may be viewed from any perspective to ensure a complete understanding. For example, a 3D model containing detailed surface and reflectance information may be "re-lit" with simulated lighting to mimic a real-life inspection. An insurance company may also compare historical records to determine whether a damage report to be submitted for repair has already been reported. A repair company may use the same similar information and capabilities to submit repair costs.

[0396] With reference to processes described herein, including but not limited to, process steps, algorithms, or the like, may be described or claimed in a particular order, although such processes may also be configured to work in different orders. In other words, the order or sequence of steps that may be illustratively described or claimed herein does not necessarily dictate that the steps must be performed in that order; rather, steps of processes described herein may be performed in any order possible. Furthermore, some steps may be performed simultaneously (or in parallel) although they are described or implied as occurring separately (e.g., because one step is described after the other). Furthermore, the illustration of a process by designations in drawings does not imply that the illustrated process is exclusive of other variations and modifications, nor does it imply that the illustrated process or any of its steps are required, nor does it imply that the illustrated process is preferred.

[0397] It will be understood that, as used herein, system, subsystem, service, logic, or like terms may be implemented as any suitable combination of software, hardware, firmware, and / or the like. It will also be understood that storage locations herein may be any suitable combination of disk drive devices, memory locations, solid-state drives, CD-ROMs, DVDs, tape backups, storage area network (SAN) systems, and / or other suitable tangible computer-readable storage media. It will also be understood that the techniques described herein may be accomplished by having a processor execute instructions, which may be tangibly stored on a computer-readable storage medium.

[0398] While several embodiments have been described, these embodiments are merely examples and are not intended to limit the scope of the present invention. Indeed, the novel embodiments described herein may be embodied in a variety of other forms, and various omissions, substitutions, and modifications of the forms of the embodiments described herein may be made without departing from the spirit of the present invention. The accompanying claims and their equivalents are intended to cover such forms or modifications as would fall within the scope and spirit of the present invention. [Explanation of symbols]

[0399] 101 scenes 103 Areas containing parts of trees 105 Area containing part of the cabinet 107 Area containing part of the hood 109 Area containing part of a window 111 Areas containing parts of mountains 113 Area containing flowers in a bottle 115 Area containing part of the sky 200 3D Image Processing System 201 Scene Recovery Engine (SRE) 203 Camera 205 Application Software 207 Data Communication Layer 209 Database 211 Input / Output Interface 213 processors 215 memory 217 User Interface Module 219 Job Script Processing Module 300 processes 401 Solid Angle Field 405 Volume Field 407 Volume Elements 500 Scene Models 501 Scene Model 503 Door with window 503, 505, 506 Kitchen window 507 Stove Hood 511 Cooking table 513 Scene Rough-in Pan 515 Short camera orbit arc 515 "Left-to-Right Short Baseline" Scan 517 workspaces 519, 521 Corridor 522 voxels 601 model 603 Bulk Glass Medium 605 Air-to-Glass Surfer 607 Bulk Air Medial 701 Opaque wall 703A, 703B, 703C Frontier 704 Scene Model Boundary 707 Mountain 709A, 709B, 709C Corridors 709D Corridor 713 area 715 Intermediate Space 717A, 717B, 717C Frontier 719A, 719B, 719C Corridors 801 voxels 803, 805, 807, 809, 811, 813 viewpoints 803, 807 Posture 819 Wood 821 Metal Surfer 901 Single-center unidirectional sail configuration 903 Single-center multi-directional configuration 905 Single-Center Omnidirectional Configuration 907 Single-Center Omnidirectional Isotropic Sail Configuration 911 Planar center unidirectional configuration 913 Sail Configuration 1001 Incident Light Field 1003 voxels 1005 BLIF 1007 Total Emitting Light Field 1009 (Optional) Emissive Light Field 1011 Response Light Field 1101~1115 Operational Module 1103 Plan Processing Module 1105 Scan Processing Module 1107 Scene Resolution Module 1109 Scene Display Module 1111 Sensor Control Module 1113 Spatial Processing Module 1115 Light Field Physics Processing Module 1117 Data Modeling Module 1201 Data Modeling Module 1203 Scene Modeling Module 1205 Sensor Modeling Module 1207 Medium Modeling Module 1209 Observation Modeling Module 1211 (Feature) Kernel Modeling Module 1213 Resolution Rule Module 1215 Merge Rules Module 1217 Operation Setting Module 1301 Light Field Physics Processing Module 1303 Microfaceted (Fresnel) Reflection Module 1305 Microfacet Integration Module 1307 Volumetric Scattering Module 1309 Emission and Absorption Module 1311 Polarization Modeling Module 1313 Spectral Modeling Module 1315 Adaptive Sampling Module 1400 processes 1500 Plan Processing Process 1600 scanning processes 1701 Acquisition Control Module 1703 Analog Control Module 1705 Binarization Control Module 1707 Optical System Control Module 1709 Data Transport Control Module 1711 Polarimetry Control Module 1713 Proprioceptive Sensor Control Module 1715 Motion Control Module 1800 Scene Resolution Process 1813 Integrity calculation 1819 Assumed Scene Update Process 1840 Process 1880 Process 1903 Collective Operations Module 1905 Geometry Module 1907 Generation Module 1909 Image Generation Module 1911 Filtering Module 1913 Surface Extraction Module 1915 Morphological Operation Module 1917 Connection Module 1919 Mass Properties Module 1921 Registration Module 1923 Light Field Operation Module 2001 Position-invariant light field generation module 2003 Incident light field generation module 2005 Outgoing / Incoming Light Field Processing Module 2007 Incoming / Outgoing Light Field Processing Module 2100 points 2101, 2102 Vector 2103 yen 2104 square 2105 points 2209 points 2211, 2212 rays 2213, 2214 rays 2310 area 2311, 2312 area 2320, 2321 images 2323 landmark points 2324 viewpoints 2324, 2325 Camera placement 2326 straight line 2327 straight line 2328, 2330 feature points 2329 points 2350 position 2359 Spherical Histogram 2440 X-axis 2441 Y axis 2442 curve 2443, 2444 points 2507 Register 2508 Registers 2509 Shift Register 2510 shift register 2511 Shifter 2512 adder 2515 Registration Process 2601 Coordinate Axis 2602 shaft 2603 nodes 2604 points 2653 points 2763, 2764, 2765 rays 2766 nodes 2767 points 2801 Position-invariant SAO sphere 2803 unit cube 2805, 2807, 2809 forward facing 2811 points 2813, 2815, 2817 cross 2819 Y axis 2821 X-axis 2823 Z axis 2979 points 2980 square 2,981 yen 2982 Ray of light 2983 nodes 2985 nodes 3090 quadtree 3091 points 3093 quadtree nodes 3092 nodes 3095 Ray of light 3096 Ray of light 3097 central ray 3098, 3099 nodes 3101 points 3103 Unit Circle 3105 Exit SAO node 3107 Ray of light 3108 Ray of light 3109 nodes 3111 points 3115 points 3117 nodes 3201, 3203, 3209, 3213 direction 3,205 yen 3207 Vector 3211 direction 3215 SAO center 3301 Fixed Gantry 3303 Enclosure 3305 Quadcopter 3309 Vehicles 3311 Light 3401 HDA System Processes 3403 Polarimetric Image Acquisition Operation 3501 HDA Inspection Process 3601 X-axis 3603 Z axis 3605 Origin 3607 External Display Screen 3609 Observer 3611 viewpoints 3613 length 3615 size 3617 Display Screen 3618 Octree Node Center 3619 z-score 3621 x-value 3623 Projection 3625 x' 3627 points 3701 Parent Node 3703 I-axis 3705 J axis 3707 Node Center 3709 Node center x value 3711 iVector 3713 j vector 3721 Axis 3723 Axis 3725 distance xi 3727 xj distance 3729 Node Center 3801 x Register 3803 Shift Register xi 3805 Shift Register xj 3807 Shift register xk 3809 Adder 3811 Transformed output x value 3901 Projection Vector 3903 X value x' 3905, 3907 Distance d 4001, 4003, 4005 windows 4007 nodes 4009 span 4011 span 4013 span 4021 Ray of light 4022 span 4023 Ray of light 4025 Window center ray 4027 QWIN 4029 QSPAN 4030 Node Center Offset 4031 ZONE3 4033 ZONE2 4035 ZONE1 4037 ZONE2 4039 ZONE3 4040 Node Center Offset 4051 New Origin 4053 New No-Shift Center 4055 New Lower Origin 4061 division 4063 division 4065 division 4103 NODE register 4111, 4113, 4115 DIF Registers 4117, 4119, 4121 adders 4123 Adder 4125 QSPAN ("Quarter Span") Register 4131, 4133, 4135 QS_DIF ("Quarter Span Difference") Registers 4137 Adder 4139 Adder 4140 Selector 4151 registers 4153, 4155, 4157 Shift Registers 4161 QNODE 4163 1-bit shifter 4171 Comparator 4181 Shift Register 4183 Comparator 4187 WIN register 4189 Adder 4191 Shift Register 4193 Comparator 4201, 4211, 4221 center 4205, 4215, 4225 points Distances of 4203, 4213, and 4223

Claims

1. 1. A volumetric scene reconstruction engine for creating a stored volumetric scene model of a real scene, comprising: at least one digital data processor communicatively connected to a digital data memory and a data input / output circuit, said at least one digital data processor comprising: accessing data defining digital images of a light field in a real scene including different types of media, wherein the digital images are formed by cameras from opposite poses, each digital image including image data elements defined by stored data representing a light field flux received by a photosensitive detector in the camera; processing the digital image to form a digital volumetric scene model representing the real scene, wherein the volumetric scene model (i) includes volumetric data elements defined by stored data representing one or more media characteristics, and (ii) includes solid angle data elements defined by stored data representing the bundle of the light field, adjacent volumetric data elements forming corridors, at least one of the volumetric data elements in at least one corridor representing a partially light-transmitting medium; a volumetric scene reconstruction engine configured to store the digital volumetric scene model data in the digital data memory;

2. The volumetric scene reconstruction engine of claim 1 , further comprising a digital camera configured to acquire and store the digital images while pointing the camera at different poses within the workspace of the real scene.

3. The volumetric scene reconstruction engine of claim 2, wherein the digital camera is a user-carried portable camera having at least one digital data processor configured to acquire the digital images and transmit at least some of the data to a remote data processing site where at least some of the volumetric scene model processing is performed.

4. 4. The volumetric scene reconstruction engine of claim 3, wherein the user-carried portable camera is mounted within a frame for user-worn eyeglasses.

5. The volumetric scene reconstruction engine of claim 1 , wherein the image data elements include data representing sensed flux values ​​of polarized light for a plurality of predetermined polarization properties.

6. The real scene includes a vehicle panel surface that has been damaged by hail, and the volumetric scene reconstruction engine comprises: comparing the reconstructed scene model of the hail-damaged vehicle panel with a corresponding undamaged vehicle panel surface model; 6. The volumetric scene reconstruction engine of claim 5, further configured to analyze the resulting comparison data to provide a quantitative measure of detected hail damage.

7. The volumetric scene reconstruction engine of claim 5 , wherein at least some of the volumetric data elements represent media surface elements having surface element normals determined using the polarization characteristics.

8. Surface normal gradients formed from adjacent surface element normals are used as localization feature descriptors to label corresponding features across multiple polarimetric images. The volumetric scene reconstruction engine of claim 7.

9. The volumetric scene reconstruction engine of claim 1 , wherein at least one of the corridors occupies an area between photosensitive detectors in a camera when a respective associated digital image is formed, and the volumetric data elements represent a media reflective surface.

10. 2. The volumetric scene restoration engine of claim 1, wherein the volumetric scene restoration engine is further configured to accept user identification of at least one scene restoration goal, and the process continues iteratively until the identified goal is achieved by the scene restoration engine in the volumetric scene model with a predetermined accuracy.

11. 11. The volumetric scene reconstruction engine of claim 10, wherein the volumetric scene reconstruction engine is further configured to iteratively repeat the process until an angular resolution of volumetric data elements and solid angle data elements in the volumetric scene model from a view frustum is better than an angular resolution of one arc minute.

12. The volumetric scene reconstruction engine of claim 10, wherein the images of the goal-defined object are captured at different focal plane distances.

13. The volumetric scene reconstruction engine of claim 1 , wherein the volumetric scene reconstruction engine is configured to organize and store the volumetric data elements in a spatially sorted hierarchical manner.

14. The volumetric scene reconstruction engine of claim 1 , wherein the solid angle data elements are stored in a solid angle octree (SAO) format.

15. 2. The volumetric scene reconstruction engine of claim 1, wherein the volumetric scene reconstruction engine is further configured to use a global, non-derivative, non-linear optimization method to calculate a minimum of a cost function used to form characteristics of volumetric data elements and solid angle data elements in the scene model.

16. The volumetric scene reconstruction engine of claim 1 , wherein the volumetric scene reconstruction engine is further configured to compare (a) projection images of a previously constructed digital volumetric scene model with (b) corresponding digital images of the real scene.

17. The volumetric scene reconstruction engine of claim 16 , wherein the volumetric scene reconstruction engine is further configured to divide the digital volumetric scene model into smaller, temporally independent models.

18. The volumetric scene reconstruction engine of claim 1 , wherein a plurality of said solid angle data elements are centered at a point in said volumetric scene model.

19. The volumetric scene reconstruction engine of claim 1 , wherein the volumetric scene reconstruction engine is further configured to perform a perspective projection of light within the scene using the span.

20. 10. The volumetric scene reconstruction engine of claim 1, implemented in a processor without using common division operations.

21. 1. A volumetric scene reconstruction engine for constructing a memorized volumetric scene model of an everyday scene, comprising: a digital camera; and at least one digital data processor communicatively connected to said digital camera, a digital data memory, and a data input / output circuit, said at least one digital camera and data processor comprising: acquiring digital images of a light field in a real scene including different types of media while pointing a camera in different directions from different positions within a workspace of the real scene, wherein each digital image includes image data elements defined by stored data representing one or more sensed radiometric properties of the light field including measured radiance values ​​of polarized light for at least one predetermined polarization property; accessing data defining said digital image; processing the digital image data into digital volumetric scene model data including (a) corridors of adjacent volumetric data elements representing the medium of the real scene, and (b) solid angle data elements representing the light field of the real scene, some of the volumetric elements in a plurality of corridors representing partially transparent media, wherein: the corridors extend outward from the optical sensor locations when the associated digital images are formed, and at least some of the volumetric data elements located at the distal ends of the corridors represent featureless, non-transmissive, reflective media surfaces when viewed through unpolarized light; a volumetric scene reconstruction engine configured to store the digital volumetric scene model data in the digital data memory;

22. 22. The volumetric scene reconstruction engine of claim 21, wherein the digital camera is a user-carried digital camera, and wherein the processing is performed at least in part at a remote site in data communication with the user-carried digital camera.

23. 1. A method for creating a stored volumetric scene model of a real scene, comprising: accessing data defining digital images of a light field in a real scene including different types of media, the digital images being formed by cameras from opposite poses, each digital image including image data elements defined by stored data representing a light field flux received by a photosensitive detector in the camera; processing the digital image with a scene reconstruction engine to form a digital volumetric scene model representing the real scene, the volumetric scene model including (i) volumetric data elements defined by stored data representing one or more media characteristics, and (ii) solid angle data elements defined by stored data representing the bundle of the light field, adjacent volumetric data elements forming corridors, at least one of the volumetric data elements in at least one corridor representing a partially light-transmitting medium; storing the digital volumetric scene model data in a digital data memory.

24. 24. The method of claim 23, further comprising acquiring and storing the digital images while pointing a camera at different poses within the workspace of the real scene.

25. 25. The method of claim 24, wherein a user-carried portable camera is used to acquire the digital images and transmit at least some of the data to a remote data processing site where at least some of the volumetric scene model processing is performed.

26. 26. The method of claim 25, wherein the user-carried portable camera is mounted within a frame for user-worn eyeglasses.

27. 24. The method of claim 23, wherein the digital image data elements include data representing sensed flux values ​​of polarized light for a plurality of predetermined polarization characteristics.

28. The real scene includes a vehicle panel surface that has been damaged by hail, and the method includes: comparing the reconstructed scene model of the hail-damaged vehicle panel with a corresponding undamaged vehicle panel surface model; 28. The method of claim 27, further comprising analyzing the resulting comparison data to provide a quantitative measure of detected hail damage.

29. 28. The method of claim 27, wherein at least some of the volumetric data elements represent media surface elements having surface element normals determined using the polarization characteristics.

30. 30. The method of claim 29, wherein surface normal gradients formed from adjacent surface element normals are used as localization feature descriptors to label corresponding features across multiple polarimetric images.

31. 24. The method of claim 23, wherein at least one of the corridors occupies an area between photosensitive detectors in a camera when a respective associated digital image is formed, and the volumetric data elements represent a media reflective surface.

32. 24. The method of claim 23, further comprising user identification of at least one scene restoration goal, wherein the process continues iteratively until the identified goal is achieved by the scene restoration engine in the volumetric scene model with a predetermined accuracy.

33. 33. The method of claim 32, wherein the processing steps are iteratively repeated until the angular resolution of volumetric data elements and the angular resolution of solid angle data elements in the volumetric scene model from a view frustum is better than 1 arcminute angular resolution.

34. 33. The method of claim 32, wherein the images of the goal-defining object are captured at different focal plane distances.

35. 24. The method of claim 23, wherein said storing step organizes and stores said volumetric data elements in a spatially sorted hierarchical manner.

36. 24. The method of claim 23, wherein the solid angle data elements are stored in a solid angle octree (SAO) format.

37. 24. The method of claim 23, wherein the processing step uses a global, non-derivative, non-linear optimization method to calculate a minimum of a cost function used to form characteristics of volumetric and solid angle data elements in the scene model.

38. 24. The method of claim 23, wherein the processing step further comprises comparing (a) projection images of the previously constructed digital volumetric scene model with (b) respective corresponding digital images of the real scene.

39. 39. The method of claim 38, wherein said processing step includes dividing said digital volumetric scene model into smaller temporally independent models.

40. 24. The method of claim 23, wherein a plurality of the solid angle data elements are centered at a point in the volumetric scene model.

41. 24. The method of claim 23, wherein the processing step includes performing a perspective projection of light in the scene using the span.

42. 24. The method of claim 23, wherein the method is implemented in a processor without using a general division operation.

43. 1. A method for constructing a memorized volumetric scene model of an everyday scene, comprising: acquiring digital images of a light field in a real scene including different types of media, from different positions within a workspace of the real scene while pointing a camera in different directions, each digital image including image data elements defined by stored data representing one or more sensed radiometric properties of the light field including measured radiance values ​​of polarized light for at least one predetermined polarization property; accessing data defining said digital image; processing the digital image data by a scene reconstruction engine into digital volumetric scene model data including (a) corridors of adjacent volumetric data elements representing the medium of the real scene, and (b) solid angle data elements representing the light field of the real scene, some of the volumetric elements in a plurality of corridors representing partially transparent media; the corridors extend outward from the optical sensor locations when the associated digital images are formed, and at least some of the volumetric data elements located at the distal ends of the corridors represent featureless, non-transmissive, reflective media surfaces when viewed through unpolarized light; storing the digital volumetric scene model data in a digital data memory.

44. 44. The method of claim 43, wherein said obtaining step is performed at a user-carried digital camera, and said processing step is performed at least in part at a remote site in data communication with said user-carried digital camera.

45. A non-volatile digital memory medium containing executable computer program instructions, the executable computer program instructions, when executed on at least one processor, accessing data defining digital images of a light field in a real scene including different types of media, the digital images being formed by cameras from opposite poses, each digital image including image data elements defined by stored data representing a light field flux received by a photosensitive detector in the camera; processing the digital image to form a digital volumetric scene model representing the real scene, the volumetric scene model including (i) volumetric data elements defined by stored data representing one or more media characteristics, and (ii) solid angle data elements defined by stored data representing the bundle of the light field, adjacent volumetric data elements forming corridors, at least one of the volumetric data elements in at least one corridor representing a partially light-transmitting medium; and storing said digital volumetric scene model data in a digital data memory.

46. The computer program instructions, when executed, 46. ​​The non-volatile digital memory medium of claim 45, providing an interface with a digital camera for further acquiring and storing said digital images while pointing the camera at different poses within the workspace of said real scene.

47. 47. The non-volatile digital memory medium of claim 46, wherein a user-carried portable camera is used to acquire the digital images and transmit at least some of the data to a remote data processing site where at least some of the volumetric scene model processing is performed.

48. 48. The non-volatile digital memory medium of claim 47, wherein the user-carried portable camera is mounted within a frame for user-worn eyeglasses.

49. 46. ​​The non-volatile digital memory medium of claim 45, wherein the digital image data elements include data representing sensed flux values ​​of polarized light for a plurality of predetermined polarization characteristics.

50. The real scene includes a vehicle panel surface that has been damaged by hail, and the computer program instructions, when executed, comparing the reconstructed scene model of the hail-damaged vehicle panel with a corresponding undamaged vehicle panel surface model; 50. The non-volatile digital memory medium of claim 49, further comprising: analyzing the resulting comparison data to provide a quantitative measure of detected hail damage.

51. 50. The non-volatile digital memory medium of claim 49, wherein at least some of said volumetric data elements represent media surface elements having surface element normals determined using said polarization characteristics.

52. 52. The non-volatile digital memory medium of claim 51, wherein a surface normal gradient formed from adjacent surface element normals is used as a localization feature descriptor for labeling corresponding features across multiple polarimetric images.

53. 46. ​​The non-volatile digital memory medium of claim 45, wherein at least one of the corridors occupies an area between photosensitive detectors in a camera when an associated digital image is formed, and the volumetric data elements represent a reflective surface of the medium.

54. 46. ​​The non-volatile digital memory medium of claim 45, wherein the program instructions, when executed, further cause user identification of at least one scene restoration goal, and the process continues iteratively until the identified goal is achieved with a predetermined accuracy in the volumetric scene model.

55. 55. The non-volatile digital memory medium of claim 54, wherein the process is iteratively repeated until the angular resolution of volumetric data elements and the angular resolution of solid angle data elements in the volumetric scene model from a view frustum is better than one arcminute angular resolution.

56. 55. The non-volatile digital memory medium of claim 54, wherein the images of the goal-defined object are captured at different focal plane distances.

57. The computer program instructions, when executed, are spatially sorted and hierarchically 46. ​​The non-volatile digital memory medium of claim 45, wherein said volumetric data elements are organized and stored in a manner that:

58. 46. ​​The non-volatile digital memory medium of claim 45, wherein the solid angle data elements are stored in a solid angle octree (SAO) format.

59. 46. ​​The non-volatile digital memory medium of claim 45, wherein the process uses a global, non-derivative, non-linear optimization method to calculate a minimum of a cost function used to form characteristics of volumetric and solid angle data elements in the scene model.

60. 46. ​​The non-volatile digital memory medium of claim 45, wherein the processing includes comparing (a) projection images of a previously constructed digital volumetric scene model with (b) respective corresponding digital images of the real scene.

61. 61. The non-volatile digital memory medium of claim 60, wherein said processing includes dividing said digital volumetric scene model into smaller temporally independent models.

62. 46. ​​The non-volatile digital memory medium of claim 45, wherein a plurality of said solid angle data elements are centered at a point in said volumetric scene model.

63. 46. ​​The non-volatile digital memory medium of claim 45, wherein the processing includes performing a perspective projection of light within the scene using the span.

64. 46. ​​The non-volatile digital memory medium of claim 45, wherein the computer program instructions are executed within a processor without the use of a common division operation.

65. A non-volatile digital memory medium containing executable computer program instructions that, when executed on at least one processor, construct a stored volumetric scene model of an everyday scene, the program instructions, when executed, acquiring digital images of a light field in a real scene including different types of media, while pointing a camera in different directions from different positions within a workspace of the real scene, each digital image including image data elements defined by stored data representing one or more sensed radiometric properties of the light field including measured radiance values ​​of polarized light for at least one predetermined polarization characteristic; accessing data defining said digital image; processing the digital image data into digital volumetric scene model data including (a) corridors of adjacent volumetric data elements representing the medium of the real scene, and (b) solid angle data elements representing the light field of the real scene, some of the volumetric elements in a plurality of corridors representing partially transparent media; processing the corridors, each extending outward from a light sensor location when an associated digital image is formed, and at least some of the volumetric data elements located at a distal end of the corridor represent a featureless, non-transmissive, reflective medium surface when viewed through unpolarized light; and storing said digital volumetric scene model data in a digital data memory.

66. The step of obtaining is performed in a user-carried digital camera, and the processing is performed in the user's 66. The non-volatile digital memory medium of claim 65, executed at least in part at a remote site in data communication with a user-held digital camera.

Citation Information

Patent Citations

  • Image supported surface reconstruction procedure uses combined shape from shading and polarisation property methods

    DE102004062461A1

  • Static scene reconstruction method for testing work pieces, involves reconstructing scene of a two-dimensional image data by method of shape from motion when scene exhibits region with intensity gradients

    DE102006013318A1

  • High-speed image generation of complex solid objects using octree encoding

    US4694404A

  • Method of and apparatus for obtaining object data by machine vision form polarization information

    US5028138A

  • Method for processing spatial-phase characteristics of electromagnetic energy and information conveyed therein

    US6810141B2