Method for extracting the geometry of a represented object from a radiance field

The procedure for extracting object geometry from neuronal radiance fields improves efficiency and precision by focusing on the object and reducing background noise through the use of virtual camera beams and Poisson meshing.

DE102023210949A1Pending Publication Date: 2025-05-08ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102023210949
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Existing methods for extracting object geometry from neuronal radiance fields (nerf) are inefficient and prone to errors, particularly when dealing with noisy fields and backgrounds, which limits the quality of the produced mesh.

Method used

A procedure that involves generating a radiance field based on image data, determining the termination distance of virtual camera beams, reconstructing the object surface, and extracting the geometry from the radiance field using Poisson meshing, thereby focusing on the object and reducing background noise.

Benefits of technology

This approach enables efficient and precise extraction of object geometry from incomplete radiance fields, improving the quality of the mesh and reducing computational resources by focusing on the object and ignoring irrelevant background information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method (100) for extracting a geometry of a represented object (1) from a radiance field, comprising the following steps: - Providing a radiance field, wherein the radiance field is generated based on image data, wherein the radiance field includes information about the geometry of the displayed object (1), - Determining a termination distance of at least one virtual camera beam, wherein the termination distance is specific for a distance at which an impact of the virtual camera beam on an object representation of the represented object (1) is determined, - Reconstructing a surface of the represented object (1) based on the determined termination distance of the at least one virtual camera beam, - Extracting the geometry of the represented object (1) from the radiance field based on the reconstructed surface. Furthermore, the invention relates to a computer program, a device and a storage medium for this purpose.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for extracting the geometry of a displayed object from a radiance field. Furthermore, the invention relates to a computer program, a device, and a storage medium for this purpose. State of the art

[0002] Neural radiance fields are a novel concept in computer visualization and 3D reconstruction that uses deep neural networks to represent the scene. NERF can generate fine details and realistic representations of complex scenes with high resolution by capturing information about the color and density of light rays at different positions and directions. The method has proven particularly useful for generating realistic 3D scenes from a limited number of 2D images.

[0003] For use in classic rendering pipelines (e.g., Blender), it is possible to extract scene geometries and textures from neural radiance fields (NERFs). Typically, an entire scene is learned, and the area of ​​the object to be extracted is subsequently processed. If NERFs are to be used specifically to digitize an object, the object is photographed or filmed from all sides as training material, for example, while freestanding. This approach has two significant disadvantages: The NERF learns not only the object but also the background, and the capacity / resources of the neural network approximating the radiance field are distributed between the object and the background scene. Thus, more is learned than is actually necessary.

[0004] Extracting objects, especially as meshes, from radiance fields is a necessary step for using implicitly learned assets in traditional rendering pipelines. Previous methods divide the probability density of the radiance field into discrete volume elements (voxels), and then generate a mesh using so-called marching cubes. This approach is slow and is particularly problematic with noisy radiance fields, which require additional filtering and limit the quality of the generated mesh. Furthermore, neural radiance fields (NERFs), as an implicit numerical solution to an optimization problem, are never completely noise-free and, by their very nature, are subject to noise and artifacts, particularly in peripheral regions with poor training coverage. Disclosure of the invention

[0005] The subject matter of the invention is a method having the features of claim 1, a computer program having the features of claim 8, a device having the features of claim 9, and a computer-readable storage medium having the features of claim 10. Further features and details of the invention emerge from the respective subclaims, the description, and the drawings. Features and details described in connection with the method according to the invention naturally also apply in connection with the computer program according to the invention, the device according to the invention, and the computer-readable storage medium according to the invention, and vice versa, so that with regard to the disclosure of the individual aspects of the invention, reference is or can always be made to each other.

[0006] The invention particularly relates to a method for extracting a geometry of a displayed object from a radiance field, comprising the following steps, wherein the steps can be carried out repeatedly and / or sequentially: - Providing a radiance field, wherein the radiance field is generated on the basis of image data, wherein the radiance field comprises information about the geometry of the displayed object, - Determining a termination distance of at least one virtual camera beam, wherein the termination distance is specific for a distance at which an impact of the virtual camera beam on an object representation of the displayed object is determined, - Reconstructing a surface of the displayed object based on the determined termination distance of the at least one virtual camera beam, - Extracting the geometry of the displayed object from the radiance field based on the reconstructed surface.

[0007] The information about the geometry of the displayed object can be abstract information on the basis of which the geometry can be extracted from the radiance field by the method. In simple terms, the termination distance can be a distance from a starting point of the virtual camera beam to a point at which the virtual camera beam impinges on the object representation of the displayed object. The object representation can be an image of the displayed object in a simulation or in a virtual model. The at least one virtual camera beam can have a specific starting point and a specific direction. In particular, the at least one virtual camera beam can be simulated on the basis of the starting point and the direction. The surface of the displayed object can advantageously be reconstructed on the basis of the starting point, the direction and the termination distance.The reconstructed surface can, for example, be a point cloud. Extracting the geometry of the displayed object from the radiance field based on the reconstructed surface can involve creating a mesh based on this point cloud of the reconstructed surface. This can be done using a so-called Poisson meshing technique.

[0008] Within the scope of the invention, the radiance field can describe a color of emitted light and a probability density for each point in a space represented by the image data, and for each viewing direction in the represented space. The probability density represents, in particular, a probability that a photon will be absorbed or emitted.

[0009] The at least one virtual camera beam can be randomly selected and rendered based on the radiance field, wherein the rendering is, in particular, a volume rendering. The termination distance can be determined based on the rendered virtual camera beam. The at least one virtual camera beam can be randomly selected, in particular with regard to a starting point and a direction.

[0010] The image data can result from detection by at least one sensor, in particular at least one camera sensor. The image data can also be, for example, images from a radar sensor and / or an ultrasonic sensor and / or a LiDAR sensor and / or a thermal imaging camera. Accordingly, the images can also be embodied as radar images and / or ultrasonic images and / or thermal images and / or LiDAR images. The image data can comprise two-dimensional images of the object in a scene from different viewing directions. Furthermore, the image data can comprise associated camera poses and camera models of a camera used to acquire the image data. A camera model serves, in particular, to provide an exact mathematical description of the imaging ratios of the camera used.

[0011] The termination distance can be determined by at least two virtual camera beams in order to reconstruct the surface of the displayed object based on the determined termination distances of the at least two virtual camera beams. In particular, the termination distance of a plurality of virtual camera beams can advantageously be determined in order to be able to infer the surface of the displayed object based on the plurality of virtual camera beams.

[0012] It is also advantageous if determining the termination distance includes the following steps: - Determining a termination probability for at least one segment of the at least one virtual camera beam on the basis of a probability density of points in the at least one segment predetermined by the radiance field, - Determining the termination distance based on a comparison of the determined termination probability of the at least one segment with a defined threshold value.

[0013] The termination probability for a respective segment could, for example, be determined based on an integral based on the probability densities of the points in the respective segment. The defined threshold is preferably 0.5. In other words, if the termination probability in a segment of a respective virtual camera beam exceeds a value of the defined threshold, in particular 0.5, then the termination distance is determined based on a distance from a starting point of the respective virtual camera beam to this segment.

[0014] Furthermore, within the scope of the invention, it can be provided that the provision of the radiance field is carried out using a machine learning model and comprises the following steps: - Training the machine learning model based on the image data in order to generate the radiance field using the trained machine learning model based on the image data.

[0015] The machine learning model is preferably a neural network, particularly preferably a Neural Radiance Field (NeRF). The radiance field is generated in particular by the machine learning model modeling a radiance and a probability density value for each point in a 3D space represented by the image data. The training of the machine learning model can be carried out as follows. The input for the machine learning model is in particular the image data, which preferably depicts a scene from different perspectives. A virtual training camera beam can be projected for each point in the 3D space. The machine learning model then takes in particular the 3D coordinates of this point and the viewing direction of the virtual training camera beam as input. The machine learning model then preferably returns a probability density and a radiance for each such point.The probability density preferably indicates how likely it is that a visible object exists at that point, and the radiance preferably indicates the color of the light at that point. To generate a 2D image from this 3D data, a method called "volume rendering" can be used. In particular, this involves tracing the virtual training camera ray through 3D space, combining the probability density and radiance values ​​along the ray to determine the final color of the pixel in the 2D image. The machine learning model is now trained, in particular, to minimize the difference between predicted 2D images and the actual input images, i.e., the image data. During this training process, weights within the machine learning model are adjusted to increase prediction accuracy.Once the machine learning model has been trained, it can be used advantageously to render images of the scene from new, previously unseen angles.

[0016] Advantageously, the invention can provide for the following steps to be carried out during the training of the machine learning model: - Determining the termination distance of at least one virtual training camera beam, - Determining a termination position of the at least one virtual training camera beam based on the determined termination distance and an orientation direction of the at least one virtual training camera beam,

[0017] The at least one virtual training camera beam can be disregarded for further training of the machine learning model if the determined termination position indicates that the at least one virtual training camera beam terminates outside a defined volume. The defined volume can, for example, be an area described by sensors, in particular cameras, which have acquired the image data with the displayed object from different directions. In other words, the defined volume is preferably a volume within a radius of said sensors. Advantageously, by disregarding the at least one training camera beam, at least one irrelevant background object that lies outside the defined volume can no longer be considered in the training.This can accelerate further training of the machine learning model because less data needs to be processed and the machine learning model does not have to additionally learn a possible background object.

[0018] Furthermore, it can be provided that the following steps are carried out during the training of the machine learning model: - Determining the termination probability for at least one segment of at least one virtual training camera beam on the basis of the probability density of points in the at least one segment specified by the radiance field, - Adjusting a sampling range of the at least one virtual training camera beam based on the determined termination probability for the at least one segment.

[0019] Based on an analysis of the termination probability for at least one segment, an approximate area can be determined in which the displayed object is highly likely to be present, since the virtual training camera beam terminates there with a high probability. The preceding steps can be performed multiple times so that the area becomes progressively more precise. Furthermore, an additional vacuum area can be defined between a starting point of the virtual training camera beam and a beginning of the respective sampling area, in which no sampling takes place. The termination probability is advantageously determined for a plurality of segments of the respective virtual training camera beam in order to enable a more differentiated analysis.By adjusting the sampling range, further training can be advantageously performed with reduced computing power requirements and also contribute to a better-trained machine learning model, as only more relevant regions—i.e., regions in which the depicted object is likely to be present—are considered. The sampling range can be a number and distribution of points rendered during training.

[0020] A further advantage within the scope of the invention can be achieved if the at least one virtual training camera beam has a plurality of segments and the adaptation of the sampling range comprises the following step: - Restricting the sampling range to at least one segment of the plurality of segments whose termination probability is above a defined limit.

[0021] Thus, only a range of segments that are likely to be close to the displayed object can be advantageously taken into account during training using samples.

[0022] The invention also relates to a computer program, in particular a computer program product, comprising instructions that, when executed by a computer, cause the computer to execute the method according to the invention. Thus, the computer program according to the invention provides the same advantages as those described in detail with reference to a method according to the invention.

[0023] The invention also relates to a data processing device configured to carry out the method according to the invention. The device can be, for example, a computer that executes the computer program according to the invention. The computer can have at least one processor for executing the computer program. A non-volatile data memory can also be provided, in which the computer program is stored and from which the computer program can be read by the processor for execution.

[0024] The invention may also provide a computer-readable storage medium that has the computer program according to the invention and / or includes instructions that, when executed by a computer, cause the computer to carry out the method according to the invention. The storage medium is designed, for example, as a data storage device such as a hard disk and / or a non-volatile memory and / or a memory card. The storage medium can, for example, be integrated into the computer.

[0025] Furthermore, the method according to the invention can also be implemented as a computer-implemented method.

[0026] Further advantages, features, and details of the invention will become apparent from the following description, which describes embodiments of the invention in detail with reference to the drawings. The features mentioned in the claims and in the description may be essential to the invention individually or in any combination. They show: Fig. 1 a schematic visualization of a method, a device, a storage medium and a computer program according to embodiments of the invention, Fig. 2 a schematic representation of a structure for carrying out the method according to the invention, Fig. 3 a schematic visualization of an exact solution for a depth measurement according to embodiments of the invention, Fig. 4 a schematic visualization of a realistic, additionally perturbed solution for a depth measurement according to embodiments of the invention.

[0027] In Fig. 1, a method 100, a device 10, a storage medium 15 and a computer program 20 according to embodiments of the invention are schematically shown.

[0028] Fig. 1 shows, in particular, an exemplary embodiment of a method 100 for extracting a geometry of a displayed object 1 from a radiance field. In a first step 101, a radiance field is provided, wherein the radiance field is generated on the basis of image data, wherein the radiance field comprises information about the geometry of the displayed object 1. In a second step 102, a termination distance of at least one virtual camera beam is determined, wherein the termination distance is specific to a distance at which an impact of the virtual camera beam on an object representation of the displayed object 1 is determined. In a third step 103, a surface of the displayed object 1 is reconstructed based on the determined termination distance of the at least one virtual camera beam.In a fourth step 104, the geometry of the displayed object 1 is extracted from the radiance field based on the reconstructed surface.

[0029] Fig. 2 shows a setup with which the method can be carried out according to exemplary embodiments. Various sensors 3, in particular cameras, are arranged around an object 1 to be displayed in order to capture the image data. A volume 4 is defined around the area enclosing the sensors 3, which volume is particularly relevant for extracting the displayed object 1. Furthermore, two differently aligned virtual training camera beams 2 are shown, one of the virtual training camera beams 2 terminating at the displayed object 1 and the other virtual training camera beam 2 terminating at a background object 5.

[0030] The invention according to exemplary embodiments serves in particular for the efficient and accurate extraction of the geometry of a represented object 1 as well as for robust distance measurement from noisy radiance fields. The invention according to exemplary embodiments further describes a more efficient and accurate learning of implicit neural representations of objects 1 for the purpose of geometry and material extraction. The invention according to exemplary embodiments also describes an acceleration of the training and an increase in the achievable quality.

[0031] One aspect of the method is, in particular, a highly performant and precise method for extracting geometry from incomplete radiance fields, based on randomly sampled virtual camera rays and their expected termination points, as well as a depth measurement tailored to the noisy nature of the radiance fields. The method according to exemplary embodiments can perform the following steps for this purpose. Virtual camera rays can be randomly selected within the environment of the object to be extracted and with random directions. Furthermore, virtual camera rays can be filtered out if a probability density above a threshold value exists in the environment of the origin. A positive density implies, in particular, spatial proximity to an object or an area with significant interference; both of these can disqualify the virtual camera beam.Furthermore, the virtual camera beams can be rendered using a volume renderer, and the termination distance at which the virtual camera beam is terminated can be determined. In contrast to the usual NERF depth measurement using the expected value, a depth measurement can be performed using the transmissibility limit of 0.5, which allows noise to be skipped to a certain extent, which is particularly useful in the . Fig. 3 and Fig. 4 is shown. Fig. 3 shows, for example, an exact solution based on a graph of a probability density W and a transmissibility T, in which the probability density W is a Dirac delta function and the distance measurement to the displayed object 1 via the expected beam length and via a method based on the transmissibility T are identical and both are exact. In Fig.Figure 4 shows an exemplary numerical solution using a further graph of a probability density W and a transmissibility T. The probability density W of the medium is particularly smooth and has a maximum at the edge of the displayed object 1 as well as a disturbance at a distance of 0.3. A distance measurement per expected beam length via the transmissibility T is therefore too short, and the method according to embodiments of the invention remains correct as long as the disturbances do not push the transmissibility below the threshold of 0.5 in front of the actual object 1. Furthermore, virtual camera beams can be filtered whose termination probability at the end of the virtual camera beam is still below 1.0, since they cannot represent a surface point. Furthermore, Poisson meshing can be performed to generate a mesh from the resulting dense point cloud.

[0032] Furthermore, the method according to embodiments comprises training in which, after a machine learning model has roughly learned the geometry, depth information, i.e. termination distance, of the virtual training camera rays is compared with the defined volume 4 enclosed by the camera centers after each training batch, and virtual training camera rays that terminate outside the defined volume 4 are deleted from the training data or are no longer taken into account in further training. This enables faster training with higher quality. Furthermore, it is also possible to proceed without a photo box with a green screen or segmentation such as classic structure from motion (SFM)-based methods. Conversely, the method according to embodiments can even benefit from a detailed scene background.During the training, the background can be increasingly masked out using the mechanism described, so that the procedure focuses particularly on the displayed object 1.

[0033] Furthermore, a sampling strategy according to embodiments of the invention is described, which is based on concentrating a sampling area per training pixel and associated visual ray during training depending on the density profile along a virtual training camera ray 2. This can be done in particular as soon as the machine learning model has learned a geometric understanding of the scene, thereby accelerating the convergence rate with a relatively low sample count. The geometric understanding can, for example, be present after a defined number of training epochs. Furthermore, a coarse prediction network can advantageously be completely dispensed with, which further reduces the computing time. A further advantage of the method according to embodiments lies in particular in the considerably faster convergence without any loss of quality.Furthermore, the method according to embodiments can advantageously be directly applied to many NERF variants to improve the convergence behavior.

[0034] (T) o,v (x) denotes in particular the probability that a photon starting from the origin o in direction v has not yet been absorbed at a distance of x.

[0035] In particular, (T) o,v (0) = 1 and (T) o,v monotonically decreasing.

[0036] After each training epoch, the sampling range of the virtual training camera beam 2 is preferably adjusted depending on T for the next training epoch. To prevent the machine learning model from storing information in areas that are no longer sampled, an additional vacuum region can be introduced between the camera center, i.e., in particular, a starting point of a respective virtual training camera beam 2, and the beginning of a respective sampling range of the virtual training camera beam 2. This approach can simplify the machine learning model architecture because a prediction network is no longer required during training. Furthermore, it can advantageously simplify the training itself because the costly addition of sampling points is no longer necessary. A further advantage can be that the machine learning model converges more quickly due to the shrinking sampling ranges.

[0037] For inference, an intermediate training state can be used as a prediction network to avoid costly oversampling, since, for example, the sampling ranges are not available during inference. This allows the inference step to be identical in terms of computation time to classic NERFs.

[0038] A further aspect of the method according to embodiments is that the expected value of the termination distance is not determined directly via the volumetric rendering weights, but rather uses the distance at which the virtual camera ray exceeds a termination probability of 0.5. This can advantageously provide greater robustness in the case of perturbed radiance fields.

[0039] According to one embodiment of the invention, the following steps could be provided. First, a volumetric rendering can be performed in the direction v starting from the origin up to the termination distance f. Further, in the sense of formula (D), o,v := min x (T) o,v (x) > 0.5, it can be checked whether the termination probability is above the value of 0.5. Furthermore, it can be assumed that if T is not below 0.5, it is implied that no object blocks the virtual camera beam up to the termination distance f, therefore an invalid measurement may be present.

[0040] The above explanation of the embodiments describes the present invention exclusively by way of examples. Of course, individual features of the embodiments can be freely combined with one another, provided they are technically feasible, without departing from the scope of the present invention.

Claims

[1] Method (100) for extracting a geometry of a displayed object (1) from a radiance field, comprising the following steps: - providing (101) a radiance field, wherein the radiance field is generated on the basis of image data, wherein the radiance field comprises information about the geometry of the displayed object (1), - determining (102) a termination distance of at least one virtual camera beam, wherein the termination distance is specific for a distance at which an impact of the virtual camera beam on an object representation of the displayed object (1) is determined, - reconstructing (103) a surface of the displayed object (1) on the basis of the determined termination distance of the at least one virtual camera beam, - Extracting (104) the geometry of the displayed object (1) from the radiance field based on the reconstructed surface. [2] Method (100) according to claim 1, characterized by , that the radiance field describes for each point in a space represented by the image data and for each viewing direction in the represented space to a respective point a colour of an emitted light as well as a probability density of the respective point, wherein the at least one virtual camera beam is randomly selected and rendered based on the radiance field, wherein the rendering is in particular a volume rendering in order to determine the termination distance based on the rendered virtual camera beam, wherein the image data result from a detection by at least one sensor (3), in particular at least one camera sensor, wherein the termination distance is determined by at least two virtual camera beams in order to reconstruct the surface of the displayed object (1) based on the determined termination distances of the at least two virtual camera beams. [3] Method (100) according to claim 1 or 2, characterized by that determining the termination distance involves the following steps: - Determining a termination probability for at least one segment of the at least one virtual camera beam on the basis of a probability density of points in the at least one segment predetermined by the radiance field, - Determining the termination distance based on a comparison of the determined termination probability of the at least one segment with a defined threshold value. [4] Method (100) according to one of the preceding claims, characterized by that the provision of the radiance field is carried out using a machine learning model and includes the following steps: - Training the machine learning model based on the image data in order to generate the radiance field using the trained machine learning model based on the image data. [5] Method (100) according to claim 4, characterized by that the following steps are performed during the training of the machine learning model: - determining the termination distance of at least one virtual training camera beam (2), - Determining a termination position of the at least one virtual training camera beam (2) on the basis of the determined termination distance and an orientation direction of the at least one virtual training camera beam (2), wherein the at least one virtual training camera beam (2) is disregarded for the further training of the machine learning model if the determined termination position indicates that the at least one virtual training camera beam (2) terminates outside a defined volume (4). [6] Method (100) according to one of claims 4 or 5, characterized by that the following steps are performed during the training of the machine learning model: - determining the termination probability for at least one segment of at least one virtual training camera beam (2) on the basis of the probability density of points in the at least one segment specified by the radiance field, - adjusting a sampling range of the at least one virtual training camera beam (2) based on the determined termination probability for the at least one segment. [7] Method (100) according to claim 6, characterized by that the at least one virtual training camera beam (2) has a plurality of segments and the adjustment of the sampling area comprises the following step: - Restricting the sampling range to at least one segment of the plurality of segments whose termination probability is above a defined limit. [8] Computer program (20) comprising instructions which, when the computer program (20) is executed by a computer (10), cause the computer (10) to carry out the method (100) according to one of the preceding claims. [9] Device (10) for data processing, which is arranged to carry out the method (100) according to one of claims 1 to 7. [10] A computer-readable storage medium (15) comprising instructions which, when executed by a computer (10), cause the computer (10) to carry out the steps of the method (100) according to any one of claims 1 to 7.