Oriented-grid encoder for 3D implicit representation

By training the encoder and decoder of the neural network, multi-resolution grid features are generated and decoded to the object surface distance, the problem of failure to fully consider the geometric characteristics of the object in the prior art is solved, and a more comprehensive three-dimensional implicit expression and better continuity are achieved.

JP2025071766APending Publication Date: 2025-05-08MITSUBISHI ELECTRIC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024105877
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-23
Filing Date
2024-07-01
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

The prior art fails to fully consider the geometric characteristics of the object when generating implicit expressions of three-dimensional scenes, resulting in insufficient expression and poor continuity.

Method used

By training a neural network that includes encoder and decoder, the encoder encodes 3D point cloud data into multi-resolution grid features, while the decoder decodes from any 3D scene point to the object surface distance to generate an implicit expression.

Benefits of technology

A more comprehensive geometric feature capture and implicit expression of three-dimensional objects is achieved, improving the continuity and accuracy of expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025071766000001_ABST
    Figure 2025071766000001_ABST
Patent Text Reader

Abstract

To provide an artificial intelligence system for producing an implicit representation of a three-dimensional (3D) scene including a 3D object by training a neural network.SOLUTION: A neural network includes an encoder configured for encoding data indicative of a 3D point cloud of a shape of an object into grid-based features capturing multiple resolutions of the object, and a decoder. A system comprises a processor, and a memory having instructions stored thereon. The instructions cause the processor to: (i) receive input data indicative of an oriented point cloud of a 3D scene including a 3D object, the input data indicating 3D locations of points of the 3D point cloud and orientations of the points defining a normal to a surface of the 3D object at locations proximate to the 3D locations of the points; and (ii) train the encoder and the decoder to produce an implicit representation of the 3D object.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure is generally directed to methods and systems for generating an implicit representation of a three-dimensional scene. [Background technology]

[0002] There are many different ways to represent three-dimensional (3D) surfaces. In an implicit surface representation, a point with coordinates x, y, and z belongs to an object if F(x,y,z)=0, and the function F(.) defines the object. This type of 3D representation is advantageous because it is concise and guarantees continuity. Most learning-based 3D implicit representations start by encoding the 3D points, then decode those features into a chosen representation, defining F(.).

[0003] Two kinds of encoders are usually used in parallel: (1) those that map only the 3D coordinates of each point into a higher dimensional vector space, denoted here as positional encoders, and (2) those that gather information about their neighbors, called lattice-based 3D points. Multilayer perceptrons (MLPs) are usually considered the preferred choice for the decoder. Previous techniques that use geometric encoders do not take into account some of the underlying geometric properties of the object, exploiting only its spatial localization, and therefore form unsatisfactory 3D representations. Summary of the Invention

[0004] Thus, there is a continuing, unmet need for methods and systems that consider all of the underlying geometric features of an object when generating an implicit surface representation.Various embodiments and implementations are directed to methods and systems for generating an implicit representation of a three-dimensional (3D) scene that includes a 3D object by training a neural network, the network including an encoder configured to encode data indicative of a 3D point cloud of the shape of the 3D object into lattice-based features that capture multiple resolutions of the object, and a decoder configured to decode the lattice-based features into a distance from any point in the 3D scene to the object.

[0005] According to one aspect, an artificial intelligence (AI) system is provided that generates an implicit representation of a three-dimensional (3D) scene including a 3D object by training a neural network including an encoder configured to encode data indicative of a 3D point cloud of a shape of the object into lattice-based features capturing multiple resolutions of the object, and a decoder configured to decode the lattice-based features into a distance from any point in the 3D scene to the object, the AI ​​system comprising at least one processor and a memory storing instructions that cause the at least one processor of the AI ​​system to: (i) generate an implicit representation of a 3D scene including a 3D object by training a neural network including an encoder configured to encode data indicative of a 3D point cloud of a shape of the object into lattice-based features capturing multiple resolutions of the object, and a decoder configured to decode the lattice-based features into a distance from any point in the 3D scene to the object; The instructions further cause the at least one processor of the AI ​​system to (ii) train the encoder and the decoder using both the positions of the points and the orientations of the points to generate an implicit representation of the 3D object, and (iii) transmit the implicit representation of the 3D object including the encoder and the decoder over a wired or wireless communication channel.

[0006] This input data represents a cloud of oriented points of a 3D scene including 3D objects and can be obtained in several ways. According to one embodiment, it is obtained from an RGB-D sensor that outputs a point cloud. The orientation of the 3D points can be obtained in several ways, for example by locally approximating nearby points by a plane or using a neural network. The same kind of 3D point cloud can be obtained from other types of sensors such as stereo cameras or monocular cameras with neural networks that predict the depth per pixel, among other possibilities.

[0007] Another way of obtaining input data is related to augmented reality / virtual reality. According to this embodiment, there is a triangular mesh that defines the object, and the 3D points can be obtained directly from the corners of the triangular mesh. The orientation of each 3D point is given by the average of the normals of the triangles in the vicinity of each 3D point.

[0008] According to one embodiment, the encoder is trained to convert one or a combination of the position of the point and the orientation of the point in the 3D point cloud into the lattice-based features capturing multiple resolutions of the object, and the decoder is trained on an interpolation of the lattice-based features to reduce a loss function of the error between the distance from the point in the 3D scene to the object generated by the decoder and the ground truth distance.

[0009] According to one embodiment, the encoder is trained to convert the positions of the points in the 3D point cloud to the lattice-based feature, and the decoder is trained on the interpolation within a set of nested shapes that surround the lattice-based feature and are oriented based on the orientation of the points near the corresponding nested shapes.

[0010] According to one embodiment, the encoder is trained to encode the location of the point as the lattice-based feature that captures multiple resolutions of the object, and the decoder is trained based on oriented features represented by interpolation within a set of nested shapes that surround the lattice-based feature and are oriented based on the orientation of the point near a corresponding nested shape.

[0011] According to an embodiment, to train the encoder and the decoder, the processor is configured to use the encoder to encode the input data into an octree representation of features capturing multiple resolutions of the shape of the 3D object, and to surround each feature of the octree representation with an oriented shape having rotational symmetry about an axis, the dimension of the oriented shape surrounding a feature being governed by the level of the surrounded feature on the octree representation and the orientation of the axis of the oriented shape being governed by a normal to the surface of a subset of points in the vicinity of the coordinates of the surrounded feature, the processor is further configured to interpolate features within each oriented shape using volumetric interpolation to update the features of the octree representation, and is configured to decode the updated octree representation of the features using the decoder to generate the distance function, and is configured to update parameters of the neural network to minimize a loss function of the error between a distance from a point in the 3D scene to the object generated by the decoder and a ground truth distance.

[0012] According to one embodiment, the oriented shapes having rotational symmetry include one or more of a cylinder and a sphere.

[0013] According to one embodiment, each of the oriented shapes having rotational symmetry is a cylinder that encloses one or more lattice-based features and is oriented such that the axis of each of the cylinders is aligned with a normal to a region of the surface governed by the dimensions of the cylinder and the positions of the enclosed features.

[0014] According to one embodiment, each of the oriented shapes having rotational symmetry is a cylinder that encloses one or more lattice-based features, and the processor is configured to orient the cylinder to align an axis of the cylinder with a normal to a region of the surface governed by the dimensions of the cylinder and the positions of the enclosed features.

[0015] According to one embodiment, the interpolation is volumetric and the processor is configured to determine cylindrical interpolation coefficients measuring the proximity of the point to an extremity of the cylindrical representation, the cylindrical interpolation coefficients comprising (i) a first coefficient calculated from the difference in volume between the distance of the point to the top surface of the cylinder and the distance of the point to the cylinder and the axis of symmetry of the cylinder, (ii) a second coefficient calculated from the difference in volume between the distance of the point to the bottom surface of the cylinder and the distance of the point to the cylinder and the axis of symmetry of the cylinder, and (iii) a third coefficient calculated from the remainder of the cylinder.

[0016] According to one embodiment, during training of the encoder of the neural network, the processor is configured to determine the cylindrical interpolation coefficients for a number of sampled points of an input point cloud.

[0017] According to one embodiment, the processor is configured to render an image of the 3D object on a display device using the implicit representation of the 3D object.

[0018] According to another aspect, an image processing system operatively connected to the AI ​​system via the wired or wireless communication channel, the image processing system configured to render an image of the 3D object on a display device using the implicit representation of the 3D object.

[0019] According to one embodiment, the image of the 3D object is rendered for varying viewing angles.

[0020] According to one embodiment, the image of the 3D object is rendered for varying viewing angles within a virtual reality or gaming application.

[0021] According to another aspect, a robotic system operatively connected to the AI ​​system via the wired or wireless communication channel, the robotic system configured to perform a task using the implicit representation of the 3D object.

[0022] According to another aspect, a display device operatively connected to the AI ​​system via the wired or wireless communication channel, wherein the processor is configured to render an image of the 3D object using the implicit representation of the 3D object, and the display device is configured to display the rendered image of the 3D object.

[0023] According to another aspect, an image processing system configured to render an image of a three-dimensional (3D) object on a display using an implicit representation of the 3D object, the image processing system comprising: a trained neural network, the trained neural network including an encoder configured to encode data indicative of a 3D point cloud of a shape of the 3D object into lattice-based features capturing multiple resolutions of the object, and a decoder configured to decode the lattice-based features into a distance to the object from any point in a 3D scene containing the 3D object, the image processing system further comprising at least one processor and a decoder configured to store instructions. and a memory configured to receive input data indicating a oriented point cloud of a 3D scene including a 3D object, the input data indicating a 3D position of a point of the 3D point cloud and an orientation of the point defining a normal to a surface of the 3D object at a position of the point proximate to the 3D position, the instructions further causing the at least one processor to generate, at the encoder, an implicit representation of the 3D object using both the position of the point and the orientation of the point, render, at the encoder, an image of the 3D object using the implicit representation of the 3D object, and display the rendered image on a display.

[0024] These and other aspects of the various embodiments will be apparent from and elucidated with reference to the embodiments described hereinafter.

[0025] In the drawings, like reference characters generally refer to the same parts throughout the different views. The figures illustrating features and aspects of implementing various embodiments should not be construed as limiting other possible embodiments falling within the scope of the claims. Also, the drawings are not necessarily to scale, with emphasis instead generally being placed upon illustrating the principles of various embodiments. [Brief description of the drawings]

[0026] [Figure 1] 1 is a schematic diagram of a 3D object representation generation method and system according to one embodiment;

[0027] [Figure 2A] FIG. 2 is a flowchart diagram of a method for training a 3D object representation generation system according to one embodiment.

[0028] [Figure 2B] 2 is a flowchart diagram of a method for generating an implicit function representation of a three-dimensional scene and decoding the representation, according to one embodiment.

[0029] [Figure 3A] FIG. 1 is a schematic diagram of a process for determining orientation when generating a 3D object representation, according to one embodiment.

[0030] [Figure 3B] FIG. 1 is a schematic diagram of a process for determining orientation when generating a 3D object representation, according to one embodiment.

[0031] [Figure 3C] FIG. 1 is a schematic diagram of a process for determining orientation when generating a 3D object representation, according to one embodiment.

[0032] [Figure 4A] FIG. 1 is a schematic diagram of a multi-resolution oriented lattice that uses a structured lattice along with object normal directions to extend an octree representation, according to one embodiment.

[0033] [Figure 4B] FIG. 1 is a schematic diagram of generating an oriented point cloud of a 3D object including point orientations according to one embodiment.

[0034] [Diagram 5] 2 is a flowchart diagram of a method for generating an implicit function representation of a three-dimensional scene according to one embodiment.

[0035] [Figure 6] FIG. 1 illustrates a proposed scheme for cylindrical interpolation, according to one embodiment.

[0036] [Figure 7] FIG. 1 shows a table of ablation study results, according to one embodiment.

[0037] [Figure 8] 1 is a schematic diagram of the results of an ablation study, according to one embodiment.

[0038] [Figure 9] FIG. 1 illustrates a table comparing structured grids versus directed grids, according to one embodiment.

[0039] [Figure 10] FIG. 1 is a schematic diagram of large-scale scene rendering according to one embodiment.

[0040] [Figure 11] FIG. 13 shows a table detailing experimental results of the method for other types of representations, according to one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0041] This disclosure describes various embodiments of systems and methods configured to generate an implicit representation of a three-dimensional (3D) scene. More generally, applicants have recognized and appreciated that it would be beneficial to provide a method and system that considers all of the underlying geometric features of an object when generating an implicit surface representation. Accordingly, one implicit representation generation system generates an implicit representation of a 3D scene including a 3D object by training a neural network. The neural network includes an encoder configured to encode data indicative of a 3D point cloud of the shape of the 3D object into lattice-based features that capture multiple resolutions of the 3D object. The neural network also includes a decoder configured to decode the lattice-based features into a distance from any point in the 3D scene to the object.

[0042] In a 3D implicit representation pipeline, a single 3D point is passed through a geometric encoder, a position encoder, or both. The features are then injected into a decoder that models the surface of the object. By repeating the process for all point cloud points, a sparse output representation is obtained for the modeled 3D surface.

[0043] Referring to FIG. 1, in one embodiment, a schematic diagram of a 3D object representation generation method and system is shown. It should be understood that the method described in connection with the figure is provided by way of example only and does not limit the scope of the present disclosure. The implicit function representation generation system 200 can be any of the systems described or otherwise contemplated herein. The implicit function representation generation system can be a single system or multiple different systems.

[0044] The 3D object representation generation system 200 is an artificial intelligence (AI) system for generating an implicit representation of a 3D scene that includes a 3D object. The system 200 trains a neural network 232 that includes (1) an encoder 233 configured to encode data indicative of a 3D point cloud 120 of the shape of the object 110 into lattice-based features that capture multiple resolutions of the object, and (2) a decoder 234 configured to decode the lattice-based features into a distance from any point in the 3D scene to the object.

[0045] The system 200 receives input data indicative of an oriented point cloud 120 of a 3D scene including a 3D object 110, the input data indicating 3D positions of points of the 3D point cloud and point orientations 130 that define normals to the surface of the 3D object at positions proximate the points' 3D positions. The system 200 uses both the point positions and the point orientations to train an encoder 233 and a decoder 234 to generate an implicit representation of the 3D object. The system can then transmit the implicit representation 140 of the 3D object (optionally including an encoder and decoder), such as via a communication interface 250, over a wired or wireless communication channel.

[0046] According to an embodiment, an implicit function representation generation system 200 is provided. Referring to an embodiment of the implicit function representation generation system 200 as depicted in FIG. 1, for example, the system comprises one or more of a processor 220, a memory 230, a user interface 240, and a communication interface 250, which are interconnected via one or more system buses 260. It will be understood that FIG. 1 constitutes an abstraction in some respects, and the actual organization of the components of the system 200 may be different and more complex than that shown. In addition, the implicit function representation generation system 200 can be any of the systems described or otherwise contemplated herein. Other elements and components of the implicit function representation generation system 200 are disclosed and / or contemplated elsewhere herein.

[0047] According to one embodiment, system 200 includes a processor 220 that can execute instructions stored in memory 230 or otherwise process data to, for example, perform one or more steps of the method. Processor 220 may be formed from one or more modules. Processor 220 may take any suitable form including, but not limited to, a microprocessor, a microcontroller, multiple microcontrollers, circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a single processor, or multiple processors.

[0048] The memory 230 may take any suitable form, including non-volatile memory and / or RAM. The memory 230 may include various memories, such as, for example, L1, L2 or L3 caches or system memory. Thus, the memory 230 may include static random access memory (SRAM), dynamic RAM (DRAM), flash memory, read only memory (ROM), or other similar memory devices. The memory may store, among other things, an operating system. The RAM is used by the processor for temporary storage of data. According to one embodiment, the operating system may include code that, when executed by the processor, controls the operation of one or more components of the system 200. It will be apparent that in embodiments in which the processor implements one or more of the functions described herein in hardware, software described as corresponding to such functions in other embodiments may be omitted. The memory 230 may be considered a non-transitory machine-readable medium. As used herein, the term non-transitory is understood to exclude transient signals, but to include all forms of storage, including both volatile and non-volatile memory. According to one embodiment, the memory 230 of the system 200 can store one or more algorithms, modules, and / or instructions to perform one or more functions or steps of a method described or otherwise contemplated herein.

[0049] User interface 240 may include one or more devices for enabling communication with a user. The user interface may be any device or system that allows information to be communicated and / or received and may include a display, a mouse, and / or a keyboard for receiving user commands. In some embodiments, user interface 240 may include a command line interface or a graphical user interface that may be presented to a remote terminal via communication interface 250. The user interface may be located with one or more other components of the system or may be located remotely from the system and communicate via a wired and / or wireless communication network.

[0050] Communications interface 250 may include one or more devices for enabling communication with other hardware devices. For example, communications interface 250 may include a network interface card (NIC) configured to communicate according to an Ethernet protocol. In addition, communications interface 250 may implement a TCP / IP stack for communication according to a TCP / IP protocol. Various alternative or additional hardware or configurations for communications interface 250 will be apparent.

[0051] Although system 200 is shown as including one of each of the described components, in various embodiments, various components may be duplicated. For example, processor 220 may include multiple microprocessors configured to independently perform the methods described herein or to perform steps or subroutines of the methods described herein, such that the multiple processors cooperate to achieve the functions described herein. Furthermore, when one or more components of system 200 are implemented in a cloud computing system, various hardware components may belong to separate physical systems. For example, processor 220 may include a first processor in a first server and a second processor in a second server. Many other variations and configurations are possible.

[0052] 2A, in one embodiment, a flowchart of a method 250 for training a 3D object representation generation system 200 to generate an implicit representation of a 3D scene that includes a 3D object 110. During training 260, as described herein, the system trains a neural network including an encoder 233 and a decoder 234 to encode data indicative of a 3D point cloud of the object's shape using an octree representation into lattice-based features that capture multiple resolutions of the object, and to decode the lattice-based features into distances from any point in the 3D scene to the object, thereby modeling the surface of the object.

[0053] 2B is a flowchart of a method for generating an implicit representation of a three-dimensional scene and decoding the representation, according to one embodiment. The neural network 232 of the system includes an encoder 233 that is trained to encode data indicative of a 3D point cloud of the shape of an object 270 into lattice-based features that capture multiple resolutions of the object. The neural network also includes a decoder 234 that is trained to decode the lattice-based features into a distance from any point in the 3D scene to the object, thereby modeling the surface of the object via the implicit representation 140 of the 3D object. This implicit representation of the 3D object can be used immediately, communicated, and / or stored for future use. During execution 280, the trained neural network is used to decode the encoded implicit representation 140 of the 3D object using the trained decoder 234.

[0054] 3A-3C, in one embodiment, a schematic diagram of a process for determining orientation when generating a 3D object representation is shown, according to one embodiment. To generate training data and encode the 3D object using a trained encoder, the system generates input data indicating an oriented point cloud of a 3D scene including the 3D object. The input data indicates 3D positions of points of the 3D point cloud and orientations of the points that define normals to the surface of the 3D object at positions proximate to the 3D positions of the points.

[0055] For example, referring to FIG. 3A, geometric encoders perform better when different levels of detail are explicitly modeled. The 3D data points are arranged at different spatial resolutions, as shown in grid cells at 310. For each resolution, there are different cells with different sizes. These different grid resolutions are used to aggregate all the data points that are within the respective cells. The numbers 330 and 340 indicate the two different resolutions.

[0056] According to one embodiment, the system utilizes five different grid resolutions to create a multi-resolution grid encoder. However, the number of different grid resolutions may be more or less than five. For example, the number of different grid resolutions may depend on the roughness of the object / surface and can therefore be adjusted accordingly. According to one embodiment, a single decoder can be used for different multi-resolution grid encoders.

[0057] 3B and 3C, tree 380 shows how the orientation is calculated for the toy example for a single rotation case. In reality, there are three rotations, increasing the tree branching factor to six. As one goes deeper into the depth of the tree, the orientation will get closer to the correct surface direction, illustrated as different greyscales at 380 (FIG. 3B) and 340 (FIG. 3C). Each level of depth in the tree sets the orientation for a single resolution grid.

[0058] 3A-3C, for example, cells 340 and 330 are rotated according to the orientation (i.e., their respective grayscale / depth in the tree) calculated in 380 to obtain rotated cells 360 and 370 (FIG. 3C), respectively. This procedure is repeated for each and every cell, including at different resolutions.

[0059] According to an embodiment, to generate input 3D data for training, 3D positions of points are projectively obtained across the entire scene by first sampling closer to the object surface and then adding more dispersion around the environment.

[0060] 4A, a schematic diagram of a process for generating a multi-resolution oriented grid that uses a structured grid along with object normal directions to extend an octree representation is shown. Cells 410 are rotated into a oriented tree and respective level of depth (LOD) levels to form an oriented grid 420.

[0061] Referring to FIG. 4B, a schematic diagram of encoding an object 430 is shown. During training, for example, the features of sampled points in a rotated cell (in the point cloud 420) are interpolated according to a cylindrical interpolation scheme shown at 422, where neighboring cell features are aggregated with a 3DCNN sparse kernel. These features can be used in decoders of current state-of-the-art for object representations such as signed distance functions (SDFs) and occupancies. It is important to note that although a cylindrical interpolation scheme is described with respect to FIG. 4B, any shape with rotational symmetry around an axis is suitable. For example, it can be a sphere or other volumetric shape with rotational symmetry around an axis.

[0062] Referring to Figure 5, in one embodiment, a method 500 for encoding an object using a trained neural network encoder is shown. The encoder includes a directed lattice geometric encoder 520 that encodes the orientation of a point in a point cloud received as input 510, and a point encoder 530 that encodes the point cloud data received as input 510. Each 3D point passes through the geometric encoder 520, the position encoder 530, or both. According to one embodiment, the input 510 is a point and a pre-initialized tree that best fits the object. The output 540 is a set of level of depth (LOD) tree features that form a 3D representation.

[0063] The point encoder 530 includes both the position encoder and the anchor normals. The directed lattice geometric encoder 520 starts with directed lattice construction at 550, which is a pre-computed step. This includes multiple resolutions of the directed lattice. For example, FIG. 5 shows three resolutions 552, but fewer and more resolutions are possible. For each resolution, directed lattice features 554 (extracted from the tree) are generated. These features are locally aggregated across neighborhood information (i.e., local feature aggregation 560). The aggregated features 562 are used in a cylindrical interpolation scheme 570 described or otherwise envisioned herein to generate the final LOD features 580. The LOD tree features 580 and the encoded points from the point encoder 530 result in the 3D representation 540. Note that while the method of FIG. 5 utilizes a cylindrical interpolation scheme 570, any shape with rotational symmetry around an axis is suitable. For example, it could be a sphere or other volumetric shape with rotational symmetry around an axis.

[0064] According to an embodiment, a directed lattice is constructed. An octree representation is used to model the 3D representation. Specifically, the system uses an octree representation to model a lattice-based 3D encoder. However, in addition to the standard eight actions to split a lattice into eight smaller lattices at subsequent depth levels, the system includes a rotation action to model the orientation of cells, and at higher levels, smaller (denser) lattices and finer alignments better represent objects. Instead of modeling each action individually, which would result in a branching factor of 56 (8 for lattice position × 7 for orientation) for each subsequent LOD, and because lattice size and orientation are independent, they are divided into two trees: (i) Tree 1: a structured octree to model the size of the lattice, and (ii) Tree 2: a direction tree to model the orientation of the cells.

[0065] For the structured octree in Tree 1, the representation consists of position LOD (without orientation) bounded in [-1,1]. A typical octree modeling is traced from existing research.

[0066] For Orientation Tree 2, for a normalized point x taken from the surface point cloud of the object, a normal, denoted as n, is associated with this query. The goal is to align the cells along the surface. To maintain consistency within the LOD, we constructed a set of normal anchors that represent a finite set of possible orientations per level. We then rotated the cell so that its z-axis coincides with the anchor that is closest to the query normal n. To model this search tree, we need to define i) node states, ii) actions, iii) state transitions, and iv) initial state.

[0067]

number

[0068]

number

[0069] To compute the state δ for each cell, a rotational anchor can align the cell's z-axis with the surface normal, up to some rotational degree of freedom, using cosine similarity.

[0070] Regarding the association of a grid with a query point, each cell in Tree 1 has a fixed orientation that is calculated from searching Tree 2. During training and evaluation, a query point is associated with a cell on a particular LOD in Tree 1 (the structured tree). The cell is then rotated using the corresponding rotation anchor. Note that a point may be outside all octree cells; in this case the query is discarded.

[0071] Trilinear interpolation has been the typical method to obtain features for structured (unoriented) grids, as shown in Figure 3A, but for oriented grids, it is not possible to use the same approach as for structured grids. Therefore, the present system uses an oriented cylinder as shown in Figure 4B (which is a 3D representation of the oriented grid in Figure 3C), which can exploit cell alignment and mitigate the lack of invariance in defining the grid orientation (invariant rotations around the normal direction), as discussed herein. This rotational invariance places an explicit smoothness constraint on the points in the grid.

[0072] Therefore, referring to Figure 6, there is a proposed scheme for cylindrical interpolation, according to one embodiment. Note that while the method of Figure 6 utilizes a cylindrical interpolation scheme, any shape with rotational symmetry around an axis is suitable. For example, it could be a sphere or other volumetric shape with rotational symmetry around an axis.

[0073] The input cell grid has a corresponding anchor normal n 610 obtained from the oriented grid (per LOD). A cylinder 620 is aligned with the grid normal anchor 610 with radius R and height H. The interpolation scheme is of volumetric interpolation type. It depends on the distances h1 and h2 of the query point x to the height boundary of the cylinder, and the distance between x and the cylindrical symmetry axis, denoted as r.

[0074] In 630, a first coefficient c0 is calculated from the distance h1 to the top of the point and the difference in volume taking into account R and the distance r to the axis of symmetry of the point. In 640, a coefficient c2 is calculated from the distance h2 to the bottom and the difference in volume taking into account R and the distance r to the axis of symmetry of the point. Finally, in 650, c1 is the remainder of the cylinder. Each coefficient is calculated by the associated learnable feature e for k={0,1,2}. k The interpolated feature f has the following structure: k , c k At 660, cylindrical interpolation features are generated from c0, c1, and c2.

[0075] Given a query point, the goal is to calculate the relative spatial volume for the feature coefficients, taking into account the relative position of the point within the cylinder. The cylindrical interpolation coefficients measure the proximity of the point to the ends of the cylindrical cell representation, as shown in Figure 6. Points closer to the top and borders of the cylinder will generate smaller bounding volumes (volume at 630 in Figure 6). Thus, the distance from its opposite face will be greater, and therefore a larger volume coefficient (volume at 640 in Figure 6). The highest coefficients in this example are opposite along the central axis (volume at 650 in Figure 6).

[0076]

number

[0077] Finally, the interpolated features from the geometric encoder in Figure 6 are a linear interpolation of the coefficients and the local features computed from the 3DCNN:

number

[0078] In our implementation, since our method focuses on a new lattice-based encoder, our system evaluates our method using state-of-the-art decoder architectures and output representations. The loss function and training procedure are also described below.

[0079] The decoder architecture utilizes a multi-layer perceptron. The decoder is trained at each level and shared across all LODs. In addition to the input interpolated features, a state-of-the-art position encoder Φ p (·) is on the point L p are added together with the frequencies, Φ n (·) is the normal to the anchor L n The points and normals are added to each position encoder, and the size is P = 3 × 2 × L p +3 and N=3×2×L n+3. The method is shown as output expressions in terms of SDF and occupancy.

[0080] During the training phase, queries N q are sampled from the input point cloud, it is determined which voxel they are in for each of the LODs, and the features are interpolated according to the selected voxel. The sum of squared errors or cross entropy of the predicted samples is calculated from the active LODs for SDF and occupancy, respectively. In addition, double backpropagation is used to find the normal. Then, the L2 norm is calculated between the calculated normal and the anchor normal as a regularization term. The two terms are added (weighted sum) to find the final loss.

[0081] During the evaluation, we apply uniformly distributed input samples to a matrix of resolution Q = 512. 3 The mesh is obtained from the unit cube of . The results are shown for the last LOD, which corresponds to the finer LOD. If the input query does not match an existing octree cell, it is discarded. Finally, the mesh from the output using marching cubes is obtained.

[0082] Working Example

[0083] The following describes example implementations and analyses using the methods and systems described or otherwise contemplated herein, it will be understood that these are provided as examples only and do not limit the scope of the invention.

[0084] According to an embodiment, the 3D reconstruction quality was evaluated for each object using Chamfer Distance (CD), Normal Consistency (NC), and Intersection over Union (IoU). CD was calculated as the inverse minimum distance between a query point and its ground truth match. CD was calculated five times and the average was presented. NC is the corresponding normal, calculated from the cosine similarity between the query normal (corresponding to the query point obtained during CD calculation) and its corresponding ground truth normal. NC is reported as the residual of the cosine similarity between both normals. IoU quantifies the overlap between the two grid sets. The meshes were 3D meshed with resolution Q=128 for IoU. 3 Rendered using a cube.

[0085] We evaluate our method on three datasets: ABC, Thingi10k, and ShapeNet. A total of 32 meshes were sampled from Thingi10k and ABC, respectively, and 150 meshes were sampled from ShapeNet. The ShapeNet meshes were made watertight, and the work was performed in PyTorch.

[0086] The decoder architecture has one hidden layer of dimension 128 with ReLU. Each voxel feature is represented as an F = 32 dimensional feature vector. The position encoding for the query points and normals is L p =L n = 6 frequencies. Sparse 3D convolutions are considered for local feature aggregation with kernel sizes Kk, Vl, k = 5. The cylinder radius R is empirically set as follows:

number

[0087] We use the Adam optimizer to optimize the model with a learning rate of 0.001 and α n =0.1 and trained for 100 epochs. 6An initial sample size of points is considered, with a batch size of 512. Resampling is performed after every epoch. Points are sampled in equal proportions from the surface and its nearby neighbours. It is also guaranteed that each voxel has at least 32 samples before surface sampling. LODs £'={3,...,7} were considered for all datasets.

[0088] For baselines, the method was compared to state-of-the-art methods BACON, SIREN, and Fourier feature functions (FF) trained on the supplied settings. Structured grid direct methods were also evaluated against the method for a fair comparison to directed grids, which required smaller changes in the pipeline. The surface reconstruction settings were the same for all methods, as described above.

[0089] For the experiments, ablation was utilized as follows: Structured versus directed grids were also compared, and the method was evaluated against methods that use different encoder strategies. These are discussed below.

[0090] Ablation

[0091] Changes were made incrementally to different pipeline blocks to analyze the relevance of each component. 10 meshes from the ABC and Thingi10k datasets are randomly sampled for training and testing. Figure 8 shows the different cases listed in Figure 7 based on the changes made to the encoder. It was noted that the use of oriented lattices with trilinear interpolation results in many holes. Since the cells are rotated per anchor normal, a rotation-invariant cylindrical representation was used for the interpolation. Although still coarser, this results in a more adaptive representation (significant improvement in CD).

[0092] Therefore, with reference to Figure 7, the results of the ablation study are presented. The table shows the different stages leading to the final encoder: starting from a directed lattice with trilinear interpolation, then the proposed cylindrical interpolation, and finally, local feature aggregation using 3DCNN. 5 Multiply by 10- and add NC to it. 4 Multiply by.

[0093] Referring to Fig. 8, the impact of ablation on rendering is shown, mirroring the numerical results in Fig. 7. Panel (a) represents a directional encoder with trilinear interpolation, (b) adds cylindrical interpolation, (c) and (d) use 3x3x3 and 5x5x5 3DCNN kernels for feature aggregation, respectively, (e) adds normal regularization to (d), and (f) shows the ground truth.

[0094] We observe a significant improvement in mesh smoothness (reflected in NC) with the addition of 3DCNN (Figure 8(c) and (d)), which effectively contributes to the local feature aggregation step. Experiments show that a 5 × 5 × 5 kernel achieves better performance and is preferred for subsequent experiments. The proposed normal regularization enhances smoothness but at the expense of accuracy.

[0095] Structured lattice vs. directed lattice

[0096] With reference to Figure 9, a comparison of structured versus directed lattices is shown along with results for the SDF and occupancy decoders. 5 Multiply by 10- and add NC to it. 4The table compares the performance of the method disclosed herein with structured grids on SDF and occupancy decoders. The SDF and occupancy decoders are trained as described above. Normal regularization remains the same for both cases. The method disclosed herein outperforms structured grids on SDF decoders on all fronts and produces smoother results on structured surfaces. Despite the underperformance of the occupancy framework, fewer holes and depressions are observed on the mesh (the latter having a significant impact on IoU). These results demonstrate the adaptability of the method to different decoder output representations.

[0097] To open up a possible extension of the method to large-scale scene representations, a rendering of the method is shown on a scene from Matterport3D, as shown in Figure 10. The scene is divided into 4x4 crops (including the ground) and a model is trained with an occupancy decoder for each crop. During inference, the mesh crops are rendered using marching cubes, as shown in Figure 10, and finally fused to generate the scene. For thin surfaces, structured grids result in a coarser and blurrier 3D representation. The method disclosed herein adapts well to thin surfaces and renders the scene with less coarseness and sharper quality.

[0098] As a result of the oriented lattice, we observe that the proposed encoder renders planes more efficiently with fewer training steps. Especially for more structured regular objects, the structured lattice produces a caustic-like effect (surface noise). We observe a noise reduction in the oriented lattice surface reconstruction immediately from the first epoch.

[0099] Baseline

[0100] The table in Fig. 11 details the experimental results of our method (feature-based geometric multiscale representation) against other types of representations, i.e., lattice-based networks like SIREN and BACON, and multiscale representations like BACON. It also compares our method with NDF, which uses an unsigned output representation to consider non-watertight objects. NDF uses a ball pivoting algorithm to obtain a mesh, which requires a lot of hand tuning. It was computationally expensive and resulted in a very discontinuous mesh representation and poor results. Instead, this step is replaced by voxelizing the point cloud and obtaining the surface using marching cubes, using the same settings as above.

[0101] Thus, we show that the lattice-based method outperforms the baselines with significant improvements on all fronts. For simple datasets consisting of planar objects, such as ABC, the encoder reconstructs a smoother plane due to the alignment of the oriented lattice. Although rendering holistic details, most baselines often have overly smoothed surfaces. We also observe a higher IoU for our method, as a result of fewer holes in the mesh and negligible splatting (many small mesh traces around the sampling region). Overall, the oriented lattice produces a more robust 3D representation with higher fidelity across all datasets.

[0102] The number of parameters required to render the mesh is also provided. The advantage of the multi-resolution lattice representation is that the size of the decoder can be reduced to only an MLP with one hidden layer. As a result, this method obtains a mesh faster than other methods.

[0103] We analyzed examples from the Thingi10k and ShapeNet datasets. BACON and FF are able to model objects with reasonable accuracy, but record a lot of splatter, resulting in undesirable noisy surfaces and artifacts. SIREN and BACON produce overly smoothed surfaces and lose intricate details on the mesh. NDF produces a lot of holes, but manages to obtain a compact mesh without splatter. BACON, SIREN, and FF collapse on the ShapeNet dataset. We tried both watertight and non-watertight versions of ShapeNet, with similar results for the failed baselines.

[0104] Small holes can arise from non-tight planes that affect both directed and structured grids. However, directed grids fill holes better than their structured counterparts. A more general problem with multi-resolution grid representations is the difficulty of modeling thin surfaces. Despite this limitation, the present method is a substantial improvement over structured grids.

[0105] Thus, the examples and experiments demonstrate that this novel approach for a 3D lattice-based encoder for 3D representations delivers state-of-the-art results while being more robust and accurate to decoder representation changes. The encoder takes into account inherent structural regularities in the object by aligning the lattice with the object surface normals and aggregating cell features from a newly developed cylindrical interpolation technique and local aggregation scheme that alleviates problems caused by the alignment.

[0106] In addition, examples and experiments demonstrate that the systems and methods disclosed or otherwise envisioned herein result in improved computer systems. The 3D object representation generation system can generate 3D object representations faster and better than prior art systems. Thus, the 3D object representation generation system described herein, including the trained neural network, is an improvement over prior art 3D object representation generation systems.

[0107] According to an embodiment, the systems and methods disclosed or otherwise contemplated herein are configured to process thousands or millions of data points during training of the neural network and during execution using the trained neural network. For example, generating a functional, skilled, trained neural network from a corpus of training data requires processing millions of data points from the input data and generated features. This may require millions or billions of calculations to generate a new trained neural network. As a result, each trained neural network is new and different based on the input data and algorithm parameters, thus improving the functionality of the system. Generating a functional, skilled, trained neural network involves a process with an amount of calculation and analysis that the human brain cannot achieve in a lifetime.

[0108] All definitions and those used herein should be understood to govern any dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.

[0109] The indefinite articles "a" and "an," as used in the specification and claims, unless clearly indicated to the contrary, should be understood to mean "at least one."

[0110] The term "and / or" as used in the specification and claims should be understood to mean "either or both" of the elements so conjoined, i.e., elements that are conjunctive in some cases and disjunctive in other cases. Multiple elements listed with "and / or" should be construed in the same manner, i.e., "one or more" of the elements so conjoined. Other elements other than the elements specifically identified by the "and / or" clause may be present as alternatives, whether related or unrelated to those elements specifically identified.

[0111] As used herein and in the claims, "or" should be understood to have the same meaning as "and / or" as defined above. For example, when separating items in a list, "or" or "and / or" should be interpreted as inclusive, i.e., the inclusion of at least one, but more than one, of some element or list of elements, and, as an option, additional unlisted items. Only language clearly indicated to the contrary, such as "only one of" or "exactly one of," or, when used in the claims, "consisting of," will refer to the inclusion of exactly one element of some element or list of elements. In general, the word "or" used herein should only be interpreted as indicating exclusive alternatives (i.e., "one or the other, but not both") when preceded by an exclusivity language, such as "either," "one of," "only one of," or "exactly one of."

[0112] As used herein and in the claims, the phrase "at least one" should be understood in reference to a list of one or more elements to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed in the list of elements, and not excluding any combination of elements in the list of elements. This definition also allows for elements other than the elements specifically identified in the list of elements to which the phrase "at least one" refers, may be present as an option, whether related or unrelated to the elements specifically identified.

[0113] It is also to be understood that, unless expressly indicated to the contrary, in any method claimed herein that includes multiple steps or acts, the order of the method steps or acts is not necessarily limited to the order in which the method steps or acts are described.

[0114] In the claims as well as in the above specification, all transitional phrases such as "comprising," "including," "having," "having," "containing," "involving," "holding," "consisting of," and the like, are to be understood to be open-ended, i.e., meaning including but not limited to. Only the transitional phrases "consisting of" and "consisting essentially of" shall be closed or semi-closed transitional phrases, respectively.

[0115] Although several inventive embodiments have been described and illustrated herein, those skilled in the art will readily envision various other means and / or structures for performing the functions and / or obtaining one or more of the results and / or advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the inventive embodiments described herein. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are meant to be exemplary, and that the actual parameters, dimensions, materials, and / or configurations will depend on the particular application for which the teachings of the present invention are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific inventive embodiments described herein. Thus, the foregoing embodiments are presented by way of example only, and it will be understood that within the scope of the claims and their equivalents, the inventive embodiments may be practiced otherwise than as specifically described and claimed. The inventive embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the inventive scope of the present disclosure, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.

Claims

1. 1. An artificial intelligence (AI) system for generating an implicit representation of a three-dimensional (3D) scene including a 3D object by training a neural network, the neural network including: an encoder configured to encode data indicative of a 3D point cloud of a shape of the object into lattice-based features capturing multiple resolutions of the object; and a decoder configured to decode the lattice-based features into a distance from any point in the 3D scene to the object, the AI ​​system comprising: at least one processor and a memory storing instructions, the instructions causing the at least one processor of the AI ​​system to: receiving input data indicative of an oriented point cloud of a 3D scene including a 3D object, the input data indicating 3D positions of points of the 3D point cloud and orientations of the points defining normals to a surface of the 3D object at positions proximate to the 3D positions of the points, the instructions further including causing the at least one processor of the AI ​​system to: training the encoder and the decoder using both the positions of the points and the orientations of the points to generate an implicit representation of the 3D object; The AI ​​system causes the implicit representation of the 3D object, including the encoder and decoder, to be transmitted over a wired or wireless communication channel.

2. 2. The AI ​​system of claim 1, wherein the encoder is trained to convert one or a combination of the positions of the points and the orientations of the points of the 3D point cloud into the lattice-based features capturing multiple resolutions of the object, and the decoder is trained on an interpolation of the lattice-based features to reduce a loss function of an error between a distance from a point in the 3D scene to the object generated by the decoder and a ground truth distance.

3. 3. The AI ​​system of claim 2, wherein the encoder is trained to convert positions of the points of the 3D point cloud into the lattice-based features, and the decoder is trained on the interpolation within a set of nested shapes that surround the lattice-based features and are oriented based on the orientation of the points near corresponding nested shapes.

4. 2. The AI ​​system of claim 1, wherein the encoder is trained to encode the positions of the points as the lattice-based features that capture multiple resolutions of the object, and the decoder is trained based on oriented features represented by interpolation within a set of nested shapes that surround the lattice-based features and are oriented based on the orientation of the points in the vicinity of corresponding nested shapes.

5. To train the encoder and the decoder, the processor configured to use the encoder to encode the input data into an octree representation of features capturing multiple resolutions of a shape of the 3D object; The method is configured to surround each feature of the octree representation with an oriented shape having rotational symmetry about an axis, the dimension of the oriented shape surrounding a feature being governed by the level of the enclosed feature on the octree representation and the orientation of the axis of the oriented shape being governed by a normal to a surface of a subset of points in a neighborhood of coordinates of the enclosed feature, the processor further comprising: configured to interpolate features within each oriented shape using volumetric interpolation to update the features in the octree representation; configured to decode, using the decoder, the updated octree representation of the features to generate the distance function; 2. The AI ​​system of claim 1, configured to update parameters of the neural network to minimize a loss function of an error between a distance from a point in the 3D scene generated by the decoder to the object and a ground truth distance.

6. The AI ​​system of claim 5 , wherein the oriented shapes having rotational symmetry include one or more of a cylinder and a sphere.

7. 6. The AI ​​system of claim 5, wherein each of the oriented shapes having rotational symmetry is a cylinder that encloses one or more lattice-based features and is oriented such that an axis of each of the cylinders is aligned with a normal to a region of a surface governed by a dimension of the cylinder and the positions of the enclosed features.

8. 8. The AI ​​system of claim 7, wherein each of the oriented shapes having rotational symmetry is a cylinder that encloses one or more lattice-based features, and the processor is configured to orient the cylinder to align an axis of the cylinder with a normal to a region of a surface governed by a dimension of the cylinder and the positions of the enclosed features.

9. 8. The AI ​​system of claim 7, wherein the interpolation is volumetric interpolation, and the processor is configured to determine cylindrical interpolation coefficients that measure the proximity of a point to an extremity of the cylindrical representation.

10. 10. The AI ​​system of claim 9, wherein the cylindrical interpolation coefficients include: (i) a first coefficient calculated from the difference in volume between the distance of the point to the top surface of the cylinder and the distance of the point to the cylinder and the axis of symmetry of the cylinder; (ii) a second coefficient calculated from the difference in volume between the distance of the point to the bottom surface of the cylinder and the distance of the point to the cylinder and the axis of symmetry of the cylinder; and (iii) a third coefficient calculated from the remainder of the cylinder.

11. The system of claim 10 , wherein during training of the encoder of the neural network, the processor is configured to determine the cylindrical interpolation coefficients for a plurality of sampled points of an input point cloud.

12. 2. The AI ​​system of claim 1, wherein the processor is configured to render an image of the 3D object on a display device using the implicit representation of the 3D object.

13. 2. An image processing system operatively connected to the AI ​​system of claim 1 via the wired or wireless communication channel, the image processing system configured to render an image of the 3D object on a display device using the implicit representation of the 3D object.

14. The image processing system of claim 13 , wherein the images of the 3D object are rendered for varying viewing angles.

15. The image processing system of claim 14 , wherein the image of the 3D object is rendered for varying viewing angles in a virtual reality or gaming application.

16. 2. A robotic system operatively connected to the AI ​​system of claim 1 via the wired or wireless communication channel, the robotic system configured to perform a task using the implicit representation of the 3D object.

17. 2. A display device operatively connected to the AI ​​system of claim 1 via the wired or wireless communication channel, wherein the processor is configured to render an image of the 3D object using the implicit representation of the 3D object, and the display device is configured to display the rendered image of the 3D object.

18. 1. An image processing system configured to render an image of a three-dimensional (3D) object on a display using an implicit representation of the 3D object, comprising: a trained neural network, the trained neural network configured to encode data indicative of a 3D point cloud of a shape of the 3D object into lattice-based features capturing multiple resolutions of the object, and a decoder configured to decode the lattice-based features into a distance to the object from any point in a 3D scene containing the 3D object, the image processing system further comprising: at least one processor and a memory storing instructions, the instructions causing the at least one processor to: receiving input data indicating an oriented point cloud of a 3D scene including a 3D object, the input data indicating 3D positions of points of the 3D point cloud and orientations of the points defining normals to a surface of the 3D object at positions proximate to the 3D positions of the points, the instructions further causing the at least one processor to: generating, in the encoder, an implicit representation of the 3D object using both the positions of the points and the orientations of the points; causing the encoder to render an image of the 3D object using the implicit representation of the 3D object; An image processing system that causes the rendered image to be displayed on a display.

19. 20. The image processing system of claim 18, wherein the encoder is trained to convert positions of the points of the 3D point cloud into the lattice-based features, and the decoder is trained on the interpolation within a set of nested shapes that surround the lattice-based features and are oriented based on the orientation of the points near corresponding nested shapes.

20. The image processing system of claim 18 , wherein the images of the 3D object are rendered for varying viewing angles.