Ray-based implicit neural representation for predicting a ray color
The method addresses the inefficiencies and quality issues in conventional ray-based INR approaches by computing point and ray feature vectors and predicting ray colors, resulting in enhanced rendering quality and reduced computational costs.
Patent Information
- Application Number
- PCT/US2023/085150
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-20
- Publication Date
- 2025-06-26
AI Technical Summary
Conventional ray-based implicit neural representation (INR) approaches face challenges in learning scene representation efficiently, often resulting in blurry or noisy renderings for novel rays, and are computationally expensive.
A method for 3D scene reconstruction that computes point feature vectors and weights for 3D query points on a camera ray, generates a ray feature vector, and predicts a ray color using a neural network, thereby improving rendering quality and reducing computational cost.
The proposed solution achieves improved rendering quality with reduced computational resources, enabling efficient learning of ray-based neural representations from sparse observations.
Smart Images

Figure US2023085150_26062025_PF_FP_ABST
Abstract
Description
RAY-BASED IMPLICIT NEURALREPRESENTATION FOR PREDICTING A RAY COLORBACKGROUND
[0001] For three-dimensional (3D) scene representation, some conventional approaches use point-based implicit neural presentation (INR), but point-based INR may have high rendering computation costs because a neural network may compute an interference for each query point of a rendered image. Some conventional approaches use ray -based INR modeling for neural rendering. Ray -based scene representation uses individual rays instead of 3D points for 3D scenes. Modeling a 3D scene from the perspective of ray, instead of a 3D point, may produce a ray-based INR, which may increase the rendering efficiency as compared with point-based INR approaches. Ray-based INR rendering may include transforming rays into a proper representation before inputting the rays into a neural network for color predictions. However, some conventional ray-based INR approaches may have difficulty learning scene representation and may be prone to produce blurry or noisy renderings for novel rays (e.g., rays from a new viewpoint), when learned from sparse observations. Some conventional ray-based scene representations may use computationally expensive components, such as hypemetworks, multiple local scenes, distillation from pretrained neural radiance fields (NeRFs), and / or transformers, etc.SUMMARY
[0002] In some aspects, the techniques described herein relate to a method for three- dimensional (3D) scene reconstruction, the method including: computing a point feature vector and a weight from multi-dimensional feature data for a plurality of 3D query points on a camera ray; generating a ray feature vector based on point feature vectors and weights; and predicting a ray color for the camera ray based on the ray feature vector.
[0003] In some aspects, the techniques described herein relate to a ray-based implicit neural representation (INR) system for three-dimensional (3D) scene reconstruction, the raybased INR system including: at least one processor; and a non-transitory computer-readable medium storing executable instructions that cause the at least one processor to: compute a point feature vector and a weight from multi-dimensional feature data for a plurality of 3D query’ points on a camera ray; generate a ray feature vector based on point feature vectors and iweights; and predict a ray color for the camera ray based on the ray feature vector.
[0004] The proposed solution in particular relates to a method for modelling a 3D scene from the perspective of a camera ray using 2D image data to generate an implicit neural representation of the 3D scene. The method may comprise, for each of a plurality of three- dimensional (3D) query7points on a camera ray, computing a point feature vector and a weight from multi-dimensional feature data, wherein the camera ray is emitted by a ray camera towards a pixel in an image plane of the 2D image data and the camera ray is defined by a line extending from a camera origin of the ray camera though the pixel and into the 3D scene. The method then further comprises generating a ray feature vector based on point feature vectors and weights for the 3D query' points; and predicting a ray color based on the ray feature vector, the ray color representing the color of light observed along a ray cast (including the camera ray) from the camera into the 3D scene. A weight may generally be a value that represents a relative importance of the 3D query' point to the 3D scene. A weight may thus, for example, be in the range of zero to a hundred or more. Multi-dimensional feature data may generally relate to a feature volume including one or more feature maps about the underlying 3D scenes. In some examples, the multi-dimensional feature data includes a multidimensional array (tensor) comprised of a plurality of voxels with feature vectors. The proposed solution may relate to a corresponding ray-based implicit neural representation (INR) system.BRIEF DESCRIPTION OF THE DRAWINGS
[0005] FIG. 1 A illustrates a ray -based implicit neural representation (INR) system according to an aspect.
[0006] FIG. IB illustrates a 3D scene with a camera ray having a plurality of 3D query points according to an aspect.
[0007] FIG. 1C illustrates a ray-based rendering engine according to an aspect.
[0008] FIG. ID illustrates point feature vectors and weights for a plurality' of 3D querypoints according to an aspect.
[0009] FIG. IE illustrates a perspective of extracting a point feature vector for a 3D query' point on a camera ray according to an aspect.
[0010] FIG. IF illustrates a perspective of generating a ray feature vector from a plurality of point feature vectors according to an aspect.
[0011] FIG. 1G illustrates an example of a point embedding function according to an aspect.
[0012] FIG. 2 illustrates a flowchart depicting example operations of predicting a raycolor using a ray-based implicit neural representation (INR) system according to an aspect.DETAILED DESCRIPTION
[0013] This disclosure relates to a ray-based implicit neural representation (INR) system that provides a technical solution that learns a ray -based neural representation of a three-dimensional (3D) scene from two-dimensional (2D) image data, and. during neural rendering, generates a ray color for each camera ray with improved quality (e.g., reducing or eliminating blurring or noisy renderings) in a computationally efficient manner (e.g., reducing computing resources) as compared to conventional approaches. In some examples, the raybased INR system can learn a ray -based neural representation from sparse observations such as a low number of images, partial point cloud(s) and / or incomplete sensor data.
[0014] The ray-based INR system is configured to model a 3D scene from the perspective of a camera ray to achieve an implicit neural representation of the 3D scene using one network inference per camera ray. The ray-based INR system may sample 3D querypoints on a camera ray, and, for each 3D query point, the ray -based INR system may execute a point embedding function to generate a point feature vector and a weight. The weight may be a value that represents a relative importance of the 3D query- point to the 3D scene. The raybased INR system may generate a ray feature vector by aggregating the point feature vectors multiplied by their corresponding weights. The ray-based INR system includes a neural network (e.g., a multilayer perceptron (MLP) network) that predicts a ray color using the rayfeature vector as an input. The ray-based INR system provides technical benefits of improved rendering quality with reduced computational cost. For example, the MLP network is relatively small, and the system may compute one inference per ray, which may increase the rendering speed.
[0015] FIGS. 1 A to ID illustrate a ray -based implicit neural representation (INR) system 100 configured to generate a ray color 136 of a camera ray 114 of a three-dimensional (3D) scene 112. In some examples, the ray-based INR system 100 may be a tensorial radiance field-based model, which models and reconstructs 3D scenes 112 using tensor decompositions. The ray-based INR system 100 samples 3D query points 116 on a camera ray 114, and, for each 3D query- point 116, the ray-based INR system 100 executes a point embedding function 122 to generate a point feature vector 124 and a weight 126. The raybased INR system 100 may generate a ray feature vector 130 by aggregating the point feature vectors 124 multiplied by their corresponding weights 126. The ray -based INR system 100includes a neural network 134 (e.g., a multilayer perceptron (MLP) neural network 134a) that predicts a ray color 136 using the ray feature vector 130 as an input. The ray-based INR system 100 may provide relatively improved rendering quality with reduced computational cost. For example, the MLP neural network 134a is relatively small, and the ray-based INR system 100 may compute one inference per ray, which may increase the rendering speed.
[0016] The ray-based INR system 100 models a 3D scene 112 from the perspective of a camera ray 114 using 2D image data 102 to achieve an implicit neural representation of the 3D scene 1 12 using one network inference per camera ray 114. The ray-based INR system 100 includes a ray-based rendering engine 108 to compute a ray color 136 of a camera ray 114 for a 3D scene 112. As shown in FIG. IB, the ray-based rendering engine 108 may generate a camera ray 114 for rendering a 3D scene 112 and sample (e.g., obtain) 3D query points 116 on the camera ray 114. In some examples, the ray -based rendering engine 108 includes a ray camera 118 that emits camera rays 114 towards each pixel in an image plane. The ray camera 118 may be a virtual camera. Each camera ray 114 may be defined by a line extending from a camera origin through the pixel and into the 3D scene 112.
[0017] Along each camera ray 114, the ray-based rendering engine 108 may select 3D query points 116 according to a sampling strategy. In some examples, the sampling strategy includes uniform sampling in which 3D query' points 116 are evenly distributed along the ray’s length. In some examples, the sampling strategy includes stratified sampling in which 3D query points 116 are selected at selected intervals based on expected scene complexity. In some examples, the sampling strategy includes hierarchical sampling with coarse and / or fine sampling to explore space around the camera ray 114.
[0018] As shown in FIG. IB, the 3D query' points 116 may include a 3D query point 116-1, a 3D query point 116-2, a 3D query’ point 116-3. a 3D query point 116-4, and a 3D query point 116-5, and a 3D query point 116-5. The number of 3D query points 116 may be dependent on the scene’s complexity', which may range from several points to hundreds or thousands.
[0019] As shown in FIG. 1C, the ray-based rendering engine 108 includes a feature retrieval engine 120. The feature retrieval engine 120 executes a point embedding function 122 for each 3D query point 116 sampled along the camera ray 114, which generates a point feature vector 124 and a weight 126 for a respective 3D query' point 116. In some examples, the point embedding function 122 uses a location of a 3D query point 116 as an input and generates the point feature vector 124 and the corresponding weight 126 as an output. A point feature vector 124 includes information (e.g., a series of values) about a texture and / or color ofa 3D query point 116. A point feature vector 124 is a feature vector. In some examples, a point feature vector 124 is a feature embedding. A point feature vector 124 includes other features such as geometric features (e.g., position, normal, tangent, depth), other appearance features (e.g., material properties, texture coordinates, emissive properties, etc.), and / or additional features (e.g., visibility7flags, motion information, and / or auxiliary data, etc.).
[0020] As shown in FIGS. ID and IE. the feature retrieval engine 120 generates a point feature vector 124-1 and a weight 126-1 for 3D query point 116-1, a point feature vector 124-2 and a weight 126-2 for 3D query point 116-2, a point feature vector 124-2 and a weight 126-2 for 3D query point 116-2, a point feature vector 124-3 and a w eight 126-3 for 3D query point 116-3, a point feature vector 124-4 and a weight 126-4 for 3D query7point 116-4. a point feature vector 124-5 and a weight 126-5 for 3D query point 116-5, and a point feature vector 124-6 and a weight 126-6 for 3D query point 116-6.
[0021] FIG. 1G illustrates an example of a point embedding function 122. A point embedding function 122 may map each 3D query point 116 in a 3D scene 112 to ahigh- dimensional vector space (e.g., multi-dimensional feature data 138). The point embedding function 122 uses multi-dimensional feature data 138 to compute the point feature vector 124 and the weight 126. In some examples, a weight 126 is a value assigned by a convolutional kernel. A weight 126 may represent a relative importance of the observed point. For example, if one of the sample points along a camera ray 114 is the actual observed point, the 3D query point 1 16 has a weight value higher than a non-observed point. A point feature vector 124 may include a sequence of values that represent the features of a particular 3D query7point 116 on the camera ray 114. The point feature vector 124 may be computed by the point embedding function 122 using the multi-dimensional feature data 138.
[0022] The multi-dimensional feature data 138 includes the learnable parameters generated by the ray-based INR system 100 during a taining phase using 2D image data 102. The multi-dimensional feature data 138 may include any type of multi-dimensional representation of a 3D scene 112. The multi-dimensional feature data 138 may include extracted features such as edges, textures, or specific object parts, and these features are represented as values in the array, with each dimension corresponding to a different aspect of the data. The multi-dimensional feature data 138 may have a depth that represents different slices taken from the 3D data and the number of slices depends on the resolution and designed level of detail. The multi-dimensional feature data 138 may have a height and width, and the height and width may represent the spatial dimensions within each slice. The multidimensional feature data 138 may have a number of channels, and the number of channelsmay represent the different features extracted from the data. The multi-dimensional feature data 138 includes values within the array, which may represent the strength or presence of the features at each point in the 3D data (e.g., higher values might indicate a stronger edge, a more prominent texture, or a higher probability of a specific obj ect being present) and the relationships between the values in different channels and across different dimensions can be used to understand the relationships between different features and their spatial distribution within the data.
[0023] The multi-dimensional feature data 138 may be learned by the ray-based INR system 100 from relatively sparse image data (e.g., the 2D image data 102). In some examples, the multi-dimensional feature data 138 may be represented as a feature volume. In some examples, the multi-dimensional feature data 138 may be represented as one or more feature maps (referred to as feature basis maps 140) about the underlying 3D scene 112. In some examples, the multi-dimensional feature data 138 includes a multi-dimensional array (tensor). In some examples, the multi-dimensional feature data 138 includes a plurality' of voxels with feature vectors.
[0024] In some examples, the multi-dimensional feature data 138 includes one or more feature basis maps 140. A feature basis map 140 may be a collection of feature-based characteristics, which are determined by a convolutional neural network, about the underlying 3D scene 112 such as color, texture, edges, and / or object parts, etc. In some examples, the feature basis maps 140 includes three 3D feature basis maps and three ID feature basis maps. For example, the feature basis maps 140 include a 3D feature basis map 140-1, a 3D feature basis map 140-2, a 3D feature basis map 140-3, a ID feature basis map 142-1, a ID feature basis map 142-2, and a ID feature basis map 142-3.
[0025] In some examples, the ray-based INR system 100 may execute a low-rank tensor decomposition to represent the feature volume as feature basis maps 140, which can reduce the memory requirements for storing the feature volume. For example, the ray-based INR system 100 may use a four-dimensional (4D) tensor to represent the 3D scene 112, which is high-dimensional representation of a radiance field (e.g., TensorRF), and the ray-based INR system 100 may generate a loyver dimensional representation of the 3D scene 112 using a low- rank tensor decomposition, yvhere the lower dimensional representation includes the feature basis maps 140. However, the ray -based INR system 100 may use other types of technology' besides TensorFR to represent the 3D scene 112 such as neural radiance fields (NeRF), plenoptic cameras and light fields, structure from motion (SfM), multi-view stereo (MVS), occupancy mapping and signed distance fields (SDFs).
[0026] The point embedding function 122 may project a 3D query' point 116 on the feature basis maps 140. For example, the point embedding function 122 may project the 3D query’ point 116 at a projected location on the 3D feature basis map 140-1, at a projected location on the 3D basis map 140-2, at a projected location on the 3D basis map 140-3, at a projected location on the ID feature basis map 142-1, at a projected location on the ID feature basis map 142-2. and at a projected location on the ID feature basis map 142-3.
[0027] The point embedding function 122 may exact initial feature vectors (e.g., 144, 154, 146, 156, 148, 152) at projected locations on the feature basis maps 140. For example, the point embedding function 122 may sample an initial feature vector 144 from the 3D feature basis map 140-1 and an initial feature vector 154 from the ID feature basis map 142-2. An initial feature vector 144 may be a sequence of values that represent the features of a respective 3D query point 116 from the 3D feature basis map 140-1 at the projected location. An initial feature vector 144 may be a sequence of values that represent the features of a respective 3D query point 116 from the ID feature basis map 142-2 at the projected location. The point embedding function 122 may sample an initial feature vector 146 from the 3D feature basis map 140-2 and an initial feature vector 156 from the ID feature basis map 142-3. An initial feature vector 146 may be a sequence of values that represent the features of a respective 3D query' point 116 from the 3D feature basis map 140-2 at the projected location. An initial feature vector 156 may be a sequence of values that represent the features of a respective 3D query point 116 from the ID feature basis map 142-3 at the projected location. The point embedding function 122 may sample an initial feature vector 148 from the 3D feature basis map 140-3 and an initial feature vector 152 from the ID feature basis map 142-1. An initial feature vector 148 may be a sequence of values that represent the features of a respective 3D query point 116 from the 3D feature basis map 140-3 at the projected location. An initial feature vector 152 may be a sequence of values that represent the features of a respective 3D query point 116 from the ID feature basis map 142-1 at the projected location.
[0028] The point embedding function 122 may generate immediate feature vectors (e.g., 154, 156. and 158) based on the initial feature vectors sampled from the feature basis maps 140. For example, the point embedding function 122 may generate an immediate feature vector from an initial feature vector of a separate pair of feature basis maps 140, where a pair includes a 3D feature basis map and a ID feature basis map. In some examples, the point embedding function 122 generates an immediate feature vector 154 by multiplying the initial feature vector 144 from the 3D feature basis map 140-1 yvith the initial feature vector 154 from ID feature basis map 142-2. In some examples, the point embedding function 122generates an immediate feature vector 156 by multiplying the initial feature vector 146 from the 3D feature basis map 140-2 with the initial feature vector 156 from ID feature basis map 142-3. In some examples, the point embedding function 122 generates an immediate feature vector 158 by multiplying the initial feature vector 148 from the 3D feature basis map 140-3 with the initial feature vector 152 from ID feature basis map 142-1. The point embedding function 122 may generate a final feature vector 160 based on the immediate feature vectors. In some examples, the point embedding function 122 may generate the final feature vector 160 by concatenating the immediate feature vectors 154, 156, and 158. For example, the final feature vector 160 may include the immediate feature vector 154, the immediate feature vector 156, and the immediate feature vector 158.
[0029] The point embedding function 122 may select a first portion 160-1 of the final feature vector 160 as the point feature vector 124. In some examples, the point embedding function 122 selects a predetermined number of first values (e.g., first x number of channels) as the first portion 160-1 of the final feature vector 160. The first portion 160-1 is a subset of the final feature vector 160. For example, the final feature vector 160 includes a certain number of channels, which may depend on the complexity of the underlying scene. In some examples, the point embedding function 122 selects the first number of channels (e.g., one- hundred and forty- four) of the final feature vector 160 as the point feature vector 124. In some examples, the point embedding function 122 may compute the weight 126 based on a second portion 160-2 of the final feature vector 160. In some examples, the second portion 1 0-2 is the portion of the final feature vector 160 that is not used for the point feature vector 124. In some examples, the second portion includes the remaining channels (e.g., forty-eight) of the final feature vector 160. In some examples, the point embedding function 122 computes the weight 126 by summing the values of the second portion 160-2 of the final feature vector 160. In some examples, the first portion 160-1 is larger than the second portion 160-2.
[0030] Referring to FIGS. 1C and IF, the ray-based rendering engine 108 includes a ray feature generator 128. The ray feature generator 128 generates a ray feature vector 130 based on the point feature vectors 124 and the weights 126 generated by the feature retrieval engine 120. In some examples, the ray feature vector 130 is a ray embedding that can be processed by a neural network. In some examples, the ray feature generator 128 generates the ray feature vector 130 by summing the point feature vectors 124 multiplied by their corresponding weights 126. For example, the ray feature vector 130 includes the sum of the product of the point feature vector 124-1 and the weight 126-1, the product of the point feature vector 124-2 and the weight 126-2, the product of the point feature vector 124-3 and theweight 126-3, the product of the point feature vector 124-4 and the weight 126-4, the product of the point feature vector 124-5 and the weight 126-5, and the product of the point feature vector 124-6 and the weight 126-6.
[0031] The ray-based rendering engine 108 includes a color predictor 132. The color predictor 132 may generate a ray color 136 based on the ray feature vector 130. A ray color 136 represents the color of light that is observed along a specific camera ray 114 that is cast from the ray camera 118 into the 3D scene 112. The color predictor 132 includes a neural network 134 such as a multi-layer perception (MLP) neural network 134a. The MLP neural network 134a may receive the ray feature vector 130 as an input and generate the ray color 136 as an output. A MLP neural network 134a is a type of artificial neural network that may be used for classification and regression tasks. The MLP neural network 134a includes a series of layers of interconnected nodes, with each node computing a weighted sum of its inputs and then applying an activation function to the result. The MLP neural network 134a may receive the ray feature vector 130 and use this information to predict the ray color 136. In some examples, the MLP neural network 134a may receive a positional encoding output of a 3D viewing direction. In some examples, the MLP neural network 134a includes three layers with ReLU activations. In some examples, the MLP neural network 134a includes an input layer, a hidden layer, and a color prediction layer. In some examples, the input layer includes Cfeat + Cdirx A , where Cfeatis the number of ray feature channels, Ciris the positional encoding output of the 3D viewing direction, and A is an integer (e.g., 128 or some other value). In some examples, the hidden layer includes 128x128. In some examples, the color prediction layer includes 128x3.
[0032] The ray-based INR system 100 may include one or more processors 101 and one or more memory devices 103. The processor(s) 101 may be formed in a substrate configured to execute one or more machine executable instructions or pieces of software, firmware, or a combination thereof. The processor(s) 101 can be semiconductor-based - that is, the processors can include semiconductor material that can perform digital logic. The memory device(s) 103 may include any type of storage device that stores information in a format that can be read and / or executed by the processor(s) 101. In some examples, the memory' device(s) 103 is / are a non- transitory computer-readable medium. The memory device(s) 103 may store executable instructions that when executed by the processor(s) 101 may execute the operations discussed with reference to the ray-based INR system 100.
[0033] FIG. 2 illustrates a flowchart 200 depicting example operations of a ray-basedINR system according to an aspect. Although the flowchart 200 of FIG. 2 illustrates the operations in sequential order, it will be appreciated that this is merely an example, and that additional or alternative operations may be included. Further, operations of FIG. 2 and related operations may be executed in a different order than that shown, or in a parallel or overlapping fashion.
[0034] Operation 202 includes, for each of a plurality of three-dimensional (3D) query points on a camera ray, computing a point feature vector and a weight from multi-dimensional feature data. Operation 204 includes generating a ray feature vector based on point feature vectors and weights for the 3D query' points. Operation 206 includes predicting a ray color based on the ray feature vector.
[0035] Clause 1. A method for three-dimensional (3D) scene reconstruction, the method comprising: computing a point feature vector and a weight from multi-dimensional feature data for a plurality of 3D query points on a camera ray; generating a ray feature vector based on point feature vectors and weights; and predicting a ray color for the camera ray based on the ray feature vector.
[0036] Clause 2. The method of clause 1, wherein computing the point feature vector and the weight includes: executing a point embedding function using a location of a 3D query point as an input.
[0037] Clause 3. The method of clause 1 or 2, wherein the multi-dimensional feature data include a plurality of feature basis maps, wherein computing the point feature vector includes: projecting a 3D query point on the plurality of feature basis maps; sampling initial feature vectors at projected locations on the plurality' of feature basis maps; and generating the point feature vector based on the initial feature vectors.
[0038] Clause 4. The method of clause 3, wherein the initial feature vectors include a first feature vector from a first feature basis map and a second feature vector from a second feature basis map, yvherein the generating the point feature vector includes: generating an immediate feature vector by multiplying the first feature vector with the second feature vector; generating a final feature vector by concatenating the immediate feature vector with other immediate feature vectors; and selecting a portion of the final feature vector as the point feature vector.
[0039] Clause 5. The method of clause 4, wherein the first feature basis map is a 3D feature basis map, and the second feature basis map is a one-dimensional feature basis map.
[0040] Clause 6. The method of clause 4, wherein the portion of the final feature vector is a first portion, the method further comprising: generating the weight by summing asecond portion of the final feature vector.
[0041] Clause 7. The method of any one of clauses 1 to 6. wherein generating the ray feature vector includes: aggregating a product of each point feature vector and corresponding weight.
[0042] Clause 8. The method of any one of clauses 1 to 7, wherein predicting the raycolor includes: inputting the ray feature vector to a multi-layer perception (MLP) neural network.
[0043] Clause 9. A computer program product having executable instructions that cause one or more processors to execute any one of clauses 1 to 8.
[0044] Clause 10. A ray-based implicit neural representation (INR) system for three- dimensional (3D) scene reconstruction, the ray-based INR system comprising: at least one processor; and a non-transitory computer-readable medium storing executable instructions that cause the at least one processor to: compute a point feature vector and a weight from multidimensional feature data for a plurality- of 3D query points on a camera ray; generate a rayfeature vector based on point feature vectors and weights; and predict a ray color for the camera ray based on the ray feature vector.
[0045] Clause 11. The ray-based INR system of clause 10, wherein the executable instructions include instructions that cause the at least one processor to: compute the point feature vector and the weight by executing a point embedding function using a location of a 3D query point as an input.
[0046] Clause 12. The ray-based INR system of clause 10 or 11, wherein the multidimensional feature data include a plurality7of feature basis maps, wherein the executable instructions include instructions that cause the at least one processor to: project a 3D query point on the plurality of feature basis maps; sample initial feature vectors at projected locations on the plurality7of feature basis maps; and generate the point feature vector based on the initial feature vectors.
[0047] Clause 13. The ray-based INR system of clause 12, wherein the initial feature vectors include a first feature vector from a first feature basis map and a second feature vector from a second feature basis map, wherein the executable instructions include instructions that cause the at least one processor to: generate an immediate feature vector by multiplying the first feature vector with the second feature vector; generate a final feature vector byconcatenating the immediate feature vector with other immediate feature vectors; and select a portion of the final feature vector as the point feature vector.
[0048] Clause 14. The ray-based INR system of clause 13, wherein the first featurebasis map is a 3D feature basis map, and the second feature basis map is a one-dimensional feature basis map.
[0049] Clause 15. The ray-based INR system of clause 13, wherein the portion of the final feature vector is a first portion, wherein the executable instructions include instructions that cause the at least one processor to: generate the weight by summing a second portion of the final feature vector.
[0050] Clause 16. The ray-based INR system of any one of clauses 10 to 15, wherein the executable instructions include instructions that cause the at least one processor to: aggregate a product of each point feature vector and corresponding weight.
[0051] Clause 17. The ray-based INR system of any one of clauses 10 to 16, wherein the executable instructions include instructions that cause the at least one processor to: input the ray feature vector to a neural network.
[0052] Clause 18. The ray-based INR system of clause 17, wherein the neural network includes a multi-layer perception (MLP) neural network.
[0053] Clause 19. The ray-based INR system of clause 18, wherein the MLP neural network includes an input layer, a hidden layer, and a color prediction layer.
[0054] Clause 20. The ray-based INR system of any one of clauses 10 to 19, wherein the multi-dimensional feature data includes a plurality of 3D feature basis maps and a plurality of one-dimensional (ID) feature basis maps.
[0055] Various implementations of the systems and techniques described here can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device. In addition, the term “module” may include software and / or hardware.
[0056] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms “machine-readable medium” “computer-readable medium” refers to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs))used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0057] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0058] The systems and techniques described here can be implemented in a computing system that includes a back end component (e g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (“LAN”), a wide area network (“WAN”), and the Internet.
[0059] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0060] A number of embodiments have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the specification.
[0061] In addition, the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. In addition, other steps may be provided, or steps may be eliminated, from the described flows, and other components may be added to, or removed from, the described systems. Accordingly, other embodiments are within the scope of the following claims.
Claims
WHAT IS CLAIMED IS:
1. A method for three-dimensional (3D) scene reconstruction, the method comprising: computing a point feature vector and a weight from multi-dimensional feature data for a plurality of 3D query points on a camera ray; generating a ray feature vector based on point feature vectors and weights; and predicting a ray color for the camera ray based on the ray feature vector.
2. The method of claim 1 , wherein computing the point feature vector and the weight includes: executing a point embedding function using a location of a 3D query point as an input.
3. The method of claim 1 or 2, wherein the multi-dimensional feature data include a plurality7of feature basis maps, wherein computing the point feature vector includes: projecting a 3D query point on the plurality of feature basis maps; sampling initial feature vectors at projected locations on the plurality of feature basis maps; and generating the point feature vector based on the initial feature vectors.
4. The method of claim 3, wherein the initial feature vectors include a first feature vector from a first feature basis map and a second feature vector from a second feature basis map, wherein the generating the point feature vector includes: generating an immediate feature vector by multiplying the first feature vector with the second feature vector; generating a final feature vector by concatenating the immediate feature vector with other immediate feature vectors; and selecting a portion of the final feature vector as the point feature vector.
5. The method of claim 4, wherein the first feature basis map is a 3D feature basis map. and the second feature basis map is a one-dimensional feature basis map.
6. The method of claim 4, wherein the portion of the final feature vector is a first portion, the method further comprising: generating the weight by summing a second portion of the final feature vector.
7. The method of any one of claims 1 to 6, wherein generating the ray feature vector includes: aggregating a product of each point feature vector and corresponding weight.
8. The method of any one of claims 1 to 7, wherein predicting the ray color includes: inputting the ray feature vector to a multi-layer perception (MLP) neural network.
9. A computer program product having executable instructions that cause one or more processors to execute any one of claims 1 to 8.
10. A ray -based implicit neural representation (INR) system for three-dimensional (3D) scene reconstruction, the ray -based INR system comprising: at least one processor; and a non-transitory computer-readable medium storing executable instructions that cause the at least one processor to: compute a point feature vector and a weight from multi-dimensional feature data for a plurality of 3D query' points on a camera ray; generate a ray feature vector based on point feature vectors and weights; and predict a ray color for the camera ray based on the ray feature vector.
11. The ray-based INR system of claim 10, wherein the executable instructions include instructions that cause the at least one processor to: compute the point feature vector and the weight by executing a point embedding function using a location of a 3D query point as an input.
12. The ray-based INR system of claim 10 or 11, wherein the multi-dimensional feature data include a plurality of feature basis maps, wherein the executable instructions include instructions that cause the at least one processor to: project a 3D query point on the plurality of feature basis maps; sample initial feature vectors at projected locations on the plurality of feature basis maps; and generate the point feature vector based on the initial feature vectors.
13. The ray-based INR system of claim 12, wherein the initial feature vectors include a first feature vector from a first feature basis map and a second feature vector from a second feature basis map, wherein the executable instructions include instructions that cause the at least one processor to: generate an immediate feature vector by multiplying the first feature vector with the second feature vector; generate a final feature vector by concatenating the immediate feature vector with other immediate feature vectors; and select a portion of the final feature vector as the point feature vector.
14. The ray-based INR system of claim 13, wherein the first feature basis map is a 3D feature basis map, and the second feature basis map is a one-dimensional feature basis map.
15. The ray-based INR system of claim 13, wherein the portion of the final feature vector is a first portion, wherein the executable instructions include instructions that cause the at least one processor to: generate the weight by summing a second portion of the final feature vector.
16. The ray-based INR system of any one of claims 10 to 15, wherein the executable instructions include instructions that cause the at least one processor to: aggregate a product of each point feature vector and corresponding weight.
17. The ray-based INR system of any one of claims 10 to 16, wherein the executable instructions include instructions that cause the at least one processor to: input the ray feature vector to a neural network.
18. The ray-based INR system of claim 17, wherein the neural network includes a multilayer perception (MLP) neural network.
19. The ray-based INR system of claim 18, wherein the MLP neural network includes an input layer, a hidden layer, and a color prediction layer.
20. The ray-based INR system of any one of claims 10 to 19, wherein the multidimensional feature data includes a plurality of 3D feature basis maps and a plurality of onedimensional (ID) feature basis maps.
Citation Information
Cited By
Remote sensing data integrated storage and calculation method and system, computer equipment and medium
CN122262072A