Neural segmentation field for representing three-dimensional scene
By combining a smaller neural network with pre-trained NeRF, the semantic segmentation of three-dimensional scenes is solved by using color texture and volume density, and the problems of rendering speed and resource consumption in the existing technology are solved, achieving efficient semantic segmentation effect.
Patent Information
- Application Number
- CN202380081505.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-27
- Filing Date
- 2023-08-31
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to quickly and efficiently render semantic segmentation diagrams of three-dimensional scenes, especially when computing resources are limited.
Smaller neural network training is combined with pre-trained neural radiation field (NeRF), and semantic segmentation is used to optimize the parameters of the neural network through loss function to achieve rapid rendering.
It realizes the rapid rendering of high-quality three-dimensional scene semantic segmentation graphs with relatively low computing resources, achieving an average cross-convergence ratio (mIoU) above 0.9 and a model size less than 0.4MB.
Smart Images

Figure CN120266156A_ABST
Abstract
Description
[0001] 1. Cross - reference to Related Applications
[0002] This application claims the benefit of priority of U.S. Provisional Application Serial No. 63 / 410,344, filed on September 27, 2022, which is incorporated herein by reference in its entirety. 2. Technical Field
[0003] Various example embodiments relate to rendering a three - dimensional (3D) scene into a two - dimensional (2D) space. 3. Background Art
[0004] The process of associating each pixel in an image with a class label is called semantic segmentation. The labels can be, for example, "person", "flower", "building", etc. Semantic segmentation can be regarded as image classification at the pixel level. Thus, in semantic segmentation, each pixel of an image is typically associated with a corresponding class label. Semantic segmentation is different from object detection in that, unlike object detection, semantic segmentation is performed at the pixel level to determine the relatively accurate contours of the objects within the image.
[0005] Unlike semantic segmentation that treats multiple objects within a class as one entity, instance segmentation identifies individual objects within a class. Instance segmentation is sometimes considered a refined version of semantic segmentation. Various computer vision applications use semantic segmentation or a combination of semantic segmentation and instance segmentation. Summary of the Invention
[0006] Disclosed herein are various embodiments of methods and apparatuses for using machine learning to render a segmentation map of a 3D scene. According to an example embodiment, a smaller neural network (NN) is trained to render a segmentation map corresponding to an arbitrarily selected view of the 3D scene, where the training is performed using a larger neural network that is pre - trained to represent the 3D scene in terms of color texture and volume density. In other words, the small neural network is configured to operate on top of the pre - trained larger neural network, thereby providing the ability to obtain segmentation information more quickly and in a manner that is relatively less burdensome on computational resources. Also disclosed herein are multiple embodiments of an entropy - based loss function and its regularization terms, which are constructed to facilitate the training process for the small neural network, for example, by leveraging auxiliary information that is combined with or available from the previous training of the larger neural network.
[0007] According to an example embodiment, an image processing method is provided. The image processing method includes training a first neural network via a processor to render a segmentation map corresponding to an arbitrarily selected view of a 3D scene. The training includes: using the processor to calculate a color texture and a volume density corresponding to a selected training view of the 3D scene, the calculation being performed using a 3D representation pre-trained to represent the 3D scene; using the processor to generate a predicted segmentation map corresponding to the selected training view of the 3D scene, the generation being performed using the first neural network based on the color texture and the volume density; and using the processor to adjust configuration parameters of NN nodes of the first neural network based on a loss function, the loss function being configured to receive a ground truth segmentation map corresponding to the selected training view as its first input and further being configured to receive the predicted segmentation map as its second input.
[0008] According to another example embodiment, a non-transitory computer-readable medium storing instructions is provided. When executed by the processor, the instructions cause the processor to perform operations including the above or the following image processing method.
[0009] According to yet another example embodiment, an image processing apparatus is provided. The image processing apparatus includes: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to cause the apparatus to at least train a first neural network to render a segmentation map corresponding to an arbitrarily selected view of a 3D scene using the at least one processor; and wherein, to train the first neural network, the apparatus is configured to: calculate a color texture and a volume density corresponding to a selected training view of the 3D scene, the calculation being performed using a 3D representation pre-trained to represent the 3D scene; generate a predicted segmentation map corresponding to the selected training view of the 3D scene, the generation being performed using the first neural network based on the color texture and the volume density; and adjust configuration parameters of NN nodes of the first neural network based on a loss function, the loss function receiving a ground truth segmentation map corresponding to the selected training view as its first input and further receiving the predicted segmentation map as its second input.
[0010] According to yet another example embodiment, there is provided an image processing method, the image processing method including testing a first neural network via a processor to render a segmentation map corresponding to an arbitrarily selected view of a 3D scene, the testing including: calculating, using the processor, a color texture and a volume density corresponding to a selected view of the 3D scene, the calculating being performed using a 3D representation pre-trained to represent the 3D scene; and generating, using the processor, a segmentation map corresponding to the selected view of the 3D scene, the generating being performed using the first neural network based on the color texture and the volume density; and wherein the first neural network has been trained using the processor with the pre-trained 3D representation.
[0011] According to yet another example embodiment, there is provided an image processing apparatus, the image processing apparatus including: at least one processor; and at least one memory including program code; wherein the at least one memory and the program code are configured to cause the apparatus, using the at least one processor, to at least test a first neural network to render a segmentation map corresponding to an arbitrarily selected view of a 3D scene; wherein, to test the first neural network, the apparatus is configured to: calculate a color texture and a volume density corresponding to a selected view of the 3D scene, the calculating being performed using a 3D representation pre-trained to represent the 3D scene; and generate a segmentation map corresponding to the selected view of the 3D scene, the generating being performed using the first neural network based on the color texture and the volume density; and wherein the first neural network has been trained using the pre-trained 3D representation. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Other aspects, features, and benefits of the various disclosed embodiments will become more fully apparent by way of example from the following detailed description and the drawings, in which:
[0013] Figures 1A to 1B is a block diagram illustrating example process flows for training and testing a neural segmentation field according to various embodiments.
[0014] Figure 2 is a block diagram illustrating a multi-layer perceptron (MLP) that can be used to implement a neural segmentation field according to an embodiment.
[0015] Figure 3 is a flowchart illustrating the process of training Figure 2 the MLP according to various embodiments.
[0016] Figure 4 is a flowchart illustrating the process of testing Figure 2 the MLP according to various embodiments.
[0017] Figure 5 is a diagram illustrating, according to some examples, inFigure 3 Flowchart of the calculation of the loss function used in the process.
[0018] Figure 6 Is a flowchart illustrating the calculation of an example regularization term of the loss function used in the process according to some examples. Figure 3 Flowchart of the calculation of an example regularization term of the loss function used in the process.
[0019] Figure 7 Is a flowchart illustrating the calculation of another example regularization term of the loss function used in the process according to some examples. Figure 3 Flowchart of the calculation of another example regularization term of the loss function used in the process.
[0020] Figure 8 Is a block diagram of a computing device according to an embodiment. Detailed implementation
[0021] The present disclosure and its various aspects can be embodied in various forms, including: hardware, devices or circuits controlled by computer-implemented methods, computer program products, computer systems and networks, user interfaces and application programming interfaces; and hardware-implemented methods, signal processing circuits, memory arrays, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc. The foregoing is only intended to give a general idea of the various aspects of the present disclosure and does not limit the scope of the present disclosure in any way.
[0022] In the following description, many details such as optical device configurations, timings, operations, etc. are set forth in order to provide an understanding of one or more aspects of the present disclosure. It will be readily apparent to those skilled in the art that these specific details are merely exemplary and are not intended to limit the scope of the present application.
[0023] In addition, although the present disclosure mainly focuses on examples of using various circuits in digital projection systems, it should be understood that these are merely examples. It should be further understood that the disclosed systems and methods can be used in any device that requires projecting light; for example, cinema, consumer, and other commercial projection systems, head-up displays, virtual reality displays, etc.
[0024] A neural network (NN) is a typical non-linear trainable circuit that includes multiple processing elements (PEs) also referred to as "neurons", "artificial neurons", or "NN nodes". In some embodiments, the neural network is a dedicated circuit, where different PEs are implemented as corresponding configurable sub-circuits that are physically connected by physical links to form a corresponding physical network. In some other embodiments, the neural network can be computer-simulated, in which case one or more electronic processors are programmed to perform signal processing similar to that of the corresponding dedicated circuit.
[0025] Each PE of a neural network typically has connections with one or more other PEs. This plurality of connections between PEs (physical or computer simulated) defines the topology of the neural network. In some topologies, PEs can be aggregated into layers. Different layers can have different types of PEs, which are configured to perform different corresponding kinds of transformations on their inputs. Signals can travel from a first PE layer (commonly referred to as the input layer) to a last PE layer (commonly referred to as the output layer). In some topologies, a neural network can have one or more intermediate PE layers (commonly referred to as hidden layers) located between the input PE layer and the output PE layer. Example PEs can scale, sum, and bias incoming signals and use an activation function to produce an output signal that is a static non - linear function of the biased sum. The resulting PE output can become one of the outputs of the neural network or be sent via the corresponding connection(s) to one or more other PEs. The corresponding weights and / or biases applied by each PE can be changed (e.g., optimized) during a training (learning) mode of operation and are typically fixed (i.e., constant) during a testing (working) mode of operation. Various embodiments disclosed herein can employ or rely on one or more neural networks.
[0026] Neural Radiance Field
[0027] A Neural Radiance Field (NeRF) implicitly represents a 3D scene, for example, using a neural network that takes 3D positions and viewing directions as inputs and generates corresponding predicted color textures and volume densities as outputs. The corresponding neural network can be trained using a set of 2D images with known camera poses and related intrinsic information. After being trained, the neural network can render the color textures and volume densities of any view of the 3D scene by querying the corresponding 3D positions and viewing directions of various pixels in the view.
[0028] In various examples, NeRF is configured to implicitly represent a scene as a function of a continuous 3D scene density σ and color c = (r, g, b) as a continuous five - dimensional (5D) input vector of spatial coordinates x = (x, y, z) and viewing direction , where x, y, z are Cartesian coordinates, and θ, are angles used in a spherical coordinate system. The continuous function that converts the 5D input (x, d) to a four - dimensional (4D) output (σ, r, g, b) is approximated using a neural network such as a Multi - Layer Perceptron (MLP).
[0029] In some NeRF configurations, the density σ is a function of only x, while the radiance c is a function of both the position x and the viewing direction d. The color of a pixel in the 2D rendering of a scene is computed via volume rendering. Given a set of images of a 3D scene with known camera parameters, the PE parameters of a neural network can be optimized via gradient descent by minimizing the photometric difference between the rendered image and the ground-truth image. A brief description of the corresponding MLP, positional encoding, and volume rendering for the example is given below in the remainder of this section.
[0030] An MLP is a fully-connected feedforward neural network that has an input layer, one or more hidden layers, and an output layer. The "fully-connected" property means that there are corresponding weighted connections between every PE from the previous layer and every PE in the adjacent subsequent layer. The entry of the corresponding weight matrix W is w ab , and this weight matrix includes the weight of the connection between the ath PE in the subsequent layer and the bth PE in the previous layer. For a layer with multiple PEs, its output can be expressed as a function f of the weighted sum of the input values , where n o is the number of PEs in the subsequent layer, and n i is the number of PEs in the previous layer. The function f is the activation function mentioned above. In mathematical terms,
[0031] O = f(WI + B) (1)
[0032] where the vector B is the bias vector. A non-limiting example of the activation function f is the ReLU (Rectified Linear Unit) function expressed as follows:
[0033]
[0034] Other suitable activation functions can also be used.
[0035] Since deep neural networks may inherently tend to preferentially learn lower-frequency functions, NeRF can be configured to additionally use a series γ of trigonometric functions to map the input to a higher-dimensional space and then pass the higher-dimensional input to the neural network to better utilize the high-frequency components to fit the output data. In various examples, the series γ is defined as follows:
[0036] γ(p) = (sin(2 0 πp), cos(2 0 πp), …, sin(2 L-1 πp), cos(2 L-1 πp)) (3)
[0037] where L is the number of the triangular components of γ. The series γ(·) is applied separately to each coordinate in x = (x, y, z) and to each component in the viewing direction. Before the positional encoding, the coordinates and the viewing direction are typically normalized to lie in the interval [-1, 1]. Note that Equation (3) describes an example implementation of the positional encoding. In some other implementations, the positional encoding is applied to some or all of the coordinates (x, y, z), but not to the viewing direction (e.g., see ). Figure 3 )
[0038] The example NeRF represents a 3D scene as the volume density and the directional emitted radiance at any point in 3D space, and uses volume rendering by numerical integration to render the color of any ray passing through the scene. Let r(t) = o + td r (t n < t < t f ) be the ray emitted from the camera through a given pixel, which traverses between the near boundary and the far boundary (t n and t f ), where o is the origin and d r is the unit vector of the direction of the ray. For a given t, r(t) is the 3D coordinate of the point on the ray whose distance from the camera origin is t. Then, for the selected K random quadrature points {t n between t f | k = 1, …, K}, the approximate expected color can be calculated using the following equation: k |k = 1, …, K}, the approximate expected color can be calculated using the following equation:
[0039]
[0040] where α(x) = 1 - exp(-x); and δ k = t k+1 - t k is the distance between adjacent quadrature points. As a simplification, Equation (4) uses c(t k ) and σ(t k ) to represent the color and the density at the point t k respectively.
[0041] In addition to color textures, some applications and / or users require and / or expect semantic segmentation of 3D scenes that the above NeRF does not provide. Advantageously, various embodiments disclosed herein address at least this problem in the prior art by providing a neural segmentation field implemented as an additional smaller neural network, which is trained and configured to use an initially pre-trained neural network (e.g., NeRF) to implicitly represent the semantic segmentation of the same 3D scene. According to an example embodiment, the training of the smaller neural network exploits the fact that the semantics, color textures, and geometric structures of 3D scenes are generally highly correlated to a large extent. Thus, given the 3D scene representation of the initially pre-trained neural network and a set of 2D segmentations with known camera poses and intrinsic parameters, the neural segmentation field can be learned in a relatively fast and efficient manner, and then the corresponding smaller neural network can be used to render a 2D segmentation map of the 3D scene corresponding to an arbitrarily selected viewpoint. At least some embodiments of the training process implement enhancements (e.g., in the form of one or more regularization terms), which can improve the segmentation quality of at least some 3D scenes. In some specific instances, example embodiments of the disclosed neural segmentation field advantageously achieve a relatively high segmentation quality (e.g., mIoU > 0.9) and a relatively small model size (e.g., < 0.4MB) for a relatively large number (e.g., > 25) of semantic classes. Various embodiments of the neural segmentation field can be applied to, for example, object editing (such as color manipulation or enhancement) and / or to 3D scene encoding and compression.
[0042] In this document, the acronym "mIoU" stands for mean intersection over union. The mIoU metric is typically calculated from a ground truth mask and a predicted mask and is conventionally used to evaluate the performance of various object detection and semantic segmentation models. Qualitatively, this metric can be understood as measuring the overlap between the ground truth mask and the predicted mask. Thus, the higher the mIoU value, the better the performance of the model. An mIoU value greater than 0.9 generally indicates excellent performance.
[0043] Neural segmentation field
[0044] Figures 1A to 1BFIG. is a block diagram illustrating example process flows (100, 101) for training and testing a neural segmentation field (110) according to various embodiments. A training phase (108) of the training process flow (100) generates a neural segmentation field (110), which is an implicit segmentation representation that represents segmentations in a 3D scene as a neural network based on a 3D scene representation (102), a set of 2D segmentation maps (104), and corresponding camera parameters (106). In some cases, the 3D scene representation (102) is a NeRF. In some other cases, the 3D scene representation (102) is another suitable 3D scene representation pre-trained to represent a 3D scene. For example, such a 3D scene representation can be a non-NN-based solution, such as spherical harmonics or another suitable implementation thereof. In various examples, semantic segmentation and instance segmentation can be represented by the neural segmentation field (110). A testing phase (114) of the testing process flow (101) uses the 3D scene representation (102), the neural segmentation field (110) generated in the training process flow (100), and camera parameters (112) for testing (e.g., new or arbitrary) viewpoints to generate 2D segmentation maps (116) corresponding to these testing viewpoints. Specific non-limiting examples of the training phase (108) and the testing phase (114) are described in more detail below, for example, with reference to Figure 3 and Figure 4 Specific non-limiting examples of the training phase (108) and the testing phase (114).
[0045] Figure 2 FIG. is a block diagram illustrating an MLP (200) that can be used to implement a neural segmentation field according to an embodiment. The MPL (200) has four layers (2101 to 2104). The first layer (2101) is an input layer. The next two layers (2102, 2103) are hidden layers. The fourth layer (2104) is an output layer. Generally, the MLP (200) can have N hidden layers, where N is a positive integer. Thus, Figure 2 the specific example of the MLP (200) illustrated in corresponds to N = 2. In some specific examples, the number N is in the range of 1 to 4.
[0046] Each of the layers (2101 to 2104) has M PEs (202), where M is a positive integer. In some specific examples, the number M is in the range of 16 to 256. Each PE (202 21 to 202 2M ) in the second layer (2102) is directly connected to receive a corresponding input from each PE (202 11 to 202 1M ) in the input layer (2101). Each PE (202 31 to 202 3M) Each PE in is directly connected to receive the corresponding input from each PE in the second layer (2102) (202 21 to 202 2M ). Each PE in the output layer (2104) (202 41 to 202 4M ) is directly connected to receive the corresponding input from each PE in the third layer (2103) (202 31 to 202 3M ). Similar PE / layer connections are implemented for additional embodiments corresponding to other values of N. In some examples, each PE in the PE (202) operates using the ReLU activation function (see Equation (2)). In other examples, other suitable activation functions may also be used.
[0047] In a specific example, the input to the MLP (200) includes the 3D coordinates of the position encoding (γ(x), γ(y), γ(z)) (see also Equation (3)) and a 4D input including radiance c = (r, g, b) and density σ. The output of the MLP (200) is a plurality of semantic logits where n s is the number of segmentation categories.
[0048] Figure 3 Shows a flowchart (300) illustrating the process of training an MLP (NN 1, 200) according to various embodiments. This training is performed using the camera parameters (302) of a set of training views. Each training view in these training views is associated (342) with a corresponding ground truth label map (340), which is used as one of the inputs to the loss function (330). Iterative adjustment (328) of the PE parameters used in the MLP (200) is performed to optimize (e.g., approximately minimize) the loss function (330) over the set of training views. Once the optimization criterion is met, the training process is terminated, and the corresponding "optimal" PE parameters are fixed and subsequently used in the test (working) operating mode, which is described in more detail below with reference to Figure 4 the test (working) operating mode.
[0049] For a selected view in the set of training views, the camera parameters (302) are used to obtain the 5D input (304, 306) of the pre-trained neural network (NN2, 310). In some specific examples, the neural network (310) implements NeRF (e.g., as described above). The 5D input (304, 306) includes the corresponding 3D spatial coordinates (x, y, z) (306) and the 2D viewing direction (304). In response to the 5D input (304, 306), the pre-trained neural network (310) generates a 4D output (r, g, b, σ) (314). The positional encoding (312) is applied to the 3D spatial coordinates (x, y, z) (306) to generate the corresponding series γ(p) (318) (e.g., see Equation (3)). The 4D output (r, g, b, σ) (314) of the neural network (310) and the series γ(p) (318) together provide the input to the MLP (200) being trained. In response to this input, the MLP (200) generates a plurality of semantic logic values (320).
[0050] Volume rendering (322) (which may be implemented, for example, in a manner consistent with Equation (4)) is used to compute the corresponding 2D semantic segmentation map of width W and height H from the logic values s (320) (324). An example mathematical equation for the computation of Figure (324) is as follows:
[0051]
[0052] Equation (5) uses the same notation as Equation (4). Then, based on the semantic segmentation map (324), the corresponding ground truth map (340), and one or more optional parameters (332), the loss function (330) is computed. In a particular example, the parameter (332) includes the RGB image corresponding to the selected view in the set of training views. Various embodiments of the loss function (330) are described in more detail below with reference to Equations (7) through (16) and Figures 5 to 7 and
[0053] Figure 4 A flowchart (400) illustrating the process of testing the MLP (NN 1, 200) according to various embodiments is shown. As already indicated above, the term "testing" refers to the operating mode of operation in which it can generate a semantic segmentation map corresponding to any view of the 3D scene. Before entering the operating mode of operation, the MLP (200) is trained as described above with reference to Figure 3 During the testing process (400), the PE parameters of the MLP (200) are constant, e.g., fixed at the optimal values determined using the training process (300). The pre-trained neural network (310) maintains the same configuration during the testing process (400) as in the training process (300).
[0054] For a selected (e.g., arbitrarily selected) new view, the corresponding camera parameters (402) are used to obtain the corresponding 5D input (304, 306) of the neural network (310). In response to the 5D input (304, 306), the neural network (310) generates a 4D output (σ, r, g, b) (314). A positional encoding (312) is applied to the 3D spatial coordinates (x, y, z) (306) to generate the corresponding series γ(p) (318). The 4D output (σ, r, g, b) (314) of the neural network (310) together with the series γ(p) (318) provides the input to the trained MLP (200). In response to this input, the trained MLP (200) generates a plurality of semantic logical values (320). Subsequent processing of the output of the MLP includes applying a softmax function (426) to the semantic logical values (320) to obtain the probability v that a ray / pixel belongs to class c. For example, for a pixel (h, w), the probability of being in class c can be calculated as follows:
[0055]
[0056] where S h,w,c is the semantic logical value of the pixel (h, w) of class c. The class probabilities representing the pixels of the new view constitute a probability map (428). The probability map (428) is further processed to generate a label map (430), where different objects within the new view are depicted and labeled according to their respective classes c (e.g., using color coding).
[0057] Example loss function
[0058] In a specific example, the loss function (330) is defined as the cross entropy between the predicted semantic segmentation map (324) and the ground truth segmentation map (340), e.g., as follows:
[0059]
[0060] where S h,w,c is the predicted logical value at the pixel (h, w) of the map (324) for the class c ∈ [0, n s - 1] which is the target label; n s is the total number of classes c, and is the ground truth label at the pixel (h, w) of the map (340) for class c. Iterative adjustment (328) is implemented using the Adam optimizer with a learning rate of 5 × 10 -4 for 20000 iterations.
[0061] In another specific example, the loss function (330) is defined as the weighted cross entropy between the predicted semantic segmentation map (324) and the ground truth segmentation map (340), for example, as follows:
[0062]
[0063] where w h,w is the weight at pixel (h, w). In some instances, the weighting function w at pixel p can be defined as follows:
[0064] w(p) = exp(M rgb (p)) (11)
[0065] where M rgb (p) ∈ [0, 1] is the magnitude of the image edge in the rendered 2D RGB image (332). Contrary to the cross entropy loss defined by equations (7) to (8) where each pixel is equally weighted, in the weighted cross entropy loss defined by equations (9) to (11), pixels on the object boundary have a higher weight compared to pixels relatively deep inside the object. The edges of the RGB image (332) are used to indicate the object boundaries in the maps (324, 340) because object boundaries typically occur on or very close to the RGB edges.
[0066] Figure 5 FIG. shows a flow chart (500) illustrating the calculation of the weighted cross entropy of the loss function (330) according to some examples. The inputs to the calculation (500) include the predicted semantic segmentation map (324), the ground truth segmentation map (340), and the RGB image (332). Edge detection (504) is performed on the RGB image (332) to detect the object edges therein. Then the weighting function w of equation (11) is calculated (502) such that object edge pixels are given a higher weight than object body pixels. In some examples, a deep learning-based edge detector can be used to implement the calculation blocks (502, 504). Then the loss function (330) is calculated according to equations (9) to (10).
[0067] In some other specific examples, the loss function (330) includes one or more weighted regularization terms added to the cross entropy of equation (8) or the weighted cross entropy of equation (10). Two illustrative examples of such regularization terms are described below with reference to Figures 6 to 7 In various other embodiments, other suitable regularization terms can also be used in addition to or in place of these example regularization terms. Equation (12) provides an example of the loss function (330) employing such a regularization term.
[0068] total_loss = weighted_cross_entropy + λ1 × pixelwise_regularization + λ2 × pairwise_regularization (12)
[0069] In this paper, pixelwise regularization and pairwise regularization are example regularization terms; λ1 and λ2 are constants, and their values can be chosen empirically. In some examples, λ1 and λ2 can be 0.01 and 0.001 respectively.
[0070] An example regularization term implements pixelwise edge regularization. The corresponding unary regularization term R(p) at pixel p is defined as:
[0071]
[0072] where M s is the segmentation edge map; and M rgb is the edge detected, for example, using the above edge detector. The map M s is calculated as the gradient magnitude on the segmentation probability map, which is expressed as follows:
[0073]
[0074] where G x (.) and G y (.) are Sobel gradient operators; and v c is the segmentation probability map for the c-th class. Intuitively, if the segmentation edge aligns with the RGB image edge, the regularization term R(p) in Equation (13) is relatively small. Otherwise, the regularization term R(p) in Equation (13) is relatively large. Therefore, the regularization term R(p) in Equation (13) helps to achieve more accurate object depiction in the predicted semantic segmentation map.
[0075] Figure 6 FIG. shows a flowchart (600) illustrating the calculation of the pixelwise regularization term (610) of the loss function (330) according to some examples. The input to the calculation (600) includes the predicted semantic segmentation map (324) and the RGB image (332). Edge detection (504) is performed on the RGB image (332) to detect the object edges therein. An edge map (604) is generated based on the object edges detected in block (504). A segmentation edge map (602) is generated based on the predicted semantic segmentation map (324). Then, the pixelwise edge regularization term (610) is calculated based on the edge maps (602, 604) according to Equation (13).
[0076] Another example regularization term implements pairwise margin regularization. The corresponding pairwise regularization term R(p, q) for pixels p and q is defined as:
[0077] R(p, q) = g(D rgb (p, q), D s (p, q)) (16)
[0078]
[0079] where D rgb and D s are the distances between two pixels p and q on the RGB image (332) and the predicted segmentation map (324), respectively; g is a function configured to penalize large segmentation distances for pixel pairs with small RGB distances. D rgb can be provided by a pre-trained neural network (310). The pairwise margin regularization term is configured to use D rgb to improve D s and thus improve S.
[0080] In a specific example, D rgb is the length of the shortest path P between two pixels p and q on a weighted graph G rgb which is a 4-neighbor graph with pixels as nodes and RGB image edge magnitudes as edge weights. The corresponding mathematical expression is shown as follows:
[0081] D rgb (p, q) = ∑ z∈P(p,q) M rgb (z) (18)
[0082] where P(p, q) is the shortest path between p and q on G rgb calculated using Dijkstra's algorithm; and D rgb (p, q) is the length of P(p, q). If there is an image edge between two pixels on G rgb , then D rgb is large. Otherwise, D rgb is small. D s (p, q) is the L2 distance between two probability vectors v, i.e., D s (p, q) = L2(v(p), v(q)).
[0083] Figure 7FIG. 700 shows a flow chart illustrating the calculation of the pairwise margin regularization term (710) of a loss function (330) according to some examples. The input to the calculation (700) includes the predicted semantic segmentation map (324) and the RGB image (332). Edge detection (504) is performed on the RGB image (332) to detect object edges therein. An edge map (604) is generated based on the object edges detected in block (504). A pairwise distance map (702) is generated based on the edge map (604) according to Equation (18). A pairwise distance map (704) is generated based on the predicted semantic segmentation map (324) and using the above definition of D s (p,q). Then, the pairwise margin regularization term (710) is calculated based on the distance maps (702, 704) according to Equation (16). In some specific examples, the pairwise margin regularization term (710) can be calculated on a local region around the pixel in question, the local region having a size of, for example, 8·8 or 16·16 pixels.
[0084] Example Hardware
[0085] Figure 8 FIG. 800 is a block diagram illustrating a computing device (800) according to an embodiment. The device (800) can be used, for example, to implement process flows (100, 101). The device (800) includes an input / output (I / O) device (810), an image processing engine (IPE, 820), and a memory (830). The I / O device (810) can be used to enable the device (800) to receive parameters of at least a 3D representation (102) and output parameters of at least a neural segmentation field (110) and a segmentation map (116). The I / O device (810) can also be used to connect the device (800) to a display.
[0086] The memory (830) may have a buffer to receive image data corresponding to a 3D scene to be rendered. The image data may be in the form of, for example, one or more image files. Once the image data is received, the memory (830) may provide portions of the data to the IPE (820) for processing therein. The IPE (820) includes a processor (822) and a memory (824). The memory (824) may store program code therein, which, when executed by the processor (822), enables the IPE (820) to perform image processing, including but not limited to image processing according to some or all of the above flowcharts (100, 101, 300, 400, 500, 600, 700). The program code may include (but is not limited to) program code for emulating various neural networks (e.g., the NN 1 (200) and NN 2 (310) described above). Once the IPE (820) generates the various above-mentioned diagrams by executing the corresponding portions of the code, the IPE (820) may perform its rendering process and provide the corresponding viewable image(s) for viewing on a display. The viewable image may be in the form of, for example, a suitable image file output via the I / O device (810).
[0087] According to the example embodiments disclosed above, for example in the Summary of the Invention section and / or with reference to any one or any combination of some or all of the figures in FIGS. 1 to Figure 8 A method of image processing is provided, which includes training, via a processor, a first neural network to render a segmentation map corresponding to an arbitrarily selected view of a 3D scene, the training including: calculating, by the processor, a color texture and a volume density corresponding to a selected training view of the 3D scene, the calculation being performed using a second neural network pre-trained to represent the 3D scene; generating, by the processor, a predicted segmentation map corresponding to the selected training view of the 3D scene, the generating being performed using the first neural network based on the color texture and the volume density; and adjusting, by the processor, configuration parameters of processing elements of the first neural network based on a loss function configured to receive a ground truth segmentation map corresponding to the selected training view as its first input and further configured to receive the predicted segmentation map as its second input.
[0088] In some embodiments of the above method, the first neural network is a multi-layer perceptron; and wherein, the second neural network is another multi-layer perceptron implementing a neural radiance field.
[0089] In some embodiments of any of the above methods, the first neural network is configured to output semantic logic values corresponding to a selected training view of the 3D scene in response to the color texture and the volume density.
[0090] In some embodiments of any of the above methods, the training further includes: calculating a cross-entropy between the predicted segmentation map and the ground truth segmentation map; and using the cross-entropy to obtain a loss function value.
[0091] In some embodiments of any of the above methods, the loss function is configured to receive a color image of the 3D scene corresponding to a selected training view as its third input.
[0092] In some embodiments of any of the above methods, the training further includes: detecting object edges in the color image; assigning different weights to object edge pixels and object body pixels; calculating a cross-entropy between the predicted segmentation map and the ground truth segmentation map using the different weights; and using the cross-entropy to obtain a loss function value.
[0093] In some embodiments of any of the above methods, the loss function includes one or more regularization terms that promote alignment of corresponding edges in the predicted segmentation map and the ground truth segmentation map.
[0094] In some embodiments of any of the above methods, generation is performed in response to position encoding of 3D spatial coordinates corresponding to a selected training view.
[0095] In some embodiments of any of the above methods, the image processing method further includes creating a segmentation map corresponding to an arbitrarily selected view of the 3D scene using a processor, and the creating is performed using a first neural network after training.
[0096] In some embodiments of any of the above methods, the image processing method further includes: applying a softmax function to the segmentation map corresponding to an arbitrarily selected view of the 3D scene using a processor to generate a probability map; and converting the probability map to a semantic label map representing the arbitrarily selected view of the 3D scene using a processor.
[0097] According to another example embodiment disclosed above, for example in the Summary of the Invention section and / or with reference to any one or any combination of some or all of FIGS. 1 to Figure 8 a non-transitory computer-readable medium storing instructions is provided, which when executed by a processor cause the processor to perform operations including any of the above methods.
[0098] According to yet another example embodiment disclosed above, for example in the Summary of the Invention section and / or with reference to any one or any combination of some or all of FIGS. 1 to Figure 8For any one or any combination of some or all of the figures, there is provided an image processing apparatus, the image processing apparatus including: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to cause the apparatus, using the at least one processor, to at least train a first neural network to render a segmentation map corresponding to an arbitrarily selected view of a 3D scene; and wherein, to train the first neural network, the apparatus is configured to: calculate a color texture and a volume density corresponding to a selected training view of the 3D scene, the calculation being performed using a second neural network pre-trained to represent the 3D scene; generate a predicted segmentation map corresponding to the selected training view of the 3D scene, the generation being performed using the first neural network based on the color texture and the volume density; and adjust configuration parameters of processing elements of the first neural network based on a loss function that receives a ground truth segmentation map corresponding to the selected training view as its first input and further receives the predicted segmentation map as its second input.
[0099] In some embodiments of the above apparatus, the first neural network is a multi-layer perceptron; and wherein the second neural network is another multi-layer perceptron implementing a neural radiance field.
[0100] In some embodiments of any of the above apparatus, the first neural network is configured to output semantic logic values corresponding to a selected training view of the 3D scene in response to the color texture and the volume density.
[0101] In some embodiments of any of the above apparatus, to train the first neural network, the apparatus is further configured to: calculate the cross entropy between the predicted segmentation map and the ground truth segmentation map; and obtain a loss function value using the cross entropy.
[0102] In some embodiments of any of the above apparatus, the loss function is configured to receive a color image of the 3D scene corresponding to the selected training view as its third input.
[0103] In some embodiments of any of the above apparatus, to train the first neural network, the apparatus is further configured to: detect object edges in the color image; assign different weights to object edge pixels and object body pixels; calculate the cross entropy between the predicted segmentation map and the ground truth segmentation map using the different weights; and obtain a loss function value using the cross entropy.
[0104] In some embodiments of any of the above apparatus, the loss function includes one or more regularization terms that promote alignment of corresponding edges in the predicted segmentation map and the ground truth segmentation map.
[0105] In some embodiments of any of the devices in the above-described apparatus, at least one memory and program code are further configured to cause the apparatus, using at least one processor, to create a segmentation map corresponding to an arbitrarily selected view of a 3D scene, the creation being performed using a first neural network after training.
[0106] In some embodiments of any of the devices in the above-described apparatus, at least one memory and program code are further configured to cause the apparatus, using at least one processor, to: apply a softmax function to a segmentation map corresponding to an arbitrarily selected view of a 3D scene to generate a probability map; and convert the probability map into a semantic label map representing the arbitrarily selected view of the 3D scene.
[0107] According to another example embodiment disclosed above, for example in the Summary section and / or with reference to any one or any combination of some or all of the figures in FIGS. 1 to Figure 8 An image processing method is provided, which includes testing a first neural network via a processor to render a segmentation map corresponding to an arbitrarily selected view of a 3D scene, the testing including: calculating, using the processor, a color texture and a volume density corresponding to a selected view of the 3D scene, the calculation being performed using a 3D representation pre-trained to represent the 3D scene; and generating, using the processor, a segmentation map corresponding to the selected view of the 3D scene, the generation being performed using the first neural network based on the color texture and the volume density; and wherein the first neural network has been trained using the processor with the pre-trained 3D representation.
[0108] According to another example embodiment disclosed above, for example in the Summary section and / or with reference to any one or any combination of some or all of the figures in FIGS. 1 to Figure 8 An image processing apparatus is provided, which includes: at least one processor; and at least one memory including program code; wherein the at least one memory and the program code are configured to cause the apparatus, using at least one processor, to at least test a first neural network to render a segmentation map corresponding to an arbitrarily selected view of a 3D scene; wherein, to test the first neural network, the apparatus is configured to: calculate a color texture and a volume density corresponding to a selected view of the 3D scene, the calculation being performed using a 3D representation pre-trained to represent the 3D scene; and generate a segmentation map corresponding to the selected view of the 3D scene, the generation being performed using the first neural network based on the color texture and the volume density; and wherein the first neural network has been trained using the pre-trained 3D representation.
[0109] Regarding the processes, systems, methods, heuristics, etc. described herein, it should be understood that although the steps of these processes, etc. have been described as occurring in a particular ordered sequence, these processes may be practiced using the described steps executed in an order different from that described herein. Further, it should be understood that certain steps may be performed simultaneously, other steps may be added, or certain steps described herein may be omitted. In other words, the process descriptions herein are provided for the purpose of illustrating certain embodiments and should in no way be construed as limiting the claims.
[0110] Accordingly, it should be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided will be apparent by reading the above description. The scope should not be determined with reference to the above description, but rather should be determined with reference to the appended claims and the full scope of equivalents to which those claims are entitled. It is expected and intended that future developments will occur in the technologies discussed herein, and the disclosed systems and methods will be incorporated into such future embodiments. In summary, it should be understood that this application is subject to modification and variation.
[0111] All terms used in the claims are intended to be given the broadest reasonable interpretation and ordinary meaning as understood by those who are knowledgeable in the technologies described herein, unless an explicit contrary indication appears herein. In particular, the use of singular articles such as "a," "the," etc. should be understood to recite one or more of the indicated elements unless the claim recites an explicit contrary limitation.
[0112] A summary of the disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. This summary is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Additionally, in the foregoing detailed description, it can be seen that for the purpose of putting the disclosure into a coherent whole, various features are grouped together in various embodiments. The methods of the disclosure should not be construed as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as reflected in the appended claims, the inventive subject matter lies in less than all of the features of a single disclosed embodiment. Accordingly, the appended claims are hereby incorporated into the detailed description, with each claim standing on its own as a separately claimed subject matter.
[0113] Although this disclosure includes references to illustrative embodiments, this specification is not intended to be construed in a limiting sense. Various modifications of the described embodiments and other embodiments within the scope of the disclosure (which are obvious to those skilled in the art to which the disclosure pertains) are considered to be within the principles and scope of the disclosure as expressed, for example, in the appended claims.
[0114] Some embodiments may be implemented as circuit-based processes, including possible implementations on a single integrated circuit.
[0115] Some embodiments may be embodied in the form of methods as well as apparatuses for practicing these methods. Some embodiments may also be embodied in the form of program code recorded in a tangible medium such as a magnetic recording medium, an optical recording medium, a solid-state memory, a floppy disk, a CD-ROM, a hard disk drive, or any other non-transitory machine-readable storage medium, wherein when the program code is loaded into and executed by a machine such as a computer, the machine becomes an apparatus for practicing the various embodiments described herein. Some embodiments may also be embodied in the form of, for example, program code stored in a non-transitory machine-readable storage medium, including being loaded into and / or executed by a machine, wherein when the program code is loaded into and executed by a machine such as a computer or a processor, the machine becomes an apparatus for practicing the various embodiments described herein. When implemented on a general-purpose processor, the program code segments combine with the processor to provide a unique device that operates in a manner similar to specific logic circuitry.
[0116] Unless otherwise expressly stated, each numerical value and range should be interpreted as approximate, as if the word “about” or “approximately” preceded the value or range.
[0117] The use of figure numbers and / or reference numerals in the claims is intended to identify one or more possible embodiments of the claimed subject matter to facilitate interpretation of the claims. Such use should not be construed as necessarily limiting the scope of these claims to the embodiments shown in the corresponding figures.
[0118] Although the elements (if any) in the following method claims are recited in a specific sequence with corresponding labels, unless the claim recitation otherwise implies a specific sequence for implementing some or all of these elements, these elements are not necessarily intended to be limited to being implemented in that specific sequence.
[0119] The mention of “one embodiment” or “an embodiment” herein means that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present disclosure. In this specification, the phrase “in one embodiment” does not necessarily refer to the same embodiment everywhere it appears, nor is it necessarily a separate or additional embodiment that is mutually exclusive of other embodiments. The same applies to the term “implementation”.
[0120] Unless otherwise specified herein, the use of ordinal adjectives such as “first,” “second,” “third,” etc. to refer to one of a plurality of similar objects only indicates that different instances of such similar objects are being referred to and is not intended to imply that the similar objects so referred to must be in a corresponding order or sequence in time, in space, in ranking, or in any other way.
[0121] Unless otherwise specified herein, in addition to its ordinary meaning, the conjunction “if” may also or alternatively be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” such construction being dependent upon the particular context. For example, the phrase “if determined” or “if detected [stated condition]” may be construed to mean “when determining” or “in response to determining” or “when detecting [stated condition or event]” or “in response to detecting [stated condition or event].”
[0122] Also for the purposes of this specification, the terms “couple,” “coupling,” “coupled,” “connect,” “connecting,” “connected” refer to any means known in the art or developed later that permits energy to be transferred between two or more elements, and the insertion of one or more additional elements may be contemplated, although this is not required. Conversely, the terms “directly coupled,” “directly connected,” etc. imply the absence of such additional elements.
[0123] The functions of the various elements shown in the drawings (including any functional blocks labeled as "processor" and / or "controller") can be provided by dedicated hardware as well as the use of hardware capable of executing software associated with appropriate software. When provided by a processor, the functions can be provided by a single dedicated processor, by a single shared processor, or by multiple individual processors (some of which may be shared). Moreover, the explicit use of the term "processor" or "controller" should not be construed to refer only to hardware capable of executing software, but can implicitly and non-limitingly include digital signal processor (DSP) hardware, network processors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), read only memory (ROM) for storing software, random access memory (RAM), and non-volatile storage devices. Other conventional and / or custom hardware may also be included. Similarly, any switches shown in the drawings are merely conceptual. Their functions can be performed by the operation of program logic, by dedicated logic, by the interaction of program control and dedicated logic, or even manually, and the specific technique can be selected by the implementer according to a more specific understanding of the context.
[0124] As used in this application, the terms "circuit" and "circuitry" can refer to one or more or all of the following: (a) only hardware circuit implementations (such as only in analog and / or digital circuits); (b) a combination of hardware circuits and software, such as the following (where applicable): (i) a combination of (multiple) analog and / or digital hardware circuits and software / firmware, and (ii) any part of (multiple) hardware processors with software (including (multiple) digital signal processors), software, and (multiple) memories, which work together to enable a device such as a mobile phone or a server to perform various functions); and (c) (multiple) hardware circuits and / or (multiple) processors, such as (multiple) microprocessors or a part of (multiple) microprocessors, which require software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation. This definition of circuit applies to all uses of the term in this application, including all uses in any claims. As a further example, as used in this application, the term circuit also encompasses implementations that are only hardware circuits or processors (or multiple processors) or parts of hardware circuits or processors and their (or their) accompanying software and / or firmware. For example and where applicable to a particular claim element, the term circuit also encompasses a baseband integrated circuit or a processor integrated circuit for a mobile device or a similar integrated circuit in a server, a cellular network device, or other computing or network devices.
[0125] Those skilled in the art should understand that any block diagrams herein represent conceptual views of illustrative circuits embodying the principles of the present disclosure. Similarly, it should be understood that any flowcharts, flow diagrams, state transition diagrams, pseudocode, etc. represent various processes that can be substantially represented in a computer-readable medium and thus executed by a computer or a processor, whether or not such computer or processor is explicitly shown.
[0126] The "Summary of the Invention" in this specification is intended to introduce some exemplary embodiments, where additional embodiments are described in the "Detailed Description" and / or with reference to one or more of the drawings. The "Summary of the Invention" is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.
Claims
1. An image processing method, including training a first neural network (NN) via a processor to render a segmentation map corresponding to an arbitrarily selected view of a three-dimensional (3D) scene, the training including: Using the processor to calculate a color texture and a volume density corresponding to a selected training view of the 3D scene, the calculation being performed using a 3D representation pre-trained to represent the 3D scene; Using the processor to generate a predicted segmentation map corresponding to the selected training view of the 3D scene, the generation being performed using the first neural network based on the color texture and the volume density; And Using the processor to adjust configuration parameters of NN nodes of the first neural network based on a loss function, the loss function being configured to receive a ground truth segmentation map corresponding to the selected training view as its first input and further being configured to receive the predicted segmentation map as its second input.
2. The image processing method according to claim 1, Among them, The first neural network is a multi-layer perceptron; and Wherein, the 3D representation is a second neural network implementing a neural radiance field.
3. The image processing method according to claim 1 or claim 2, wherein, The first neural network is configured to output semantic logic values corresponding to the selected training view of the 3D scene in response to the color texture and the volume density.
4. The image processing method according to any one of claims 1 to 3, wherein, The training further includes: Calculating a cross entropy between the predicted segmentation map and the ground truth segmentation map; and Using the cross entropy to obtain a loss function value.
5. The image processing method according to any one of claims 1 to 4, wherein, The loss function is configured to receive a color image of the 3D scene corresponding to the selected training view as its third input.
6. The image processing method according to claim 5, wherein, The training further includes: Detecting object edges in the color image; Assigning different weights to object edge pixels and object body pixels; Calculating a cross entropy between the predicted segmentation map and the ground truth segmentation map using the different weights; and Using the cross entropy to obtain a loss function value.
7. The image processing method according to any one of claims 1 to 6, wherein, The loss function includes one or more regularization terms that promote alignment of corresponding edges in the predicted segmentation map and the ground truth segmentation map.
8. The image processing method according to any one of claims 1 to 7, wherein, The generation is performed in response to position encoding of 3D spatial coordinates corresponding to the selected training view.
9. The image processing method according to any one of claims 1 to 8, further including: Using the processor to create a segmentation map corresponding to an arbitrarily selected view of the 3D scene, the creation being performed using the first neural network after the training; Using the processor to apply a softmax function to the segmentation map corresponding to the arbitrarily selected view of the 3D scene to generate a probability map; And Using the processor to convert the probability map into a semantic label map representing the arbitrarily selected view of the 3D scene.
10. A non-transitory computer-readable medium storing instructions, the instructions when executed by the processor cause the processor to perform operations including the method according to any one of claims 1 to 9.
11. An image processing apparatus, including: At least one processor; And At least one memory, the at least one memory including program code; wherein the at least one memory and the program code are configured to cause the apparatus, using the at least one processor, to at least train a first neural network (NN) to render a segmentation map corresponding to an arbitrarily selected view of a three-dimensional (3D) scene; and wherein, to train the first neural network, the apparatus is configured to: calculate a color texture and a volume density corresponding to a selected training view of the 3D scene, the calculation being performed using a 3D representation pre-trained to represent the 3D scene; generate a predicted segmentation map corresponding to the selected training view of the 3D scene, the generation being performed using the first neural network based on the color texture and the volume density; and adjust configuration parameters of NN nodes of the first neural network based on a loss function that receives a ground truth segmentation map corresponding to the selected training view as its first input and further receives the predicted segmentation map as its second input.
12. The image processing apparatus according to claim 11, Among them, wherein the first neural network is a multi-layer perceptron; and wherein the 3D representation is a second neural network implementing a neural radiance field.
13. The image processing apparatus according to claim 11 or claim 12, wherein, The first neural network is configured to output semantic logic values corresponding to the selected training view of the 3D scene in response to the color texture and the volume density.
14. The image processing apparatus according to any one of claims 11 to 13, wherein, To train the first neural network, the apparatus is further configured to: calculate a cross entropy between the predicted segmentation map and the ground truth segmentation map; and use the cross entropy to obtain a loss function value.
15. The image processing apparatus according to any one of claims 11 to 14, wherein, The loss function is configured to receive a color image of the 3D scene corresponding to the selected training view as its third input.
16. The image processing apparatus according to claim 15, wherein, To train the first neural network, the apparatus is further configured to: detect object edges in the color image; assign different weights to object edge pixels and object body pixels; calculate a cross entropy between the predicted segmentation map and the ground truth segmentation map using the different weights; and use the cross entropy to obtain a loss function value.
17. The image processing apparatus according to any one of claims 11 to 16, wherein, The loss function includes one or more regularization terms that promote alignment of corresponding edges in the predicted segmentation map and the ground truth segmentation map.
18. The image processing apparatus according to any one of claims 11 to 17, wherein, The at least one memory and the program code are further configured to cause the apparatus, using the at least one processor, to: create a segmentation map corresponding to an arbitrarily selected view of the 3D scene, the creation being performed using the first neural network after the training; apply a softmax function to the segmentation map corresponding to the arbitrarily selected view of the 3D scene to generate a probability map; and convert the probability map to a semantic label map representing the arbitrarily selected view of the 3D scene.
19. An image processing method, comprising testing a first neural network via a processor to render a segmentation map corresponding to an arbitrarily selected view of a three-dimensional (3D) scene, the testing including: Compute, using the processor, a color texture and a volume density corresponding to a selected view of the 3D scene, the computation being performed using a 3D representation pre-trained to represent the 3D scene; and Generate, using the processor, a segmentation map corresponding to the selected view of the 3D scene, the generation being performed using the first neural network based on the color texture and the volume density; and wherein the first neural network has been trained using the processor with the pre-trained 3D representation.
20. An image processing apparatus, comprising: At least one processor; and At least one memory including program code; wherein the at least one memory and the program code are configured to cause the apparatus, using the at least one processor, to at least test a first neural network to render a segmentation map corresponding to an arbitrarily selected view of a three-dimensional (3D) scene; wherein, to test the first neural network, the apparatus is configured to: Compute a color texture and a volume density corresponding to a selected view of the 3D scene, the computation being performed using a 3D representation pre-trained to represent the 3D scene; and Generate a segmentation map corresponding to the selected view of the 3D scene, the generation being performed using the first neural network based on the color texture and the volume density; and wherein the first neural network has been trained using the pre-trained 3D representation.