Scalable three-dimensional scene representation with neural field modeling
By adopting a two-layer system in the three-dimensional scene representation, the basic layer of MPI representation and the enhancement layer of neural field encoding are used to solve the contradiction between terminal device computing capabilities and three-dimensional reconstruction quality, and an efficient and scalable three-dimensional scene representation is achieved.
Patent Information
- Application Number
- CN202380071209.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-08
- Filing Date
- 2023-09-05
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to handle the computing power of the terminal device and the required three-dimensional reconstruction quality while ensuring communication bandwidth, especially in the ability to provide high-quality specular highlights and non-Lambertian reflections.
A two-layer system is adopted, where the base layer ensures the baseline quality of low computational load through multi-plane imaging (MPI) representation, and the enhancement layer encodes the differential codes of specular highlights and non-Lambertian reflections through neural field encoding, and uses residual neural field networks to improve rendering quality.
It realizes the provision of scalable three-dimensional scene representation on different terminal devices, which not only ensures baseline quality, but also provides higher rendering quality and wider color gamut when computing power allows.
Smart Images

Figure CN120019657A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Application No. 63 / 404,885, filed on September 8, 2022, which is hereby incorporated by reference in its entirety. Technical Field
[0003] This document generally relates to images. More specifically, embodiments of the present invention relate to scalable three-dimensional scene representation using a two-layer approach, where neural field modeling is used to enhance layer information. Background Art
[0004] In recent years, there has been a growing interest in efficient modeling and representation of three-dimensional scenes. Three-dimensional scenes can be used for a variety of applications, including volumetric imaging, virtual reality, or augmented reality. Deep learning techniques have shown promising results in three-dimensional scene representation and reconstruction; however, not all devices can handle the computational load associated with such approaches. As the inventors of the present invention have appreciated, it is desirable to provide scalable three-dimensional scene representation under various scalability standards, and thus improved techniques for three-dimensional scene representation are described herein.
[0005] The term "metadata" in this article refers to any auxiliary information transmitted as part of the encoded bitstream and that can help a decoder render a decoded image or 3D scene. Such metadata may include, but is not limited to, color space or gamut information, reference display parameters, camera parameters, neural network parameters, etc.
[0006] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Therefore, unless otherwise indicated, it should not be assumed that any approach described in this section is considered prior art simply because it is included in this section. Similarly, problems associated with one or more approaches identified on the basis of this section should not be assumed to have been recognized in any prior art unless otherwise indicated. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Embodiments of the invention are illustrated by way of example and not limitation in the accompanying drawings in which like reference numerals indicate similar elements and in which:
[0008] Figure 1A Describes an example of an encoder for scalable three-dimensional scene representation under a general scalability framework according to an embodiment of the present invention;
[0009] Figure 1B Describes an example of a decoder for scalable three-dimensional scene representation under a general scalability framework according to an embodiment of the present invention;
[0010] Figure 1C Describes an example of an encoder for scalable three-dimensional scene representation under the PSNR standard according to an embodiment of the present invention;
[0011] Figure 1D Describes an example of a decoder for scalable three-dimensional scene representation under the PSNR standard according to an embodiment of the present invention;
[0012] Figure 2A Depicts an example of an encoder for scalable three-dimensional scene representation under a PSNR standard and a multi-plane image (MPI) representation according to an embodiment of the present invention; and
[0013] Figure 2B Depicted is a decoder example for scalable three-dimensional scene representation under the PSNR standard and MPI representation according to an embodiment of the present invention. DETAILED DESCRIPTION
[0014] Example embodiments related to scalable three-dimensional scene representation are described herein. In the following description, for the purpose of explanation, many specific details are listed to provide a thorough understanding of various embodiments of the present invention. However, it is apparent that various embodiments of the present invention can be implemented without these specific details. In other examples, well-known structures and devices are not described in detail to avoid unnecessarily blocking, obscuring or confusing embodiments of the present invention.
[0015] Overview
[0016] The example embodiments described herein relate to scalable three-dimensional scene representation. In an embodiment, to generate a scalable three-dimensional scene representation in an encoder, a processor performs the following operations:
[0017] accessing a first set of images in a first format for a scene (102);
[0018] Based on the first set of images, generating a first three-dimensional scene representation (107) for the scene;
[0019] accessing a second set of images in a second format for the scene (104);
[0020] generating a second three-dimensional scene representation for the scene based on the second set of images (112), wherein the second three-dimensional scene representation is superior to the first three-dimensional scene representation according to one or more quality criteria;
[0021] Generate an output image residual (122) based on the first three-dimensional scene representation and the second three-dimensional scene representation using a set of original viewpoint positions and a set of new viewpoint positions;
[0022] Using the output image residual, a residual neural field network (125) is trained to generate a predicted residual image that approximates the output image residual;
[0023] A first three-dimensional scene representation (107) for the scene is transmitted as a base layer; and information about the trained residual neural field network is transmitted as an enhancement layer.
[0024] In an embodiment, to generate an output three-dimensional scene in a decoder, the processor performs the following operations:
[0025] Receiving a base layer bitstream (107), the base layer bitstream comprising a first three-dimensional scene representation (107) for a scene;
[0026] Receiving an enhancement layer bitstream (127), the enhancement layer bitstream including information for reconstructing a trained residual neural field network;
[0027] Given an observer position:
[0028] generating a first three-dimensional output of the scene based on the first three-dimensional scene representation (132);
[0029] generating an image residual (145) using the observer position and the trained residual neural field network; and
[0030] The first three-dimensional output of the scene is combined with the image residual to generate an enhanced three-dimensional output of the scene.
[0031] 3D scene representation and neural fields
[0032] Introduction
[0033] There are currently a variety of 3D scene representation models, including multi-view plus depth (MVD) methods (see reference [6]), multi-planar imaging (MPI) (see reference [5]), and neural radiance field (NeRF) (see reference [2]) representations. Among all these methods, there are three main evaluation criteria: (1) computational complexity during training and testing; (2) bit size (bandwidth requirement) and model size of the scene representation; and (3) quality of 3D scene reconstruction. In practice, there are a variety of end devices, and applications need to handle the computational power of the end application and the required 3D reconstruction quality while ensuring communication bandwidth. Some devices can only withstand a lower computational load, but their users can accept lower quality. For high-end devices, it is feasible to increase the computational load to obtain better quality. In order to cover a wide range of needs and requirements, the embodiments of this article propose a two-layer system with a base layer (BL) that can meet a set of baseline requirements and an enhancement layer (EL) that can enhance the user experience. The proposed framework can also incorporate various scalability criteria based on peak signal-to-noise ratio (PSNR), dynamic range, color gamut, spatial resolution, temporal frame rate, etc.
[0034] For example, an MPI representation can be used for the base layer because it has extremely low decoding computational overhead. Such a base layer ensures that the encoded bitstream can be widely deployed on a variety of devices and that baseline quality can be maintained. However, MPI lacks the ability to provide a large number of specular highlights (non-Lambertian reflectance; for example, transparent materials belong to the non-Lambertian system). To provide these specular highlights, the difference between a 3D scene with specular highlights and MPI can be encoded into an enhancement layer using neural field coding. The base layer can be encoded (compressed) using traditional codecs such as AVC, HEVC, VVC, AV1, while the enhancement layer can carry the neural network coefficients representing the neural field. Once the device has more computational power, the enhancement layer can be decoded and added on top of the base layer to provide better rendering quality.
[0035] Compared to using a single-layer solution neural network (such as NeRF) to directly encode the 3D scene, there are some other benefits including the following. Neural network solutions (such as NeRF), or MLP (Multi-layer Perceptron) as it is commonly known, need to be trained for specific scenes, which can be a problem for some applications. In contrast, MPI can use pre-trained networks. If the MLP is only used for the residual, the neural network (NN) can be greatly simplified and the training time can be significantly reduced.
[0036] For MLP (e.g. NeRF), the model size is about 5M bytes per image scene. Directly transmitting such a model for video sequences would put a heavy burden on the network. Moreover, the compressibility of such NN representations is still under research. If MLP is used for the residual layer, the transmission bitrate can be significantly reduced.
[0037] In some embodiments, the enhancement layer is outside the encoding loop. Therefore, quality enhancement can be provided by simply adding NN residual information to a scene rendered using only the base layer. In addition, operations outside the encoding loop do not require bit-exact processing. The platform can choose floating-point operations or fixed-point operations to suit its computing environment.
[0038] In an embodiment, without limitation, the NN coefficients may be carried in the bitstream or downloaded from an external device, for example using the syntax defined in reference
[13] (see also reference [4]).
[0039] Extensibility allows the generation of enhancement layers with a variety of quality criteria, including:
[0040] PSNR: Improves 3D rendering quality by adding an enhancement layer residual on top of a lower quality base layer.
[0041] Dynamic range: A high dynamic range (HDR) 3D scene can be generated by adding an enhancement layer residual on top of a standard dynamic range (SDR) 3D scene to enhance the dynamic range of the rendered image;
[0042] Color gamut: A wider color gamut 3D scene can be generated by adding an enhancement layer residual on a narrower SDR color gamut 3D scene to enhance the color gamut;
[0043] Spatial resolution: Using an upscaled version of the base layer, enhancement layer information can be added to enhance the details of the final scene at a resolution higher than the base layer resolution;
[0044] Temporal frame rate: Frame rate interpolation can be applied on the base layer and then the neural field residual is added to generate a higher frame rate output;
[0045] Any combination of the above scalability criteria.
[0046] Neural Field
[0047] The term "neural field" refers to a coordinate-based fully connected neural network (see reference [1]). A neural network connects multiple layers of artificial neurons to learn to nonlinearly map a fixed-size input to a fixed-size output. A multilayer perceptron (MLP) neural network can approximate any function through learned parameters. Therefore, a multilayer perceptron (MLP) can be used to build a neural field. In the following discussion, the terms MLP and neural field will be used interchangeably.
[0048] The end-to-end MLP network consists of K layers of weights {W k} and bias {b k} parameters. These parameters are expressed as Φ = {{W k},{b k}}. The input of the MLP network is x, and the output is in,
[0049]
[0050] With the ground truth signal y gt In the case of , the formal problem statement for optimizing the parameter set Φ is given as:
[0051]
[0052] Where D() represents the loss / error function.
[0053] In some 3D scene representations, such as NeRf (see Ref. [2]), the input x consists of a spatial position (x, y, z) and a viewpoint direction (θ, φ), and the outputs are the volume density (σ) and the view-dependent emission amplitude (r, g, b) at these coordinates.
[0054] Positional encoding
[0055] Neural scenes lose frequency details. To solve this problem, applying position encoding is a common solution. In position encoding, the network input is mapped to a higher dimensional space. This is because neural networks are more biased towards learning low-frequency functions. Therefore, typical neural networks cannot represent high-frequency changes in color and geometry. For neural scene representation, by mapping the position coordinate p from R to R 2L (where L is the frequency number), which can significantly improve the performance of the neural network. The typical mapping γ acting on the coordinate p can be expressed as:
[0056]
[0057] Among them, {l0,l1,…,l L-1} is an integer. In a typical setting, l k =k.
[0058] Alternatively, parameter encoding can also be applied, i.e., additional trainable parameters (besides weights and biases) are arranged in an auxiliary data structure (such as a grid or a tree), and these parameters are looked up and (optionally) interpolated based on the input vector.
[0059] To help alleviate high frequency modeling, alternative solutions include using periodic functions as activation functions (see SIREN in reference [7]).
[0060] Forward Mapping
[0061] In some applications, the output of the MLP is not directly needed, but another mapping is needed. For example, in NeRF, the output of the MLP at the coordinate query point (x, y, z, θ, φ) is (σ, r, g, b). In order to construct the projected 2D image, volume rendering is required by querying all particles on each ray and calculating the RGB values of the final rendering.
[0062] In the proposed embodiment, the output of the neural residual network is already the rendered RGB residual. The RGB residual can be directly added on top of the rendered new view of the base layer. Different architectural designs will be discussed next.
[0063] A general framework for scalable 3D representations
[0064] Consider a set of images The same scene is captured from several different viewpoint positions and denoted as {t g In an embodiment, the collected image set can be used to construct a first three-dimensional scene representation algorithm (The parameter is Φ b ), the algorithm will be used as the base layer. Given the query viewpoint position, then the original viewpoint position {t g} and the new viewpoint position {t n} Render the scene image. g The rendered image at} is represented as And will be in {t n The rendered image at} is represented as The base layer should provide a minimum (base level) 3D representation quality suitable for typical decoding environments.
[0065] Next, a second 3D scene representation algorithm may be used that can provide a higher level of quality than the base level. As previously mentioned and discussed in detail later, this level of quality improvement may include PSNR improvements, higher bit depth, wider color gamut, etc. Depending on the scalability criteria, the same training data set or a different training data set may be applied to obtain the model (with parameter Φ s ). As mentioned above, you can use At the original viewpoint position {t g} and the new viewpoint position {t n} renders the image. g The rendered image at} is represented as And will be in {t n The rendered image at} is represented as
[0066] In an embodiment, the residual image can be obtained by and a second 3D scene representation Extract the rendering difference to generate. At the original viewpoint position {t g} and the new viewpoint position {t n}, where
[0067]
[0068] In an embodiment, the two sets of residual images and Both are used to train the third neural residual network MLP (Parameter Φ r ). Note that in this case, the MLP takes as input the image coordinates (x,y) with position encoding and the viewpoint position t; and outputs the RGB value for the pixel position (x,y) as
[0069]
[0070] Wherein, as mentioned above, γ() represents the position encoding function.
[0071] Unlike NeRF, which requires volume rendering, neural residual does not need to go through forward mapping to obtain a rendered 2D image. The output of the MLP is already in the RGB domain. The main goal of the neural residual network is to take an arbitrary viewpoint position {t} and output a predicted residual image The optimization process can be expressed as follows:
[0072]
[0073] In an embodiment, the base model parameter sets Φ and Φ can be compressed by MPEG NNC (see reference [4]). b and residual model parameter set Φ r Other embodiments may use a three-dimensional representation that does not involve a neural network. For example, the base model parameter set Φ b It can represent Multi-View Texture (MVC), Multi-View Texture plus Depth (MVC+D or MVD), or MPI formats. These formats can be used to render 3D scenes and can be compressed using existing single-layer or multi-layer codecs such as AVC, HEVC, VVC, MIV (MPEG Immersive Video), etc.
[0074] Figure 1A Depicts an example processing pipeline for encoding scalable 3D representations using a general framework that supports a variety of scalable standards such as:
[0075] a.PSNR Scalability
[0076] b. Dynamic range scalability
[0077] c. Color gamut scalability
[0078] d. Spatial resolution scalability
[0079] e. Time frame rate scalability
[0080] like Figure 1A As shown, the base layer includes a first unit (105) for generating a first baseline three-dimensional scene representation (107). The input of this unit is a first set of reference input images (102) for the scene, and the first set of reference input images is in a first format. The three-dimensional scene representation can also be compressed using traditional image and video coding tools or alternative NN representation coding tools (not shown).
[0081] To generate the enhancement layer (127), a second set of reference input images (104) for the same scene (but in a second format) is fed to a second unit (110), which generates a second three-dimensional representation (112). For example, according to the scalability criteria and without loss of generality, the two sets of reference images (102, 104) may represent:
[0082] a.PSNR scalability: the first set of images is the same as the second set of images;
[0083] b. Dynamic range scalability: the first set of images is SDR, and the second set of images is HDR;
[0084] c. Color gamut scalability: the first set of images is R.709, and the second set of images is R.2020;
[0085] d. Spatial resolution scalability: the first set of images is 1080p, and the second set of images is 2160p;
[0086] e. Temporal frame rate scalability: 24 frames / second for the first set of images and 48 frames / second for the second set of images;
[0087] In unit 110, for the second 3D scene representation, if the rendered scene is located at the original camera position where the true value image is available, a 1:1 bypass can be used. In other words, by using a 1:1 bypass, the pose {t g} has the advantage of having a true value image by directly using the true value image to generate the residual, rather than using a second model (110), where the output of the second model may still contain artifacts / distortion.
[0088] like Figure 1A As shown, in some embodiments, when there is a spatial and / or temporal misalignment between the base layer and enhancement layer outputs (107 and 112) (such as case d) and case e) discussed above, a reformatter (115) may be required. To achieve scalability of spatial resolution, the reformatter can perform spatial upscaling or downscaling. To achieve scalability of temporal frame rate, the reformatter can drop frames or perform inter-frame interpolation. Such a reformatter can be used in encoders and decoders (see Figure 1B ). In some embodiments, the reformatter may be used in the enhancement layer following the second / enhancement layer representation unit (110).
[0089] Given two scene representations, a residual (122) is generated by a residual generator (120), the residual representing the difference between the two scenes. All residuals from different views are encoded by a neural field (125). The neural network representation of the residual neural field (125) is compressed and transmitted as a neural network residual bitstream output (127).
[0090] On the decoder side, Figure 1B As shown, the decoder receives bitstreams (107) and (127) representing baseline and enhancement information. Note that if the bitstream (107) has been compressed before transmission, it should also be appropriately decompressed in the decoder (not shown). Some decoders may use only the baseline information and ignore any enhancement information. Figure 1B As shown, given a user-specified viewpoint position to render a scene, a base layer unit (130) will reconstruct a rendered baseline view (132). As previously described, depending on the scalability standard, if the decoder will use residual information, the baseline view (132) may need to be processed by a reformatter (115). The enhancement layer bitstream (127) will be decoded along with the user's viewer position input to render the residual (145) generated using the neural field (140). The output of the reformatter will be added to the residual to generate a new improved view (150).
[0091] In embodiments, it may be desirable to reduce the computational complexity of generating the neural field (125) and / or reduce the size of the neural field model, for example by training the neural field 125 using input residuals (122) of lower spatial resolution. This step of reducing the spatial resolution of the residual may be a separate processing unit (not shown) located after the residual generator (120) and before the residual neural field (125), or may be absorbed by the structure of the residual field (125). In the decoder, a spatial upscaling unit may be added after the neural field 140. Alternatively, since the neural field is a continuous functional block, during inference, a higher resolution output may be queried even if the neural residual network was trained using lower resolution grid data. Thus, the residual decoded neural field (140) may absorb spatial interpolation operations without the need for a separate spatial / temporal interpolation module.
[0092] Leveraging the scalability of the PSNR standard
[0093] Figure 1C and Figure 1D Depicted when the scalability criterion is PSNR Figure 1A and Figure 1B A simplified version of Figure 1C and Figure 1D As shown, the reformatter (115) is removed and the encoder is trained based on a single set of reference views and scenes (108).
[0094] exist Figure 1D In the example, given a new viewpoint position t, the basic 3D scene representation The output base image (132) is denoted as And the neural residual model (140) outputs the prediction residual (145) as The final improved rendered image (150) will be a combination of these two images:
[0095]
[0096] In an embodiment, the baseline representation may be based on a multi-view plus depth (MVD) format. In this scenario, the original view image is the input coded image, and the new view image is the image generated by depth based image rendering (DIBR) (see reference [6]).
[0097] In another embodiment, the baseline representation can be based on the MPI representation (see reference [5]). It should be noted that the term "MPI" is typically used to process face-forward scenes. When processing 360-degree videos, the term used is MSI (multi-sphere imaging) (see reference [8]). But the concept is the same. Next, as an example and not limiting, additional details of the MPI representation are provided, assuming that the scene is a face-forward scene where all camera poses are on the same plane. As an example, only image scenes will be considered, but similar concepts can also be applied to video scenes.
[0098] Figure 2A An example embodiment of an encoder for scalable 3D scene representation under the PSNR standard and MPI representation is depicted. Given a reference multi-camera captured image (202), the base layer (207) contains an MPI bitstream. To reduce the burden on the decoder, in unit (205), for each camera position (camera pose or viewpoint) t∈{t g}, you can use the pre-trained NN to transform the image Convert to D-layer MPI format where i=0,...,D-1, where is the i-th texture layer and is the i-th transparent layer. The preprocessing step of converting the MPI of multiple camera poses to a codec suitable for traditional codecs (such as AVC, HEVC, and VVC, etc.) is not shown in this figure. For each new viewpoint position t∈{t n}, a predefined number (usually 4 or 8) of the closest original camera viewpoints are selected for interpolation and the set is denoted as T tEach selected original viewpoint s will be warped to the new viewpoint position t as a new MPI representation: Where i = 0, ..., D-1. Each layer from the current viewpoint position s to the new viewpoint position t The deformation process can be expressed as:
[0099]
[0100] Deformation function T s,t () can be expressed as
[0101]
[0102] Among them, (u s ,v s ) is the pixel coordinate when the posture is s, (u t ,v t ) is the pixel coordinate when the posture is t. K s and K t is the camera intrinsic parameter model for the reference view and the target view. R and t are the camera extrinsic parameter models for rotation and translation. n is the normal vector [0 0 1] T a is the depth σd i The distance to the plane parallel to the front face of the source camera.
[0103] Rendered image I from s to t (s→t) Can be calculated as a deformed texture and an alpha channel:
[0104]
[0105] You can also add all transparent layers together, which will be used during the blending process.
[0106]
[0107] The new viewpoint is obtained by transforming the selected neighborhood T t The weighted combination is given by the following (see reference [9]):
[0108]
[0109] The weight factor is expressed as
[0110]
[0111] Here f represents the focal length of the camera, p s and p t represents the pose of camera s and new viewpoint t.
[0112] In an embodiment, the enhancement layer (227) contains a neural network encoded bitstream that includes NN MLP model parameters (e.g., for model 225). The input to the NN MLP (225) is (x, y, m, n), where (x, y) is the pixel location of the image and (m, n) represents the pose coordinates. The output of the NN MLP is the RGB value for any given (x, y, m, n). On the encoder side, MLP training is required for each NN residual scene. The training residual image (220) can be generated (e.g., in unit 215) using the reference view and the new view as follows:
[0113] - First, a second 3D scene representation algorithm (such as NeRF 210) is applied using the same dataset to obtain the model (with parameter Φ s ). You can use At the original viewpoint position {t g} and the new viewpoint position {t n}. g The rendered image (212) at} is represented as And will be in {t n The rendered image (212) at} is represented as
[0114] -For the original camera pose t g , Right now
[0115] -For new posture n , Right now
[0116] Residual Image(220) It is then used in NN 225 to train the MLP function (The parameter is Φ r ):
[0117]
[0118] The network parameters can be obtained by optimization
[0119]
[0120] This method of generating residual images does not take into account compression artifacts. If compression is taken into account, one or more parameters such as (x, y, m, n, Qp) can be added to the NNMLP input function, where for example Qp is the average quantization parameter used to encode the base layer. Additional parameters can also be used to indicate the quality level. When generating training data, the MPI rendered images can be replaced with compressed MPI rendered images.
[0121] Figure 2B An example embodiment of a corresponding decoder is depicted. Given a baseline input 207, for any given new pose (m, n), a base layer NN (230) decodes the required multiple (e.g., four) camera poses MPI and renders a base layer image (232). For the enhancement layer, a residual image (245) is generated using the trained MLP (240) given the input 227. The base layer (232) and the enhancement layer (245) are then added to generate a final image (250).
[0122] Here is an example of an MLP function using PyTorch code:
[0123] Table 1. Examples of neural fields for residual coding
[0124]
[0125]
[0126] During training, the normalized root mean square error can be used to define the loss function. Other loss functions can also be applied.
[0127]
[0128] For example, in an embodiment, for the training parameters: the learning rate is 1e-3, and Adam optimization (an alternative optimization algorithm to stochastic gradient descent for training deep learning models) is used.
[0129] In an alternative embodiment, the base layer three-dimensional representation may be a baseline NeRF with a smaller model size, and the enhancement layers may be created via a higher-level (or higher-precision) NeRF.
[0130] The input x of each NeRF consists of a spatial position (x, y, z) and a viewpoint direction (θ, φ), and the output is the volume density (σ) and view-dependent emission amplitude (r, g, b) at these coordinates.
[0131] For each viewpoint direction, a parameter Φ is used b The base image for the NeRF rendered model is:
[0132]
[0133] The parameter used is Φ s Smaller models of NeRF can generate higher quality images
[0134]
[0135] The residual (e.g. 220) can be generated using the original view and the new view in the following way:
[0136] -For the original camera pose t g , Right now
[0137] -For new posture n , Right now
[0138] Note: By using 1:1 bypass, you can take advantage of the g} by using the ground-truth image directly to generate the residual, rather than using a larger NeRF model which may still contain artifacts / distortions.
[0139] Neural Residual Network Training can be performed using methods similar to those mentioned for the MPI base layer.
[0140] In another embodiment, a scene-independent NeRF (such as PixelNeRF, see reference [3]) can be used to generate the basic 3D representation. The training process for this scenario is the same as the training process for the scene-dependent NeRF discussed above.
[0141] Messaging considerations
[0142] The proposed method is outside the encoding loop. In an embodiment, the syntax related to the system parameters can be conveyed using metadata such as Supplemental Enhancement Information (SEI) used in the MPEG video coding standard. The syntax can also be carried as part of a video program sequence (VPS), a slice program sequence (SPS), a picture program sequence (PPS), a picture header, a slice header, etc.
[0143] For example, SEI messages may carry information for cameras, base layers, and enhancement layers. Camera information should include camera parameters and camera positions. For the base layer, since the base layer bitstream is the codec bitstream, only additional information that is not carried by the codec bitstream needs to be signaled. This information may include base layer representations such as MPI, MVD, etc. For each input representation, some additional information may be required. For example, for MPI, the number of cameras, the number of MPI layers, and the tiles-assembly may need to be transmitted. For the enhancement layer, syntax elements need to specify NN MLP parameters. NN parameters may be carried by external means or using a neural network represented by an ISO / IEC 15938-17 bitstream. As in the NNPFC SEI (see reference
[10] ), the enhancement layer information may include input and output formatting information.
[0144] Table 2 provides an example of an SEI message used to convey syntax parameters related to the neural field used in the scalable 3D scene representation. To avoid repetition, only the additional information is listed, that is, the information not carried in the NNPFC SEI. The definition of the descriptor information (such as ue(v), u(n), etc.) is the same as in the NNPFC SEI.
[0145] Table 2. Example of SEI messaging for scalable 3D scene representation
[0146]
[0147]
[0148]
[0149] For camera viewpoint information, depending on the setting, for a one-dimensional setting, the multiview acquisition information SEI and multiview viewpoint position SEI in VSEI, HEVC and AVC can be used. For a general setting, the viewport camera parameters SEI and viewport position SEI in visual volume-based video coding (V3C) (see reference
[11] ) and MPEG immersive video (MIV) (see reference
[12] ) can be used. Table 3 below shows an example. Note that the following SEIs: multiview_acquisition_info(), multiview_view_position(), viewport_camera_parameters(), and viewport_position() do not need to be included in the camera_viewport_info() SEI message shown in Table 3. These SEIs can also be sent outside the camera_viewport_info() SEI message. In another embodiment, since the 6DoF (6 degrees of freedom) setting includes a one-dimensional setting, the 6DoF setting case can always be used, but this may require more bits.
[0150] Table 3. Camera viewport SEI message examples
[0151]
[0152] cv_idc specifies the settings for the camera viewpoint information, as shown in the following table.
[0153] cv_idc Camera viewpoint settings 0 One-dimensional level 1 6DoF or Universal Settings
[0154] nnr_purpose_idc specifies the purpose of the neural network residual output. This is a 5-bit switch signaling, where: bit 0 signals PSNR quality enhancement; bit 1 signals dynamic range enhancement; bit 2 signals color gamut enhancement; bit 3 signals spatial resolution enhancement; bit 4 signals temporal frame rate. Each bit can be 0 or 1, and there are 2 5 Combination of.
[0155] nnr_output_pic_width_in_luma_samples specifies the width of the output picture referring to SEI in units of luma samples.
[0156] nnr_output_pic_height_in_luma_samples specifies the height of the output picture referring to the SEI in units of luma samples.
[0157] Note: Using this SEI, the base layer decoded picture resolution may be different from the final output resolution.
[0158] The following syntax elements are used for color space, where the color space that the 3D neural network residual layer system can apply is different from the color space used for the coded layer video sequence (CLVS) layer. For example, the bitstream is in YCbCr color space, while the two-layer system can operate in RGB color space.
[0159] nnr_output_colour_description_present_flag equal to 1 indicates that different combinations of primary colors, transfer characteristics and matrix coefficients of the output picture derived by SEI are specified in the SEI message syntax structure.
[0160] nnr_output_colour_description_present_flag equal to 0 indicates that the combination of colour primaries, transfer characteristics and matrix coefficients of the output picture derived by the SEI is the same as indicated in the VUI parameters for CLVS.
[0161] The semantics of nnr_colour_primaries are the same as specified for the vui_colour_primaries syntax element in clause 7.3, with the following exceptions:
[0162] -nnr_colour_primaries specifies the primaries of the resulting picture after applying the SEI message, instead of the primaries used for CLVS.
[0163] - When nnr_colour_primaries is not present in the SEI message, the value of nnr_colour_primaries is inferred to be equal to vui_colour_primaries.
[0164] The semantics of the nnr_transfer_characteristic are the same as specified in clause 7.3 for the vui_transfer_characteristics syntax element, with the following exceptions:
[0165] -nnr_transfer_characteristics specifies the transfer characteristics of the picture after applying the SEI message, instead of the transfer characteristics used for CLVS.
[0166] - When nnr_transfer_characteristics is not present in the SEI message, the value of nnr_transfer_characteristics is inferred to be equal to vui_transfer_characteristics.
[0167] The semantics of nnr_matrix_coeffs are the same as specified in Section 7.3 for the vui_matrix_coeffs syntax element, with the following exceptions:
[0168] -nnr_matrix_coeffs specifies the matrix coefficients of the picture obtained after applying the SEI message, instead of the matrix coefficients used for CLVS.
[0169] - When nnr_matrix_coeffs is not present in the SEI message, the value of nnr_matrix_coeffs is inferred to be equal to vui_matrix_coeffs.
[0170] - The allowed values for nnr_matrix_coeffs are not restricted by the chroma format of the decoded video pictures, which is indicated by the value of ChromaFormatIdc for the semantics of the VUI parameter.
[0171] nnr_num_cameras_minus1 plus 1 specifies the number of viewport cameras.
[0172] The following is the semantics of the base layer related information.
[0173] nnr_bl_idc specifies the base layer signal input format for 3D scene representation as follows:
[0174] nnr_bl_idc Base layer input format 0 MPI 1 MSI 2 MVC or MVD
[0175] nnr_mpi_layer_minus1 plus 1 specifies the number of MPI layers for the MPI representation.
[0176] nnr_sf_value specifies the scaling factor used for view rendering, in units of 0.001.
[0177] It should be noted that for depth_representation_info(), it does not need to be signaled within the proposed SEI. It can be signaled outside the current SEI.
[0178] The following are the semantics for enhancement layer NNR.
[0179] nnr_mode_idc equal to 0 specifies that the neural network residual is determined by an external device not specified in this specification. nnr_mode_idc equal to 1 specifies that the neural network residual is a neural network represented by the ISO / IEC 15938-17 bitstream contained in this SEI message. nnr_mode_idc equal to 2 specifies that the neural network residual is a neural network identified by the specified tag uniform resource identifier (URI) (nnr_uri_tag[i]) and the neural network information URI (nnr_uri[i]). The value of nnr_mode_idc should range from 0 to 255 (inclusive). Values of nnr_mode_idc greater than 2 are reserved for future ITU-T|ISO / IEC specifications and shall not appear in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification should ignore SEI messages containing reserved values of nnr_mode_idc.
[0180] nnr_input_dimension_minus3 specifies the input signal dimension by adding 3.
[0181]
[0182]
[0183] nnr_position_encoding_freq[i] specifies the input position encoding frequency for the i-th dimension.
[0184] nnr_normalized_weight specifies the weight value, and the weight value is in units of 0.001.
[0185] nnr_abs_normalized_offset specifies the absolute offset value. The absolute offset value is in units of 0.001.
[0186] nnr_sign_normalized_offset specifies the sign of the offset value.
[0187] nnr_sign_normalized_offset=((nnr_sign_normalized_offset>=0)?1:-1)*nnr_abs_normalized_offset
[0188] The residual input x to the neural network is scaled to y in [0 1]: y = (nnr_normalized_weight*x+nnr_normalized_offset)*0.001.
[0189] For boundless scenes, it may also be necessary to specify minimum and maximum values for pose coordinates and / or depth information (for zoom-in / zoom-out effects) to avoid over-rendering.
[0190] In one example, let the pose of the reference camera be (x si ,y si ,z si ), where i is the index of the reference camera. The pose of the reference camera forms a 3D volume in 3D space, whose minimum and maximum values can be calculated as follows:
[0191] x_min=min(x si ); x_max=max(x si );
[0192] y_min=min(y si ); y_max = max(y si );
[0193] z_min=min(z si ); z_max=max(z si );
[0194] When rendering a new viewpoint (x t ,y t ,z t ), the camera pose of the new viewpoint needs to be bounded by the minimum and maximum values of the reference pose in three-dimensional space, that is,
[0195] x tb =min(x_max,max(x_min,x t ,));
[0196] y tb =min(y_max,max(y_min,y t ,));
[0197] z tb =min(z_max,max(z_min,z t ,));
[0198] In another example, a scaling factor may be further applied to the minimum and maximum values of the reference camera pose to adjust the bounded region.
[0199] References
[0200] The entire contents of all references listed herein are incorporated by reference in their entirety.
[0201] [1] Yiheng Xie et al., “Neural Fields in Visual Computing and Beyond,” Eurographics 2022 / CGF, Frontier Technology Reports, Vol. 41(2022), No. 2, 2022.
[0202] [2] Ben Mildenhall et al., “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis,” ECCV 2020, also available at arXiv:2003.08934v2, 5 April 2022.
[0203] [3] Alex Yu et al., “pixelNerF: Neural Radiance Fields from One or Few Images,” CVPR 2021, also available at arXiv:2012.02190v3, May 30, 2021.
[0204] [4] Heiner Kirchhoffer et al., “Overview of the Neural Network Compression and Representation (NNR) Standard,” IEEE Transactions on Circuits and Systems for Video Technology, Vol. 32, No. 5, May 2022, pp. 3203–3216.
[0205] [5] Richard Tucker and Noah Snavely, “Single-view view synthesis with multiplane images,” CVPR 2020.
[0206] [6] “Test Model 11 of 3D-HEVC and MV-HEVC,” JCT3V-K1003, Geneva, Switzerland, February 2015.
[0207] [7] Vincent Sitzmann et al., “Implicit Neural Representations with Periodic Activation Functions,” NeurIPS 2020, also available at arXiv:2006.09661v1, June 17, 2020.
[0208] [8] Benjamin Attal, “MatryODShka: Real-time 6DoF Video View Synthesis using Multi-Sphere Images,” ECCV 2020.
[0209] [9] Ben Mildenhall et al., “Local Light Field Fusion: Practical View Synthesis with Prescriptive Sampling Guidelines,” ACM Transactions on Graphics, vol. 38, no. 4, p. 29, July 2019.
[0210]
[10] Sean McCarthy et al., “Additional SEI messages for VSEI (Draft 2)”, JVET-AA2006v2, 27th JVET Meeting, July 13–22, 2022.
[0211]
[11] ISO / IEC 23090-5, Information technology—Coded Representation of Immersive Media—Part 5: Visual Volumetric Video-based Coding (V3C) and Video-based Point Cloud Compression (V-PCC).
[0212]
[12] ISO / IEC 23090-12, Information technology—Coded representation of immersive media—Part 12: MPEG Immersive video.
[0213]
[13] ISO / IEC 15938-17:2022, “MPEG NNC specification: Information technology—Multimedia content description interface—Part 17: Compression of neural networks for multimedia content description and analysis”.
[0214] Example Computer System Implementation
[0215] Embodiments of the present invention may be implemented by computer systems, systems configured in electronic circuits and components, integrated circuit (IC) devices (such as microcontrollers, field programmable gate arrays (FPGAs) or other configurable or programmable logic devices (PLDs)), discrete time or digital signal processors (DSPs), application specific integrated circuits (ASICs), and / or devices including one or more of such systems, devices, or components. Computers and / or integrated circuits may implement, control, or execute instructions related to scalable three-dimensional scene representations (such as the instructions described herein). Computers and / or integrated circuits may calculate any of the various parameters or values related to the scalable three-dimensional scene representations described herein. Image and video implementations may be implemented by hardware, software, firmware, and various combinations thereof.
[0216] Some embodiments of the present invention include a computer processor that executes software instructions so that the processor executes the method of the present invention. For example, one or more processors in a display, encoder, set-top box, transcoder or similar device can implement the method related to the scalable three-dimensional scene representation described above by executing software instructions in a program memory accessible to the processor. Embodiments of the present invention can also be provided in the form of a program product. The program product can include any non-transient tangible medium that carries a set of computer-readable signals, including instructions that, when executed by a data processor, cause the data processor to execute the method of the present invention. The program product according to the present invention can be any of various non-transient and tangible forms. For example, the program product can include physical media, such as magnetic data storage media including floppy disks, hard disk drives, optical data storage media including CD ROMs, DVDs, electronic data storage media including ROMs, flash RAMs, etc. The computer-readable signals on the program product can be optionally compressed or encrypted. In the content mentioned above about components (such as software modules, processors, components, devices, circuits, etc.), unless otherwise indicated, reference to the component (including reference to "means") should be interpreted as including any component that performs the function of the component as equivalent to the component (for example, functional equivalents), including components that are not structurally equivalent to the disclosed structures that perform the functions in the example embodiments shown in the present invention.
[0217] Equivalents, Extensions, Alternatives, and Miscellaneous
[0218] Thus, example embodiments related to scalable three-dimensional scene representations are described. In the foregoing description, embodiments of the present invention have been described with reference to numerous specific details, which may vary from one implementation to another. Therefore, the sole and exclusive indicator of the content of the present invention and the scope of the invention intended by the applicant is the set of claims finally granted by this application, which is subject to the specific form finally determined by these claims, including any subsequent amendments. Any explicit definitions made by this application for the terms contained in the claims apply to the meanings of the terms used in the claims. Therefore, any limitation, element, property, feature, advantage or attribute not expressly stated in a claim shall not limit the scope of the claim in any way. Accordingly, the description and drawings should be regarded as illustrative rather than restrictive.
Claims
1. A method for generating a scalable three-dimensional scene representation in an encoder, the method comprising: accessing a first set of images in a first format for a scene (102); generating a first three-dimensional scene representation (107) for the scene based on the first set of images; accessing a second set of images in a second format for the scene (104); generating a second three-dimensional scene representation for the scene based on the second set of images (112), wherein the second three-dimensional scene representation is superior to the first three-dimensional scene representation according to one or more quality criteria; Generate an output image residual based on the first three-dimensional scene representation and the second three-dimensional scene representation using a set of original viewpoint positions and a set of new viewpoint positions (122); Using the output image residual to train a residual neural field network (125) to generate a predicted residual image that approximates the output image residual; transmitting the first three-dimensional scene representation (107) for the scene as a base layer; as well as Information about the trained residual neural field network is transferred as an enhancement layer. 2 . The method of claim 1 , further comprising reformatting an output of the first three-dimensional scene representation or the second three-dimensional scene representation before generating the image residual.
3. The method of claim 2, wherein reformatting comprises image upscaling, image downscaling, frame dropping, frame interpolation, or dynamic range / color gamut expansion.
4. The method according to any one of claims 1 to 3, wherein the one or more quality criteria include PSNR scalability, dynamic range scalability, color gamut scalability, spatial resolution scalability and temporal frame rate scalability.
5. The method according to any one of claims 1 to 4, wherein the first set of images is identical to the second set of images.
6. A method according to any one of claims 1 to 4, wherein the first set of images differs from the second set of images in terms of dynamic range or bit depth, color gamut, spatial resolution or frame rate.
7. The method according to any one of claims 1 to 6, wherein the three-dimensional scene representation can be one of a multi-view plus depth (MVD) representation, a multi-planar imaging (MPI) representation, or a neural radiance field (NeRF) neural network representation.
8. The method of claim 5, wherein the first 3D scene representation comprises a first NeRF model, and the second 3D scene representation comprises a second NeRF model, wherein the second NeRF model renders a higher quality image than the first NeRF model, and generating the output image residual comprises: Calculate the first image residual as well as Calculate the second image residual Among them, t g represents the original camera pose, t n represents the new camera pose, and represents images rendered for spatial position (x, y, z) and viewpoint direction (θ, φ) based on the first NeRF model and the second NeRF model, respectively, and represents an image in the first set of images.
9. The method of claim 8, wherein during training, the parameters of the residual neural field network are optimized by Generated, where Φ r * represents the optimized set of parameters of the residual field network, represents the output of the trained residual field network at viewpoint t, denotes the image residual at viewpoint t, and D() denotes the loss function to be minimized during training.
10. A method for generating an output three-dimensional scene in a decoder, the method comprising: Receiving a base layer bitstream (107), the base layer bitstream comprising a first three-dimensional scene representation (107) for a scene; Receiving an enhancement layer bitstream (127), the enhancement layer bitstream including information for reconstructing a trained residual neural field network; Given an observer position: generating a first three-dimensional output of a scene based on the first three-dimensional scene representation (132); generating an image residual (145) using the observer position and the trained residual neural field network; and The first three-dimensional output of the scene and the image residual are combined to generate an enhanced three-dimensional output of the scene.
11. The method of claim 10, further comprising reformatting the first three-dimensional output or the image residual of the scene before combining it.
12. The method of claim 11, wherein the reformatting comprises image enlargement, image reduction, frame dropping or frame interpolation.
13. The method of claim 1, wherein information about the trained residual neural field network includes one or more of the following: a quality parameter (nnr_purpose_idc) specifying the one or more quality criteria, Camera viewport information (viewport_camera_info_present_flag parameters), a first model parameter (nnr_bl_idc) for the first three-dimensional representation model, The number of hidden layers in the residual neural field (this is related to the NN topology and can be carried in the NNR bitstream or externally based on nnr_mode_idc), Input position encoding method (nnr_position_encoding_freq[i]), The activation function (which is related to the NN topology and can be carried in the NNR bitstream or an external device based on nnr_mode_idc, which can be ReLU in Table 1 or SIREN in reference [7], etc.), Parameters related to residual rescaling (nnr_normalized_weight / offset and nnr_sign_normalized_offset) the descriptor of the input coordinate parameter (nnr_input_dimension_minus3), and Descriptors of the output parameters (nnr_colour_primaries, nnr_output_pic_width_in_luma_samples, nnr_output_pic_height_in_luma_samples, etc.), Residual rescaling parameters (nnr_normalized_weight, nnr_abs_normalized_offset and nnr_sign_normalized_offset).
14. The method of claim 13, wherein the information is transmitted as part of a Supplemental Enhancement Information message.
15. The method of claim 1, wherein the residual neural field network (125) is trained in a first spatial resolution using the output image residual, and the method further comprises: The residual neural field network is trained by inputting the output image residual at a second spatial resolution lower than the first spatial resolution and outputting the predicted residual image at the first spatial resolution.
16. A non-transitory computer-readable storage medium having computer-executable instructions stored thereon, the computer-executable instructions for executing the method according to any one of claims 1 to 15 using one or more processors.
17. An apparatus comprising a processor, and the processor is configured to perform the method according to any one of claims 1 to 15.