Nerve radiation field coding and decoding method and device

By fine-tuning an image encoder to align with neural radiance field characteristics, the method effectively compresses NRF feature data, addressing storage inefficiencies and enhancing compression efficiency.

CN120318343APending Publication Date: 2025-07-15ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410060585.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-15
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The prior art cannot effectively compress the characteristic data of the neural radiation field, resulting in a large storage occupancy, and the existing end-to-end image codecs cannot be directly used for the characteristic data compression in the neural radiation field.

Method used

By fine-tuning the feature codec, it is more suitable for the feature data in the neural radiation field, the feature data is constructed and characterized, and the coded feature data is formed to form a code stream, and the fine-tuned feature codec is used to achieve efficient compression.

Benefits of technology

It realizes efficient compression of the neural radiation field, reduces storage usage, and improves coding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318343A_ABST
    Figure CN120318343A_ABST
Patent Text Reader

Abstract

The invention discloses a neural radiation field coding and decoding method and device, and the method comprises the steps: 1) in a coding part, constructing characterization which comprises the fine tuning parameters of a plane grid, a line segment grid, a multilayer perceptron and a feature decoder of a neural radiation field; encoding fine tuning parameters in the neural radiation field and the feature decoder to form a code stream; 2) in a decoding part, decoding the code stream, and constructing a feature decoder; decoding by using the constructed feature decoder to obtain a reconstructed plane grid; decoding the code stream to obtain a reconstructed line segment grid; decoding the code stream to obtain a reconstructed multi-layer perceptron; the reconstructed plane grids, the reconstructed line segment grids and the reconstructed multi-layer perceptron form a reconstructed neural radiation field. According to the invention, the coding performance of the feature codec on the feature data in the nerve radiation field can be enhanced in a targeted manner by finely adjusting the parameters of the feature codec, so that the efficient compression of the nerve radiation field is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video coding and decoding, and particularly relates to a neural radiance field coding and decoding method and device. Background Art:

[0002] A neural radiance field is a three-dimensional scene representation that can generate highly realistic images from arbitrary viewpoints. The idea of a neural radiance field is to represent a three-dimensional scene as a continuous volume function and use a neural network or other data structures that can be learned through backpropagation of gradients to learn this function. The learning of a neural radiance field, which can also be referred to as the training of a neural radiance field, requires multi-view images of the three-dimensional scene as training data. When training a neural radiance field, a loss function needs to be defined, and generally the L2 norm between the real image and the rendered image of the neural radiance field is used to represent it. During the training process, the neural network adjusts its parameters in the direction of minimizing the loss function to minimize the difference between the rendered image of the neural radiance field and the real image. This process usually requires multiple iterations until the loss function reaches a relatively small value, and then the neural radiance field can be considered trained. The trained neural radiance field can then be used to render images from arbitrary viewpoints.

[0003] Similar to traditional visual media such as images and videos, a neural radiance field, as a novel visual media, allows viewers to freely and smoothly change the viewing angle and has extremely high application prospects in scenarios such as immersive free viewpoint videos.

[0004] As a visual media, how to ensure the compression of the media as much as possible while maintaining a certain visual quality is an important issue, which will affect the bandwidth resources consumed during the transmission and distribution of this visual media and also affect the storage occupancy consumed during its storage, directly affecting its actual application prospects. Therefore, how to achieve efficient compression of neural radiance fields is an important research issue with practical application requirements.

[0005] Currently, the mainstream data form for implementing a neural radiance field is feature data plus a multi-layer perceptron (MLP). The neural radiance field implemented with this mainstream data form has advantages such as short training time, fast rendering speed, and high rendering quality, but it has a disadvantage of large storage occupancy, generally ranging from 70MB to 1GB. Usually, in the neural radiance field implemented with this mainstream data form, the feature data occupies more than 99% of the storage occupancy, while the multi-layer perceptron only occupies less than 1% of the storage occupancy. Therefore, it is necessary to compress the feature data.

[0006] In the neural radiance field composed of feature data and MLP, numerous studies have shown that the underlying data structure for storing and representing feature data can take various forms, such as feature voxel grids, multi-scale hash feature grids, three groups of feature plane grids, and feature line segment grids. Compared with the neural radiance field using feature voxel grids or multi-scale hash feature grids as the underlying data structure, the neural radiance field using three groups of plane grids and line segment grids as the underlying data structure of feature data is widely used in academia and industry due to its lower storage occupancy and ease of use.

[0007] Schematic in the accompanying drawings Figure 1 Schematically shows a neural radiance field using three groups of plane grids and line segment grids as the underlying data structure of feature data. When using a neural radiance field with three groups of plane grids and line segment grids as the underlying data structure of feature data, the storage occupancy of the neural radiance field mainly consists of plane grids, line segment grids, and MLP. Among them, the plane grids occupy the vast majority of the storage occupancy of the neural radiance field. Therefore, the plane grids are the main compression targets.

[0008] Masked Wavelet NeRF (see Rho, Daniel, et al. "Masked Wavelet Representation for Compact Neural Radiance Fields." Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2023.) is a research work published in CVPR 2023. It performs efficient compression on the neural radiance field using three groups of plane grids and line segment grids as the underlying data structure of feature data. Inspired by the field of image coding compression, it uses operations such as wavelet transform, quantization, applying masks to filter out invalid information, and entropy coding on the feature data, achieving efficient compression of the neural radiance field composed of feature data and MLP.

[0009] However, this technique fails to consider the constraint on the information entropy of the object to be encoded during the training optimization process, which may lead to certain information redundancy, resulting in a relatively long final coding length. In addition, the wavelet transform used in this technique is a linear transform. Although this transform can achieve efficient compression, there is still room for further improvement in its compression efficiency.

[0010] The latest progress in the field of image coding and compression shows that non-linear transformation is a more efficient transformation method compared to transformations such as wavelet transformation. The end-to-end image codec is an image codec based on non-linear transformation and is a more advanced solution in the current field of image coding and decoding. The end-to-end image codec is a data-driven non-linear transformation learned from a large number of natural images, which can bring better rate-distortion performance.

[0011] The end-to-end image codec is an image codec composed of a neural network. The end-to-end image codec takes the image to be encoded as input and can output a bitstream representing this image and a reconstructed image after encoding and decoding. An advanced end-to-end image codec can achieve the best codec rate-distortion performance in the world currently.

[0012] The end-to-end image codec needs to be trained before it can be used. Specifically, the end-to-end image codec learns the neural network parameter values inside each component by training on an image dataset composed of a large number of images. Once the training is completed, the neural network parameters in the end-to-end image codec will be solidified and deployed to the encoding end and the decoding end. In actual use, the image to be encoded will be sent into the encoding part of the end-to-end image codec at the encoding end, and a bitstream will be obtained through the inference of each sub-component; the bitstream will be transmitted and received by the decoding end, and then sent into the decoding part of the end-to-end image codec, and a reconstructed image will be obtained through the inference of each sub-component.

[0013] However, although the end-to-end image codec can achieve efficient image compression, it cannot be simply and directly used for compressing the feature data in the neural radiance field. Even if the feature data in the neural radiance field is reorganized to form feature mosaic data similar to an image, and then the end-to-end image codec is used to compress the feature mosaic data, efficient compression of the feature data in the neural radiance field cannot be achieved. Summary of the Invention

[0014] To overcome the above defects of the prior art, the present invention proposes a method for encoding and decoding a neural radiance field, which can fine-tune the parameters in the feature codec to make the fine-tuned feature codec more adaptable to the plane grid in the neural radiance field, so as to be able to use the fine-tuned feature codec to achieve efficient compression of the feature data in the neural radiance field.

[0015] The main idea of the present invention is:

[0016] Considering the similarity between the feature data in the image and the neural radiance field, as well as the powerful compression efficiency of advanced image codecs, the present invention conceives that, based on the image codec, with relatively low cost, some modifications and customized fine-tuning can be carried out to obtain a feature codec for the neural radiance field after fine-tuning. Here, the fine-tuning is to fine-tune the parameters in the feature codec according to the inherent characteristics of the feature data in the neural radiance field, so that the feature codec can efficiently compress the feature data in the neural radiance field.

[0017] To this end, the first object of the present invention is to propose a neural radiance field encoding method, which includes:

[0018] 1) Constructing representations, including the plane grid, line segment grid, multi-layer perceptron of the neural radiance field, and the fine-tuning parameters of the feature decoder;

[0019] 2) Encoding the neural radiance field and the fine-tuning parameters in the feature decoder to form a bitstream.

[0020] Preferably, the constructing representations includes:

[0021] By training the neural radiance field, obtaining the required representations, including the plane grid, line segment grid, multi-layer perceptron of the neural radiance field;

[0022] By jointly fine-tuning the feature decoder and the neural radiance field, obtaining the required representations, including the fine-tuning parameters of the feature decoder.

[0023] Preferably, the encoding the neural radiance field and the fine-tuning parameters in the feature decoder includes:

[0024] Encoding the plane grid in the neural radiance field to obtain a plane grid bitstream;

[0025] Encoding the line segment grid in the neural radiance field to obtain a line segment grid bitstream;

[0026] Encoding the multi-layer perceptron in the neural radiance field to obtain a multi-layer perceptron bitstream;

[0027] Encoding the fine-tuning parameters in the feature decoder to obtain a feature decoder parameter bitstream.

[0028] The second object of the present invention is to propose a neural radiance field decoding method, which includes:

[0029] Decoding the bitstream and constructing a feature decoder;

[0030] Using the constructed feature decoder to decode to obtain a reconstructed plane grid;

[0031] Decoding the bitstream to obtain a reconstructed line segment grid;

[0032] Decode the bitstream to obtain a reconstructed multi-layer perceptron;

[0033] The reconstructed planar grid, the reconstructed line segment grid, and the reconstructed multi-layer perceptron form a reconstructed neural radiance field.

[0034] Preferably, when decoding the bitstream to construct a feature decoder, it further includes:

[0035] Decode the bitstream to obtain the fine-tuning parameters of the feature decoder;

[0036] Construct a feature decoder according to the default parameters of the feature decoder and the fine-tuning parameters decoded from the bitstream;

[0037] Use the feature decoder to decode the bitstream to obtain a reconstructed planar grid.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] The present invention is a method for encoding and decoding the feature data of a neural radiance field by using feature encoding and decoding, so as to realize the efficient encoding and decoding of the neural radiance field. This method reuses an image codec, and based on the image codec, makes appropriate modifications to it so that it can become a feature codec for efficiently encoding and decoding the feature data of the neural radiance field. Due to the targeted modification, the encoding performance of the feature data in the neural radiance field is enhanced, thereby realizing the efficient compression of the neural radiance field.

[0040] The present invention enhances the encoding performance of the feature codec for the feature data in the neural radiance field by fine-tuning the parameters of the feature codec, thereby realizing the efficient compression of the neural radiance field. Description of the Drawings

[0041] Figure 1 A neural radiance field using three sets of planar grids and line segment grids as the underlying data structure of feature data in the prior art;

[0042] Figure 2 A schematic structural diagram of the feature codec according to an embodiment of the present invention;

[0043] Figure 3 A schematic structural diagram of the encoding part of the feature codec according to an embodiment of the present invention;

[0044] Figure 4 A schematic structural diagram of the decoding part of the feature codec according to an embodiment of the present invention. Detailed Embodiments

[0045] To further understand the present invention, the preferred embodiments of the present invention will be described below in conjunction with the embodiments and the accompanying drawings. However, it should be understood that these descriptions are only for further explaining the features and advantages of the present invention, rather than limiting the claims of the present invention.

[0046] Embodiment 1

[0047] The embodiment of the present invention proposes a neural radiance field encoding method, and its main work includes the following two aspects:

[0048] 1) By using an efficient feature codec to encode and decode the feature data in the neural radiance field, the overall storage occupancy of the neural radiance field can be significantly reduced. At the same time, the present invention also compresses other data in the neural radiance field except for the feature data, which can further reduce the overall storage occupancy of the neural radiance field.

[0049] 2) The feature codec proposed in the present invention is customized for each three-dimensional scene. Here, the customization is reflected in that all the parameters of the encoder in the feature codec and a part of the parameters of the decoder part are customized for the three-dimensional scene to be compressed. Since there are some customized parameters related to the content of the three-dimensional scene in the decoder, these parameters also need to be transmitted to the client, that is, these customized parameters also need to be encoded and compressed.

[0050] It should be noted that before performing the above compression in this embodiment, it is assumed that this embodiment has already obtained in advance the data used to train this neural radiance field, and these data need to be used in the subsequent steps. Among them, the data in the neural radiance field includes the feature data and MLP in the neural radiance field. In this embodiment, the feature data selects plane grids and line grids.

[0051] The neural radiance field encoding method based on feature encoding and decoding in this embodiment includes the following steps:

[0052] S1. Construct representations, including the plane grids, line grids, multi-layer perceptron of the neural radiance field, and the fine-tuning parameters of the feature decoder.

[0053] Constructing representations means constructing the neural radiance field and simultaneously constructing the fine-tuning parameters of the feature decoder. These contents are all objects to be encoded at the encoding end, and are collectively referred to as representations.

[0054] The representations can be divided into two types of objects. One type is to construct the neural radiance field, and the other type is to construct the fine-tuning parameters of the feature decoder. Among them, the neural radiance field is the main object to be compressed, and the fine-tuning parameters of the feature decoder are additional encoding objects introduced to more efficiently encode and decode the neural radiance field. Therefore, constructing the neural radiance field and constructing the fine-tuning parameters of the feature decoder are described in two steps.

[0055] S11. Obtain the required representations by training the neural radiance field, including the plane grid, line segment grid, and multi-layer perceptron of the neural radiance field.

[0056] The operation performed in this step is to construct the neural radiance field. The essence of constructing the neural radiance field is to obtain the required representations by training the neural radiance field, including the plane grid, line segment grid, and multi-layer perceptron of the neural radiance field. The training of the neural radiance field is prior art and will not be elaborated here.

[0057] S12. Obtain the required representations by jointly fine-tuning the feature codec and the neural radiance field, including the fine-tuning parameters of the feature decoder.

[0058] After constructing the neural radiance field, the next step is to construct the fine-tuning parameters of the feature decoder. To construct the fine-tuning parameters of the feature decoder, the feature codec needs to be jointly fine-tuned with the neural radiance field obtained in the previous step. The purpose of jointly fine-tuning the feature codec and the neural radiance field is to adjust the parameters in the feature codec so that it can encode and decode the neural radiance field with higher encoding efficiency.

[0059] It should be noted that in the fine-tuning of this step, the parameters involved in the fine-tuning are: the parameters of the last multi-layer convolutional layer in the last layer of the feature decoder, all the parameters in the feature encoder, and the parameters in the MLP of the neural radiance field.

[0060] Here is an explanation for selecting the above parameters for fine-tuning: Theoretically, the best means of fine-tuning the model is to perform full-parameter fine-tuning on all the parameters in the feature decoder and the neural radiance field. However, if full-parameter fine-tuning is directly performed, the parameters of the decoding end model in the feature decoder will change significantly, and all the parameters of the decoding end model need to be transmitted, and the cost of transmitting the parameters will be very high. Therefore, the present invention chooses to only fine-tune the parameters of the last layer or multi-layer convolutional layer in the feature decoder, rather than fine-tuning the parameters of other layers except the last layer or multi-layer convolutional layer in the feature decoder.

[0061] It should be noted that here, one layer or multiple layers are selected depending on the degree of fine-tuning desired and the transmission cost of the fine-tuning parameters. At least the parameters of the last convolutional layer need to be transmitted. If you hope to achieve as full a fine-tuning as possible and do not mind the additional overhead brought by transmitting more layers, then on the basis of transmitting the last layer, additional layers can be transmitted. In addition, since the feature encoder is not used during decoding and does not need to be transmitted, the present invention chooses to fine-tune all the parameters in the feature encoder. Finally, since the MLP in the neural radiance field must be transmitted and fine-tuning it will not bring additional transmission overhead, therefore, the present invention chooses to fine-tune the parameters in the MLP of the neural radiance field to make it more adaptable to the codec.

[0062] The specific fine-tuning process of this embodiment is as follows:

[0063] S121. Design the objective function used in fine-tuning.

[0064] In the fine-tuning process of this embodiment, a loss function needs to be defined to guide the direction of parameter update in the fine-tuning process. In this embodiment, the objective function mainly consists of two terms. One is the reconstruction loss of rendered pixels, and the other is the bitrate loss.

[0065] The reconstruction loss of rendered pixels is a common and necessary loss in neural radiance field training, which is used to measure the difference between the color of the pixels rendered by the neural radiance field and the color of the pixels at the corresponding positions in the real image.

[0066] The bitrate loss is a common and necessary loss in the training of the feature codec, which is used to measure the lowest bitrate that can be achieved under the current feature codec.

[0067] The objective function during fine-tuning will be the weighted sum between the reconstruction loss of rendered pixels and the bitrate loss.

[0068] S122. According to the designed objective function, fine-tune the parameters in the feature codec and the pre-trained neural radiance field that need to participate in fine-tuning.

[0069] S123. Calculate the value of the objective function by simulating the encoding and decoding process and the rendering process.

[0070] There are two parts that need to be calculated in the objective function. One is the reconstruction loss of rendered pixels, and the other is the bitrate loss. These can be calculated by simulating the encoding and decoding process and the rendering process, which can also be called the forward propagation process.

[0071] The calculation of the reconstruction loss of rendered pixels is as follows: The neural radiance field first undergoes the simulated encoding and decoding process to obtain the reconstructed neural radiance field. Following the rendering process in the neural radiance field, the reconstructed neural radiance field is rendered to obtain the color of the pixels rendered by the reconstructed neural radiance field. Finally, the mean square error function between the color of the pixels rendered by the reconstructed neural radiance field and the color of the pixels at the corresponding positions in the real image is the reconstruction loss of rendered pixels.

[0072] The calculation of the bitrate loss is as follows: During the simulated encoding and decoding process of the neural radiance field, the planar grid is sent into the codec. The feature codec will output the probabilities of each element in the hidden layer features and the hyperprior features. These probabilities will be used to calculate the entropy of this batch of hidden layer features and hyperprior features, and the calculated entropy is the bitrate loss.

[0073] S124: After obtaining the objective function value, use the gradient descent algorithm to optimize the model parameters designated for update.

[0074] This process is the process of backpropagating the gradient of the objective function. The direction of backpropagation is exactly opposite to the direction of calculating the objective function. The gradients propagated to each model are used to update the model parameters. Here, the model parameters to be updated are the parameters of the last layer or multiple layers in the feature decoder, all the parameters in the feature encoder, and the parameters in the MLP of the neural radiance field.

[0075] S125: Repeat the above two steps of S123 and S124 until the performance of the system composed of the feature encoder-decoder and the neural radiance field tends to be stable, and the optimization is completed.

[0076] To determine whether the performance of the system tends to be stable, it can be to see whether the relative change rate of the objective function is lower than a certain threshold. If it is lower than a certain threshold, it means that the model is basically converged and the optimization ends. It can also be to limit the optimization time and directly end the optimization when the time ends, and use the model parameters that obtain the optimal objective function as the model parameters finally fine-tuned.

[0077] It should be particularly noted that the feature encoder-decoder used later refers to the feature encoder-decoder adjusted by fine-tuning in this step.

[0078] In addition, it should be particularly noted that different from the encoding and decoding methods of the image codec in the prior art, since the parameters of the last layer or multiple layers in the decoder of the feature encoder-decoder in this embodiment need to be fine-tuned, the parameters of the last layer or multiple layers in the decoder need to be encoded.

[0079] S2: Encode the fine-tuning parameters in the neural radiance field and the feature encoder-decoder to form a bitstream.

[0080] After finishing the fine-tuning step of S2, the fine-tuning parameters in the neural radiance field and the feature encoder-decoder can be encoded to form a bitstream. Among them, what needs to be encoded in the neural radiance field are: feature data and MLP. The fine-tuning parameters in the feature encoder-decoder refer to the parameters of the last layer or multiple layers in the decoder of the feature encoder-decoder, and these parameters are hereinafter referred to as feature decoder fine-tuning parameters.

[0081] Encode the feature data to obtain a feature data bitstream, encode the MLP to obtain an MLP bitstream, and encode the fine-tuning parameters in the feature decoder to obtain a feature decoder fine-tuning parameter bitstream.

[0082] Among them, in this embodiment, encoding the feature data includes the following steps:

[0083] S21: Obtain a plane grid bitstream for the plane grid in the neural radiance field.

[0084] The purpose of this step is to encode the feature data in the neural radiance field. The feature data can include a planar grid and a line segment grid. Since the encoding of the planar grid requires a feature codec, the feature codec will be introduced here.

[0085] The feature codec includes the following parts: an encoder, an image decoder, a quantizer, an arithmetic encoder, an arithmetic decoder, and a hidden layer feature probability estimator. The hidden layer feature probability estimator further includes the following parts: a hyperprior encoder, a hyperprior decoder, a quantizer, an implicit probability estimator, an arithmetic encoder, and an arithmetic decoder.

[0086] After introducing the structure of the feature codec, we will introduce how to use the encoder part of the feature codec to encode the planar grid to form a planar grid bitstream.

[0087] Generally speaking, the specific operations of the encoder part during encoding include that after three groups of planar grids are successively fed into the encoding part of the feature codec, three corresponding hidden layer feature bitstreams and hyperprior feature bitstreams will be formed. Below, we will introduce the process of a single group of planar grids.

[0088] 1) The planar grid to be encoded is fed into the feature encoder as input, and the feature encoder outputs hidden layer features.

[0089] After the feature encoder outputs the hidden layer features, they are divided into two branches. The purpose of branch 1 is to estimate the probability of the hidden layer features, and the purpose of branch 2 is to obtain the object to be encoded, that is, the quantized hidden layer features.

[0090] Below, we will first introduce that steps 2)-7) are the steps of branch 1, starting from the hidden layer features being fed into the hyperprior encoder until the hyperprior decoder outputs the probability statistical characteristic parameters of the hidden layer features.

[0091] 2) The hidden layer features are fed into the hyperprior encoder, and hyperprior features are output;

[0092] 3) The hyperprior features are fed into the quantizer, and quantized hyperprior features are output;

[0093] 4) The quantized hyperprior features are fed into the implicit probability estimator, and the probability of each element in the hyperprior features is output.

[0094] 5) The quantized hyperprior features and the probability of each element in the hyperprior features are fed into the arithmetic encoder. After arithmetic encoding, a hyperprior bitstream is output.

[0095] 6) The hyperprior bitstream is fed into the arithmetic decoder and, combined with the probability of each element in the hyperprior features, decodes the hyperprior features are output.

[0096] 7) The decoded hyper-prior features are fed into the hyper-prior decoder, and the probability statistical characteristic parameters of the hidden layer features are output.

[0097] Up to here, branch 1 ends. Next, the steps of branch 2 are introduced. Before the convergence of branch 1 and branch 2, branch 2 only has step 8).

[0098] 8) The hidden layer features are fed into the quantizer, and the quantized hidden layer features are output.

[0099] Up to here, branch 2 ends. Next, the outputs of branch 1 and branch 2 will jointly participate in the operations of the next module.

[0100] 9) The quantized hidden layer features, combined with the probability statistical characteristic parameters of the hidden layer features obtained in step 7), are fed into the arithmetic encoder for arithmetic coding to obtain the hidden layer feature bitstream.

[0101] Up to here, the encoding of the plane grid by the encoder part ends, and the plane grid bitstream composed of the hyper-prior bitstream and the hidden layer feature bitstream is output.

[0102] S22: Encode the line grid in the neural radiance field to obtain the line grid bitstream.

[0103] For the encoding of the line grid, due to its small data volume and simple data representation form, quantization and entropy coding means are directly used for encoding to form the line grid bitstream. Among them, the entropy coding method can be Huffman coding.

[0104] So far, the encoding operations of the feature data in the neural radiance field have all been completed, and the feature data bitstream is output. The feature data bitstream can be composed of the plane grid bitstream and the line grid bitstream.

[0105] S23: Encode the MLP in the neural radiance field to obtain the MLP bitstream.

[0106] The data volume of the MLP in the neural radiance field is small, so simple encoding means can be used for encoding. In this embodiment, quantization and entropy coding means are used for encoding to form the MLP bitstream.

[0107] S24: Encode the fine-tuning parameters in the feature decoder to obtain the feature decoder parameter bitstream.

[0108] Since the parameters of the last layer or multiple layers of the decoder part in the codec are fine-tuned in step S1 of this embodiment, in order to ensure correct decoding, the fine-tuning parameters of the decoder part also need to be encoded to form the feature decoder parameter bitstream and transmitted to the decoding end. In this embodiment, the fine-tuning parameters of the decoder part can be encoded by using quantization and entropy coding means.

[0109] So far, step S2 is completed, and the fine-tuning parameters in the neural radiance field and the feature codec are encoded to form a bitstream.

[0110] Embodiment 2

[0111] This embodiment proposes a neural radiance field decoding method based on feature encoding and decoding. It uses the following steps to parse and decode the bitstream to obtain the reconstructed neural radiance field.

[0112] The bitstream output in step S2 of the above Embodiment 1 contains the necessary information for reconstructing the neural radiance field at the decoding end. Therefore, this embodiment uses the bitstream output in step S2 of Embodiment 1 to obtain the reconstructed neural radiance field. Specifically, it includes the following steps:

[0113] P1. Parse and decode the bitstream, construct a feature decoder, and use the feature decoder to decode to obtain the reconstructed planar grid.

[0114] The goal of step P1 is to use the feature decoder to decode the planar grid bitstream to obtain the reconstructed planar grid. Before the feature decoder is officially used for decoding, it needs to obtain its parameters. Since the parameters of the feature decoder include default parameters and fine-tuning parameters, and the fine-tuning parameters need to be obtained by decoding the fine-tuning parameter bitstream. Therefore, step P1 is further refined into the following steps:

[0115] P11. Parse the bitstream to obtain the fine-tuning parameters of the feature decoder.

[0116] In this step, the fine-tuning parameter bitstream of the feature decoder is decoded to obtain the fine-tuning parameters of the feature decoder. The obtained fine-tuning parameters will be used for subsequent operations of constructing the feature decoder.

[0117] P12. Based on the default parameters of the feature decoder and the fine-tuning parameters obtained from the bitstream, construct the feature decoder.

[0118] In this step, based on the default parameters of the part except the last few layers in the feature decoder, and the fine-tuning parameters of the part belonging to the last few layers obtained in P11, a feature decoder that can correctly decode the planar grid bitstream is constructed.

[0119] P13. Use the feature decoder to parse and decode the bitstream to obtain the reconstructed planar grid;

[0120] When the correct feature decoder is constructed in the previous step, it can be used to decode the planar grid bitstream to obtain the reconstructed planar grid.

[0121] The specific decoding process of the feature decoder in this embodiment is as follows:

[0122] 1) The super-a priori feature code stream is sent to the arithmetic decoder, and after arithmetic decoding, the decoded super-a priori features are obtained;

[0123] 2) The decoded super-prior features are sent to the super-prior decoder, which outputs the probability statistical characteristic parameters of the hidden layer features, which will be used for arithmetic decoding of the hidden layer feature code stream later;

[0124] 3) The hidden layer feature code stream, combined with the probability statistical characteristic parameters of the hidden layer features, is sent to the arithmetic decoder, and after arithmetic decoding, the decoded hidden layer features are obtained;

[0125] 4) The decoded hidden features are fed into the feature decoder and the reconstructed plane grid is output.

[0126] At this point, the decoding process of the feature decoder ends and outputs the reconstructed plane raster.

[0127] P2. Parse and decode the bitstream, and obtain the reconstructed line segment grid and reconstructed MLP.

[0128] After obtaining the reconstructed plane grid, the goal of this step is to obtain the reconstructed line segment grid and MLP to obtain the necessary elements for reconstructing the neural radiation field. Specifically, the line segment grid code stream is first entropy decoded to obtain the reconstructed line segment grid. The entropy decoding method here must correspond to the entropy coding method in step S2, such as Huffman decoding. Next, the MLP is entropy decoded to obtain the reconstructed MLP. The entropy decoding method here must correspond to the entropy coding method in step S2, such as Huffman decoding. At this point, step P2 ends, and the reconstructed line segment grid and reconstructed MLP are obtained.

[0129] P3, the reconstructed plane grid, the reconstructed line segment grid, and the reconstructed MLP constitute the reconstructed neural radiation field.

[0130] After the reconstructed plane grid, the reconstructed line segment grid and the reconstructed MLP are obtained in the above steps, this step organizes them according to the data structure organization form of the neural radiation field to form a reconstructed neural radiation field.

[0131] At this point, the decoding end has obtained the reconstructed neural radiation field through the decoding operation, which includes the reconstructed plane grid, the reconstructed line segment grid, and the reconstructed MLP.

[0132] The reconstructed neural radiation field obtained by the decoding operation at the decoding end can be used for rendering at the decoding end.

[0133] This embodiment also proposes a neural radiation field encoding and decoding device. The specific implementation methods of the various components or programs of the device can be referred to the steps of the aforementioned neural radiation field encoding and decoding method, which are the same and will not be repeated here.

[0134] The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A method for encoding a neural radiance field, characterized in that Comprising: 1) Construction representations, including the fine-tuning parameters of the planar grid, line segment grid, multi-layer perceptron, and feature decoder of the neural radiance field; 2) Encoding the fine-tuning parameters in the neural radiance field and the feature decoder to form a bitstream.

2. The neural radiance field encoding method according to claim 1, wherein The construction representations include: By training the neural radiance field, obtaining the required representations, including the planar grid, line segment grid, and multi-layer perceptron of the neural radiance field; By jointly fine-tuning the feature decoder and the neural radiance field, obtaining the required representations, including the fine-tuning parameters of the feature decoder.

3. The neural radiance field encoding method according to claim 1, wherein The encoding of the fine-tuning parameters in the neural radiance field and the feature decoder includes: Encoding the planar grid in the neural radiance field to obtain a planar grid bitstream; Encoding the line segment grid in the neural radiance field to obtain a line segment grid bitstream; Encoding the multi-layer perceptron in the neural radiance field to obtain a multi-layer perceptron bitstream; Encoding the fine-tuning parameters in the feature decoder to obtain a feature decoder parameter bitstream.

4. A neural radiance field decoding method, characterized in that Comprising: Decoding the bitstream to construct a feature decoder; Using the constructed feature decoder to decode to obtain a reconstructed planar grid; Decoding the bitstream to obtain a reconstructed line segment grid; Decoding the bitstream to obtain a reconstructed multi-layer perceptron; The reconstructed planar grid, reconstructed line segment grid, and reconstructed multi-layer perceptron form a reconstructed neural radiance field.

5. The neural radiance field decoding method according to claim 4, wherein The decoding of the bitstream to construct a feature decoder further includes: Decoding the bitstream to obtain the fine-tuning parameters of the feature decoder; Constructing a feature decoder according to the default parameters of the feature decoder and the fine-tuning parameters decoded from the bitstream; Using the feature decoder to decode the bitstream to obtain a reconstructed planar grid.