Method for compressing and transferring dynamic three-dimensional spatial information

The dynamic volume coding method addresses the inefficiencies in encoding dynamic 3D videos by employing tensor decomposition and 2D video codecs, resulting in enhanced encoding efficiency and video quality, as well as reduced bandwidth and energy consumption.

WO2025127479A1PCT designated stage expired Publication Date: 2025-06-19HYUNDAI MOTOR CO LTD +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/018513
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-20
Filing Date
2024-11-21
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing methods for encoding dynamic 3D videos face challenges in efficiently representing and compressing dynamic 3D spatial information, leading to decreased encoding efficiency and video quality.

Method used

A dynamic volume coding method and device that utilizes tensor decomposition to divide dynamic 3D spatial information into low-dimensional matrices and vectors, and then compresses these using a 2D video codec, allowing for efficient reconstruction of dynamic 3D videos at the decoder side.

Benefits of technology

This approach improves the encoding efficiency and quality of dynamic 3D videos, reduces network bandwidth requirements, and decreases energy consumption for video playback devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024018513_19062025_PF_FP_ABST
    Figure KR2024018513_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are a method and device for dynamic volume coding using the compression of dynamic three-dimensional spatial information. In the present embodiment, a volume decoding device decodes tensor features representing an I-volume, and weights of a tensor generation model. The volume decoding device reconstructs the tensor generation model by using the weights of the tensor generation model. The volume decoding device reconstructs tensor features of a B-volume on the basis of two nearest decoded I-volumes by using the tensor generation model. The volume decoding device generates an image of the current point in time on the basis of the tensor features of the I-volume when the current time corresponds to decoding of the I-volume. When the current time corresponds to decoding of the B-volume, the volume decoding device generates an image of the current point in time on the basis of the reconstructed tensor features of the B-volume.
Need to check novelty before this filing date? Find Prior Art

Description

Method for compressing and transmitting dynamic three-dimensional spatial information

[0001] The present disclosure relates to a dynamic volume coding method and device utilizing compression of dynamic three-dimensional spatial information.

[0002] The content described below merely provides background information related to the present invention and does not constitute prior art.

[0003] Recently, research on implicit neural network representation models, which represent various data, including images, using neural network structures, has been active. To represent conventional 3D (dimensional) video, multi-viewpoint and depth video are acquired simultaneously, or light-field video or 360-degree multi-viewpoint video are acquired. Conventional 3D video representation methods typically use explicit representations, which use RGB pixel values ​​for each pixel location. To replace these explicit representations, implicit neural representations (INRs) are being introduced. These models represent functions that output (r, g, b) values ​​and volume density from five-dimensional inputs consisting of spatial pixel locations (x, y, z) and viewpoint directions (θ, φ) using neural networks. Compared to explicit representations, implicit neural representations can instantly restore (r, g, b) values ​​based on coordinates at any point in time, offering practical advantages over existing decoders.

[0004] Among implicit neural representation methods, NeRF (Neural Radiance Fields) is a representative model used for 3D video synthesis. Static NeRF, which represents static space, performs preprocessing by applying positional embedding to the aforementioned 5-dimensional input vector and inputs the preprocessed input vector into a simple multilayer perceptron (MLP) to output pixel values ​​and volume density. However, expressing the spatial information of the target 3D image based on a small network size has limitations. Furthermore, expressing temporally varying spatial information can be difficult. Dynamic NeRF technology has been proposed to overcome the aforementioned problems, but the network size can become large and encoding efficiency can decrease. Furthermore, the development of a dedicated codec for the dynamic NeRF model may be impractical due to compatibility with existing codec technologies.

[0005] Therefore, in order to improve the efficiency of dynamic 3D video encoding representation and enhance the quality of dynamic 3D video, a method for efficiently reconstructing spatial information of dynamic 3D video needs to be considered.

[0006] The present disclosure aims to provide a dynamic volume coding method and device that divides dynamic 3D spatial information into low-dimensional matrices and vectors based on a tensor decomposition technique, compresses the divided information based on a 2D video codec, and generates a video of an arbitrary point in time at a decoder side based on the restored dynamic 3D spatial information.

[0007] According to an embodiment of the present disclosure, a method for restoring a multi-view video, performed by a volume decoding device, comprises the steps of: decoding tensor features representing an I volume, weights of a tensor generation model, and side information from a bitstream, wherein the tensor features of the I volume are represented by decomposed tensors, the tensor generation model is used to reconstruct a B volume, and the side information includes information related to a quantization parameter and information related to a structure of the tensor generation model; reconstructing the tensor generation model so as to conform to the information related to the structure of the tensor generation model using the weights of the tensor generation model and the information related to the quantization parameter; reconstructing tensor features of the B volume based on two nearest decoded I volumes using the tensor generation model; obtaining a current time and a current view point; And, if the current time corresponds to decoding of the I volume, a step of generating an image of the current time based on tensor features of the I volume; If the current time corresponds to decoding of the B volume, a step of generating an image of the current time based on tensor features of the reconstructed B volume is provided.

[0008] According to another embodiment of the present disclosure, a method for encoding a multi-view video, performed by a volume encoding device, is provided, comprising: obtaining a current time and a current view point; when the current time corresponds to encoding of an I volume, obtaining an I volume and decomposing the I volume to generate low-dimensional tensors representing tensor features of the I volume; when the current time corresponds to encoding of a B volume, reconstructing tensor features of the B volume based on two nearest decomposed I volumes using a tensor generation model; encoding tensor features of the I volume; and encoding weights of the tensor generation model.

[0009] According to another embodiment of the present disclosure, a method for providing video data to a video decoding device is provided, comprising: encoding the video data into a bitstream; and transmitting the bitstream to the video decoding device, wherein the encoding the video data comprises: obtaining a current time and a current viewpoint; when the current time corresponds to encoding of an I volume, obtaining an I volume and decomposing the I volume to generate low-dimensional tensors representing tensor features of the I volume; when the current time corresponds to encoding of a B volume, reconstructing tensor features of the B volume based on two nearest decomposed I volumes using a tensor generation model; encoding tensor features of the I volume; and encoding weights of the tensor generation model.

[0010] As described above, according to the present embodiment, by providing a dynamic volume coding method and device that divides dynamic 3D spatial information into low-dimensional matrices and vectors based on a tensor decomposition technique, compresses the divided information based on a 2D video codec, and generates a video of an arbitrary point in time on the decoder side based on the restored dynamic 3D spatial information, it is possible to improve dynamic 3D video encoding efficiency and dynamic 3D video quality.

[0011] In addition, according to the present embodiment, by providing a dynamic volume coding method and device for efficiently compressing and restoring dynamic 3D spatial information, it is possible to reduce the burden on the network based on bit rate reduction in various contents such as UHD (Ultra High Definition) video, game broadcasting, 360-degree video streaming, VR / AR (Virtual Reality / Augmented Reality) video, online lectures, etc., and to reduce energy consumption for video playback capable devices.

[0012] FIG. 1 is a block diagram illustrating a dynamic volume encoding device based on multi-view video coding according to one embodiment of the present disclosure.

[0013] FIG. 2 is a block diagram illustrating a multi-view video coding-based dynamic volume decoding device according to one embodiment of the present disclosure.

[0014] Figure 3 is an example diagram showing a forward network as a deep learning-based neural network.

[0015] FIG. 4 is an exemplary diagram showing the operation of a convolutional layer according to one embodiment of the present disclosure.

[0016] Figure 5 is an example diagram showing a SISR (Single Image Super Resolution) network.

[0017] Figure 6 is an example diagram showing a residual block used in SISR.

[0018] Figure 7 is an example diagram showing coordinate systems related to extraction of arbitrary points in NeRF (Neural Radiance Fields).

[0019] FIG. 8 is an exemplary diagram showing a GOV (Group of Volume) according to one embodiment of the present disclosure.

[0020] FIG. 9a and FIG. 9b are exemplary diagrams showing tensor decomposition according to one embodiment of the present disclosure.

[0021] FIG. 10 is an exemplary diagram illustrating a tensor generation model according to one embodiment of the present disclosure.

[0022] FIG. 11 is an exemplary diagram illustrating a tensor generation model according to another embodiment of the present disclosure.

[0023] FIGS. 12A to 12C are exemplary diagrams showing the creation of a B volume according to one embodiment of the present disclosure.

[0024] FIG. 13 is an exemplary diagram showing an encoding pipeline of a neural network compression method according to one embodiment of the present disclosure.

[0025] FIG. 14 is an exemplary diagram showing a decoding pipeline of a neural network compression method according to one embodiment of the present disclosure.

[0026] FIG. 15 is a flowchart illustrating a method for a volume encoding device to encode multi-view video according to an embodiment of the present disclosure.

[0027] FIG. 16 is a flowchart illustrating a method for a volume decoding device to restore multi-view video according to an embodiment of the present disclosure.

[0028] Hereinafter, embodiments of the present invention will be described in detail with reference to exemplary drawings. When designating components in each drawing, it should be noted that, where possible, identical components are given the same reference numerals, even if they appear in different drawings. Furthermore, in describing the present embodiments, detailed descriptions of related known structures or functions will be omitted if they are deemed to obscure the gist of the present embodiments.

[0029] The present embodiment provides a dynamic volume coding method and device that divides dynamic 3D spatial information into low-dimensional matrices and vectors based on a tensor decomposition technique, compresses the divided information based on a 2D video codec, and generates a video of an arbitrary point in time on the decoder side based on dynamic 3D spatial information restored from the compressed information.

[0030] FIG. 1 is a block diagram illustrating a dynamic volume encoding device based on multi-view video coding according to one embodiment of the present disclosure.

[0031] In the example of FIG. 1, a dynamic volume encoding device based on multi-view video coding (hereinafter, referred to as a 'volume encoding device') generates (i.e., trains) a multi-view video coding model to learn to generate dynamic 3D video from multiple views, compresses the learned multi-view video coding model to generate a bitstream, and transmits the generated bitstream. The volume encoding device includes a multi-view video encoding unit (110) including a multi-view video coding model and a feature compressor (120). Here, the components included in the volume encoding device according to the present disclosure are not necessarily limited thereto. The volume encoding device may additionally include a training unit (not shown) for training the multi-view video coding model, or may be implemented in a form linked with an external training unit.

[0032] Each component of the volume encoding device may be implemented in hardware, software, or a combination of hardware and software. Furthermore, the functions of each component may be implemented in software, with a microprocessor executing the software functions corresponding to each component.

[0033] The volume encoding device can store a bitstream of encoded video data on a non-transitory recording medium or transmit it to a multi-view video coding-based dynamic volume decoding device using a communication network.

[0034] FIG. 2 is a block diagram illustrating a multi-view video coding-based dynamic volume decoding device according to one embodiment of the present disclosure.

[0035] In the example of FIG. 2, a dynamic volume decoding device based on multi-view video coding (hereinafter, referred to as a "volume decoding device") decompresses a bitstream to restore a multi-view video coding model, and generates a video from multiple views using the restored multi-view video coding model. The volume decoding device includes a feature decompressor (210) and a multi-view video decoding unit (220) including the restored multi-view video coding model.

[0036] Similar to the volume encoding device illustrated in Fig. 1, each component of the volume decoding device may be implemented in hardware, software, or a combination of hardware and software. Furthermore, the functions of each component may be implemented in software, with a microprocessor configured to execute the software functions corresponding to each component.

[0037] Hereinafter, volume and 3D space are used interchangeably, and dynamic volume, 4D space-time, 3D video, and multi-view video are used interchangeably.

[0038] Below, before describing the components illustrated in FIGS. 1 and 2, the deep learning-related element technologies used in the present disclosure are described.

[0039] I-1. Multilayer Perceptron (MLP)

[0040] A multilayer perceptron (MLP) is composed of multiple neurons and edges connecting the neurons, as shown in the example in Figure 3. In addition to an input layer and an output layer, an MLP may include one or more hidden layers. Each hidden layer includes one or more hidden units (or hidden nodes). Additionally, different weights may be assigned to each edge. In addition, an activation function may be used in the process of propagating output values ​​from one layer to the next. Representative activation functions include the sigmoid function, the tangent hyperbolic function, or the Relu (Rectified Linear Unit) function.

[0041] As illustrated in the example of Fig. 3, a two-layer MLP with a forward network structure passes an input vector x through an intermediate hidden unit to provide an output vector y. In the aforementioned forward network structure, any node value located in the output layer can be expressed as in mathematical equation 1.

[0042]

[0043] In mathematical equation 1, class is a weight matrix, corresponding to the edges connecting the input vectors and hidden units, and the edges connecting the hidden units and the output vectors. w ji (1) and w kj(2) are each class Indicates the component of y k represents the components of the output vector y. x0 and z0 represent the bias units of the input layer and the hidden layer, respectively. Therefore, D is the dimension of the input vector x, M represents the number of hidden units, and K is the dimension of the output vector y. In addition, h and σ are the activation functions applied to the hidden layer and the output layer, respectively. Meanwhile, the activation function is not necessarily applied to the output of the hidden layer and the output layer. If the activation function is not used, mathematical expression 1 becomes class It includes an operation corresponding to the weight matrix multiplication between the two.

[0044] The process of calculating weights based on training data and labels is called training. Training can typically be performed using the stochastic gradient descent (SGD) algorithm, which utilizes backpropagation. The process of calculating outputs in a feed-forward manner using the weights calculated through training is called inference or testing.

[0045] Hereinafter, multilayer neural networks, MLPs, or forward-facing networks can be used interchangeably.

[0046] I-2 CNN(Convolutional Neural Network)

[0047] CNN, a neural network comprised of multiple convolutional layers and pooling layers, is a deep learning technique known to be ideal for image processing. Convolutional layers extract feature maps (also known as "features" interchangeably) using multiple kernels or filters. The kernel coefficients that make up the filters are parameters determined during the learning process.

[0048] Among the convolutional layers of CNN, the front layer closer to the input extracts feature maps that respond to simple, low-level image features such as lines, points, or surfaces, while the back layer closer to the output extracts feature maps that respond to higher-level features such as textures and object parts.

[0049] FIG. 4 is an exemplary diagram showing the operation of a convolutional layer according to one embodiment of the present disclosure.

[0050] A convolutional layer generates a feature map from an input image using a convolution operation. The example in Fig. 4 illustrates a kernel (or filter) with a kernel size of 3×3. The kernel size is also referred to as the kernel size or filter size. The kernel has kernel parameters, also called weights. The kernel illustrated in Fig. 4 has a total of nine kernel parameters. The kernel parameters are initially set to arbitrary values, and their values ​​can be updated based on training.

[0051] The convolution layer performs convolution operations using blocks of the same size as the kernel size in the input image. At this time, the blocks of the same size as the kernel size in the input image are referred to as windows.

[0052] When filtering an input image in raster scan order, the window movement size is called the stride. In the example of Fig. 4, the stride is 1. If the stride is set to 2, the convolution operation is performed by spacing the window by 2 samples, and as a result, the width and height of the feature map become half the width and height of the input image.

[0053] As mentioned above, a single convolutional layer can include multiple filters. The number of filters or kernels is called a channel. In other words, the number of channels is equal to the number of filters. Furthermore, the number of filters determines the dimensionality of the feature map.

[0054] Padding refers to a method of expanding input data by filling the area around it with a specific value before performing a convolution operation. Padding is primarily used to adjust the spatial size of the output data. The padding value can be determined by hyperparameters, but zero-padding is commonly used. Without padding, the spatial size of the output data decreases with each convolutional layer, potentially causing boundary information to disappear. Therefore, padding is used to prevent this problem. Specifically, padding can be used to match the spatial sizes of the output data from the convolutional layer to the input data.

[0055] The deconvolution layer performs the opposite operation to the convolution layer. It generates the desired data image from the input feature map as output.

[0056] The pooling layer performs pooling, a process of subsampling the feature map generated by the convolutional layer. The pooling layer uses a 2×2 window to select samples so that the output is half the width and height of the input. In other words, the pooling layer is used to reduce the size of the input image or input feature map by condensing a 2×2 region into a single sample.

[0057] The opposite concept of a pooling layer is defined as an unpooling layer. Unpooling layers, in contrast to pooling layers, function to expand the dimensionality and are primarily used after deconvolution layers.

[0058] The convolutional encoder-decoder architecture is a network structure composed of pairs of convolutional layers and deconvolutional layers. The convolutional encoder consists of a convolutional layer and a pooling layer, and outputs a feature map (or feature vector) from the input image. The final output vector of the convolutional encoder is also referred to as a latent vector. The convolutional decoder consists of a deconvolutional layer and an unpooling layer, and generates an output image from the feature map or latent vector.

[0059] The inputs and outputs of a convolutional encoder-decoder can be configured in various ways depending on the application and network purpose. For example, the inputs and outputs can be optical flow maps, saliency maps, image frames, etc.

[0060] Figure 5 is an example diagram showing an SISR network.

[0061] One example of CNN application is Single Image Super Resolution (SISR). The SISR network generates a high-resolution image from a low-resolution input image. As illustrated in Figure 5, the SISR network may include multiple convolutional layers. Each convolutional layer includes an activation function, such as the Rectified Linear Unit (ReLU). The parameters of the SISR network can be trained so that the resulting SR (Super Resolution) image approximates the Ground Truth (GT).

[0062] SR methods using CNN can improve SR performance by increasing the depth (e.g., increasing the number of convolutional layers). To overcome the overfitting problem that may occur in learning due to the increase in depth, a residual block that can perform skip connection and residual learning can be used in the SISR network. The residual block, as illustrated in Fig. 6, is a block that stores input features x l In addition to the path that applies the convolution operation, it includes a skip path. In addition, the residual block outputs x l+1 When generating, a path or skip path for applying convolution operations can also be selected based on learning efficiency. In the example of Fig. 6, the residual block includes a BN (Batch Normalization) layer.

[0063] For example, Enhanced Deep Residual Networks (EDSR) improves network performance by continuously connecting residual blocks to increase their depth. Another example is Accurate Image Super-Resolution Using Very Deep Convolutional Networks (VDSR), a CNN model based on the Visual Geometry Group (VGG) network. It uses residual learning, a method that adds residual frames to the final output. VDSR adds residual signals to the very end of the network, thereby adding them to the input signal.

[0064] I-3. Implicit Neural Representation Model

[0065] Recently, there has been active research on implicit neural representation (INR) models, which represent various data, including images, using neural network structures. Conventional video representation methods use explicit representations, which represent RGB pixel values ​​at each pixel location. To replace this explicit representation, implicit neural representation models are being introduced, which represent functions that generate (r,g,b) values ​​from (x,y) coordinates (pixel locations) in an image using neural networks. Compared to explicit representations, implicit neural representation models can be used regardless of image resolution.

[0066] As an example of 3D video coding that utilizes an implicit neural network representation model, there is NeRF (Neural Radiance Fields) technology.

[0067] By learning from images captured from multiple viewpoints of a multi-view video, NeRF synthesizes an image of an object viewed from any arbitrary position and orientation. After receiving a 5-dimensional coordinate consisting of a 3D position (x, y, z) that samples ray information in the direction the camera is looking and a viewing direction (θ, φ), NeRF outputs the RGB values ​​(r, g, b) and volume density σ corresponding to the input coordinates. Here, the 3D position (x, y, z) is the position in the world coordinate system, and θ and φ that constitute the viewing direction can represent the azimuth and altitude, respectively.

[0068] Compared to existing implicit neural network representations, NeRF offers the following improvements to better represent high-resolution, complex scenes. NeRF embeds positional information using positional encoding, and then uses the embedded index as input, as shown in Equation 2.

[0069]

[0070] As shown in the example in Equation 2, each spatial coordinate can be mapped to a 2L-dimensional vector γ(p). Instead of directly inputting the 3D input coordinates, each spatial coordinate can be mapped to a 2L-dimensional vector γ(p) using the sin and cos functions. By utilizing the embedded spatial index in training the NeRF model, the NeRF model can better predict data containing high-frequency variations.

[0071] NeRF is an implicit neural network representation model that uses a multi-layer perceptron (MLP) network containing multiple fully-connected layers (FC). NeRF represents an arbitrary scene by outputting RGB values ​​(r, g, b) and volume density σ corresponding to high-dimensional input coordinates. Then, the color C(r) of an arbitrary ray r passing through the scene can be rendered according to a volume rendering technique based on the RGB values ​​(r, g, b) and volume density. Since C(r) is differentiable, it is used for training NeRF together with GT (Ground Truth). At this time, an image that captures world coordinates (x, y, z) along the view direction (θ, φ) is used as GT for each 5-dimensional coordinate.

[0072] Meanwhile, to improve rendering efficiency, NeRF includes two networks, a coarse network and a fine network, instead of using a single MLP. NeRF uses stratified sampling to achieve N c After sampling the positions of the dogs, N c The output of the coarse network is calculated at each location. To generate more samples to apply to the fine network from the output of the coarse network, N c Composite color C based on the coarse network output of the dog c (r) is expressed as in mathematical formula 3.

[0073]

[0074] In mathematical expression 3, c i is N c The color of the coarse network calculated from the dog's location. w iis a value that depends on the volume density and the distance between sampled locations as a corresponding weight. Using inverse transform sampling, N with high weights from the aforementioned composite color f After sampling the location of the dog, NeRF c +N f Calculate the output of the fine network at the location of the dog, and color C based on the output of the calculated fine network. f (r) can be rendered. At this time, NeRF can be trained using the MSE (Mean Square Error) loss function between the color and GT according to the output of the coarse network and the fine network.

[0075] As mentioned above, NeRF outputs a rendered result at an arbitrary point in time by using a learning process to reconstruct all pixels in the image into world coordinates. Therefore, a process is needed to transform all points projected to a point in the image coordinate system (x, y) into world coordinates. Therefore, the input of NeRF includes the camera parameters required for projection into the world coordinate system along with the index in the world coordinate system.

[0076] The process of extracting camera parameters is to convert pixels in an image into a world coordinate system [U, V, W] through a coordinate system [X, Y, Z] with a focal length Z = -1, as shown in the example of Fig. 7. At this time, prior information about the camera is essential, and the prior information, the camera parameters, includes internal parameters and external parameters. The internal camera parameters include the focal length, the principal point on the image plane, and the skew coefficient. Depending on the internal camera parameters, the object can be projected into the camera coordinate system [X, Y, Z]. In order to convert the information projected into the camera coordinate system into the world coordinate system [U, V, W], translation information about where the camera is located in the world coordinate system and rotation information about where it is looking are required. The external camera parameters include the aforementioned translation information and rotation information. When capturing an image, a matrix [R|t], which is a combination of a rotation matrix R and a translation vector t, is used to calculate the projection from the world coordinate system to the image coordinate system. Therefore, the camera extrinsic parameters can be extracted by calculating the inverse matrix of the combination matrix [R|t].

[0077] Meanwhile, the 5-dimensional coordinates, which are the inputs of NeRF, can be derived from the extracted camera parameters. As described above, the camera parameters can be derived based on the captured 2D images and the positions in the world coordinate system corresponding to the 2D images. When capturing images for GT use, the internal camera parameters can be fixed. When the internal camera parameters are used while being fixed, the 5-dimensional coordinates can be derived from the external camera parameters.

[0078] Alternatively, in addition to using the camera parameters used when shooting the video, a method such as colmap can be used to extract feature points from re-view images and reconstruct information about the position and direction of the camera within the world coordinate system.

[0079] As another example, grid-based static NeRF is a hybrid NeRF that uses explicit representations in addition to implicit neural network representations. Grid-based static methods use a voxel grid for rapid reconstruction of 3D volumes. However, voxel-based methods require large amounts of memory because they must store features (e.g., color and volume density) per voxel. To improve memory efficiency, methods that decompose the component parts of the feature grid or feature tensor are used. Commonly used decomposition methods are vector-matrix (VM) decomposition or canonical polyadic (CP) decomposition, which can be viewed as generalizations of matrix singular value decomposition (SVD).

[0080] Grid-based static methods, similar to the aforementioned NeRF, explicitly train tensors representing color and volume density, and calculate the color of a target pixel by integrating the colors and volume densities constituting the target viewpoint. At this time, the volume density can be used without additional steps. To enhance the global characteristics, implicit representations can be added to the color by inputting the color to a small-sized MLP. Furthermore, the viewpoint direction can be reflected in the color by inputting the output of the MLP and the viewpoint direction to a shader. The shader can be implemented as an MLP or a Spherical Harmonics (SH) function.

[0081] Grid-based static NeRF has the advantage of reducing training time compared to static NeRF, but still has the disadvantage of not being able to handle dynamic volumes.

[0082] As another example, dynamic NeRF (D-nerf) adds a time dimension to the aforementioned NeRF. Conventional NeRF learns spatial information from static images without considering time, which leads to large errors for moving images. To consider time, dynamic NeRF takes (x, y, z, t) as input and trains a neural network that implicitly represents 4D space-time. The implicit neural network includes two modules. The first module calculates the difference (δx, δy, δz) between the standard space at a reference time (e.g., t=0) and the space at time t. Similar to conventional NeRF, the second module generates RGB values ​​(r, g, b) and volume density from the five-dimensional information of (x+δx, y+δy, z+δz) and the viewpoint direction (θ, φ). Afterwards, the color of the target pixel can be calculated according to a volume rendering technique based on the RGB values ​​(r, g, b) and the volume density.

[0083] Dynamic NeRFs, similar to static NeRFs, have the disadvantage of taking a long time to train.

[0084] As another example, grid-based dynamic NeRF is a hybrid dynamic NeRF that uses explicit representations in addition to implicit neural network representations. Grid-based dynamic methods use a voxel grid for rapid reconstruction of time-sensitive volumes, i.e., 4D space-time. For example, a 4D spatiotemporal grid is decomposed into six feature planes, each containing a pair of coordinate axes (e.g., XY, ZT). Here, the features explicitly represent color and volume density. By projecting a 4D space-time point onto each feature plane and fusing the projections onto the six planes, a feature vector can be reconstructed for the 4D space-time point. Grid-based dynamic methods train six planes representing color and volume density, and calculate the color of a target pixel by integrating the colors and volume densities constituting the target viewpoint. In this case, the volume density can be used without additional steps. To enhance global features, an implicit representation can be added to color by inputting color into a small-sized MLP. Additionally, by inputting the output of the MLP and the viewpoint direction into the shader, the viewpoint direction can be reflected in the color. The shader can be implemented as an MLP or SH function.

[0085] Grid-based dynamic NeRF has the advantage of reducing training time compared to dynamic NeRF, but has the disadvantage of not properly utilizing the correlation between volumes that make up 4D space-time.

[0086] Below, we describe a dynamic volume coding method that can efficiently utilize the correlation between volumes compared to grid-based dynamic NeRF while reducing training time compared to dynamic NeRF.

[0087] Hereinafter, the multi-view video coding model and the multi-view model can be used interchangeably.

[0088] Below, feature maps and features can be used interchangeably.

[0089] Hereinafter, volume density and density may be used interchangeably.

[0090] Hereinafter, voxels represent elements that constitute a three-dimensional grid, and pixels represent elements that constitute a two-dimensional plane.

[0091] II. Embodiments according to the present disclosure

[0092] The volume encoding device illustrated in FIG. 1 trains a multi-view video coding model in a multi-view video encoding unit (110) using a video having multiple views, thereby enabling the multi-view model to generate dynamic 3D videos of various views. At this time, information indicating 4D spatiotemporal coordinates (x, y, z, t) and viewpoint directions (θ, φ) are used as inputs of the multi-view model, and (r, g, b) values ​​corresponding to target pixels are used as GT (Ground Truth). Here, t represents time corresponding to a volume in the dynamic 3D video, and (x, y, z) represents 3D coordinates in the world coordinate system. As described above, a transformation between "(x, y, z), (θ, φ)" and the target pixel in the image plane can be defined based on camera parameters.

[0093] The multi-view model according to the present disclosure can include both explicit and implicit neural network representations of 4D space-time. The training unit uses the difference between the output of the multi-view model and the GT to update features representing the explicit representation and parameters representing the implicit neural network representation, thereby enabling the multi-view model to learn how to generate images from various viewpoints and times. Upon completion of training, the multi-view model can express 4D spatiotemporal information embedded in dynamic 3D video based on the features and parameters.

[0094] The feature compressor (120) compresses the features and parameters of the trained multi-view model to generate a bitstream. The volume encoding device transmits the generated bitstream to the volume decoding device. The volume encoding device may include, as additional information, structural information of the multi-view model, quantization parameters, camera parameters, and information about motion compensation.

[0095] Below, a multi-view model that utilizes predictive coding used in existing video compression methods is described.

[0096] FIG. 8 is an exemplary diagram showing a Group of Volume (GOV) according to one embodiment of the present disclosure.

[0097] The multi-view model uses GOV (Group of Volumes) to represent and encode 4D spatiotemporal information. GOV includes I (Intra) volumes and B (Bidirectional) volumes. By using I and B volumes, time t can be processed among the 6-dimensional input. As shown in Figure 8, the I volume can represent static 3D spatial information for that time without reference. The B volume can be generated by referencing two adjacent I volumes.

[0098] Information related to motion compensation as additional information may be information related to the structure of the GOV described above. Information related to the structure of the GOV may include the size of the GOV, the location of the I volume (i.e., time information of the I volume), etc.

[0099] Since the B volume references two adjacent I volumes, the volume encoding device can first encode the I volumes within the GOV and then encode the B volumes.

[0100] In terms of encoding, the multi-view model includes an I volume decomposer and a B volume generator, as illustrated in Fig. 1.

[0101] The I volume has a larger amount of information compared to the B volume. By adjusting the number of I volumes or increasing or decreasing the grid resolution / channel count, the size of the features and the corresponding bitstream size can be determined. Furthermore, a trade-off regarding the resiliency of the reconstructed volume can be adjusted. For example, increasing the number of I volumes increases the size of the bitstream and enhances its resiliency.

[0102] As an example, the training unit can learn I volumes and B volumes for various GOV structures and determine a GOV structure that provides optimal performance. The volume encoding unit can transmit information related to the structure of the GOV to the volume decoding unit as additional information related to motion compensation.

[0103] FIG. 9a and FIG. 9b are exemplary diagrams illustrating tensor decomposition according to one embodiment of the present disclosure.

[0104] Since the I-volume does not reference other volumes, it contains tensors that can form a voxel grid of a single, complete 3D space. To improve memory efficiency, as shown in Figures 9a and 9b, the I-volume decomposer can reduce the grid of the I-volume based on a conventional tensor decomposition method and explicitly represent the I-volume using the reduced grid.

[0105] As an example of tensor decomposition, a method for reducing a feature map of an I volume based on vector-matrix (VM) decomposition is described using the example of Fig. 9a.

[0106] The I-volume decomposer can decompose the feature map of the I-volume into a linear combination of vectors and matrices using VM decomposition, as illustrated in Fig. 9a. Here, the vector corresponds to a rank 1 tensor, and the matrix corresponds to a rank 2 tensor. The feature map of the I-volume, i.e., the grid, is assumed to have a size of n1×n2×n3. That is, the feature map of the I-volume can be approximated based on a combination of vectors / matrices P1+P2+P3 (where P1, P2, and P3 are natural numbers), as shown in Equation 4.

[0107]

[0108] In Equation 4, F represents the feature map of the I volume, and the operator o represents the outer product. v x,i (1≤i≤P1), v y,i (1≤i≤P2), v z,i (1≤i≤P3) represents vectors with dimensions n1, n2, and n3, respectively. M yz,i (1≤i≤P1), M xz,i (1≤i≤P2), M xy,i (1≤i≤P3) represents matrices with dimensions n2×n3, n1×n3, and n1×n2, respectively.

[0109] The volume encoding device is M yz,i (1≤i≤P1), M xz,i (1≤i≤P2), M xy,i (1≤i≤P3) matrices can be encoded. In addition, the volume encoding device concatenates rank 1 tensors of the same size, as shown in mathematical expression 5, to generate, for example, rank 2 tensors A, B, and C, and then encodes the rank 2 tensors.

[0110]

[0111] In mathematical expression 5, A, B, and C can represent rank 2 tensors having sizes n1×P1, n2×P2, and n3×P3, respectively.

[0112] According to VM decomposition, the data size n1×n2×n3 of the original feature map is reduced to ‘(n1+ n2×n3)P1+ (n2+ n1×n3)P2+ (n3+ n1×n2)P3’.

[0113] As an example, when VM decomposition is applied to the color and density of the I volume, the matrices P1+P2+P3, A, B, and C according to Equation 5, can constitute the channels of the I volume.

[0114] As another embodiment of tensor decomposition, we describe a method for reducing the feature map of an I volume based on Canonical Polyadic (CP) decomposition using the example of Fig. 9b.

[0115] The I-volume decomposer can decompose the feature map of the I-volume into rank-1 tensors, i.e., linear combinations of vectors, using CP decomposition, as illustrated in Fig. 9b. Here, the feature map of the I-volume is assumed to have a size of n1×n2×n3. That is, the feature map of the I-volume can be approximated based on P (where P is a natural number) rank-1 tensors, as shown in Equation 6.

[0116]

[0117] In Equation 6, F represents the feature map of volume I, and the operator o represents the outer product. v x,i , v y,i , v z,i (1≤i≤P) represents a rank 1 tensor with sizes n1, n2, and n3, respectively.

[0118] The volume encoding device, as shown in mathematical expression 7, combines rank 1 tensors of the same size to generate, for example, A, B, and C, and then encodes them.

[0119]

[0120] In mathematical expression 7, A, B, and C can represent rank 2 tensors having sizes of n1×P, n2×P, and n3×P, respectively.

[0121] According to CP decomposition, the data size n1×n2×n3 of the original feature map is reduced to ‘(n1+ n2+ n3)P’.

[0122] As an example, when CP decomposition is applied to the color and density of the I volume, A, B, and C according to Equation 7 can constitute the channels of the I volume.

[0123] The I-volume determines the length of the GOV. The I-volume can be located anywhere in the dynamic volume stream and can be used for keyframe search. Similar to the grid-based static method, the multi-view model explicitly trains tensors representing colors and tensors representing densities. Here, the tensors represent decomposed tensors as described above. The multi-view model computes colors and densities at voxels corresponding to the target viewpoint based on the decomposed tensors of the I-volume, and integrates the computed colors and densities to compute the color of the target pixel (i.e., render the video or image corresponding to the target viewpoint). At this time, the density can be used without additional steps. To enhance the global characteristics, an implicit representation can be added to the color by inputting the color to a small-sized MLP. Furthermore, the viewpoint direction can be reflected in the color by inputting the output of the MLP and the viewpoint direction to a shader. The shader can be implemented as an MLP or a Spherical Harmonics (SH) function.

[0124] As an example, tensor decomposition can be performed on each I volume within the aforementioned GOV.

[0125] The B volume generator generates the B volume based on bidirectional prediction by referencing the two most adjacent I volumes at the current time. That is, the B volume generator utilizes the I volume that is temporally preceding and the I volume that is temporally succeeding the current B volume. Including more B volumes reduces the capacity of the entire dynamic volume, but may degrade the image quality of the rendered image compared to the I volume. Based on a trainable neural network, tensors representing the B volume (tensors for color and volume density) can be generated using information from existing decomposed I volumes. Here, the tensors represent decomposed tensors. The B volume generator can reconstruct the B volume based on the generated tensors.

[0126] Using the reconstructed B volume, the multi-view model calculates the colors and densities of the B volume at voxels corresponding to the target viewpoint, and integrates the calculated colors and densities to calculate the color of the target pixel (i.e., renders the video or image corresponding to the target viewpoint). At this time, the densities of the B volume can be used without any additional steps. To enhance the global characteristics, an implicit representation can be added to the color by inputting the colors of the B volume to a small-sized MLP. In addition, the viewpoint direction can be reflected in the color by inputting the output of the MLP and the viewpoint direction to the shader. The shader can be implemented as an MLP or a SH function.

[0127] B volume uses a fixed tensor and a generated tensor for one or more axes. Here, the generated tensor can exist on an axis orthogonal to the fixed tensor. That is, a volume suitable for a viewpoint can be generated based on the generated tensor and the fixed tensor. The generated tensor refers to a tensor generated for a target viewpoint. The tensor can be generated by the aforementioned neural network, i.e., the B volume tensor generation model (hereinafter, "tensor generation model").

[0128] A fixed tensor refers to a tensor that is not a generated tensor. A single, learnable tensor shared across all viewpoints can be used as a fixed tensor. Alternatively, a tensor corresponding to an axis orthogonal to the generated tensor of the B volume in an adjacent I volume can be used as a fixed tensor. Therefore, a tensor with the same axis as the fixed tensor of the B volume in an adjacent I volume can be used as a fixed tensor.

[0129] A tensor generative model is used to infer a generative tensor for spatial information at arbitrary time. The tensor generative model takes as input two tensors with the same shape along the same axis of adjacent I-volumes. For example, two vectors with the same axis or two planar tensors with the same axis are used. The encoder and decoder of the tensor generative model can be composed of trainable neural networks, such as multi-layer profiling (MLP) or convolutional neural networks (CNN). A neural network with the same parameters can be used for each tensor for color and density, or different neural networks can be used for color and density. A trainable neural network consists of one or more of the following modules:

[0130] A learnable neural network-based encoder / decoder is a neural network composed of MLP layers, convolutional layers, and deformable convolutional layers, and may include N layers in succession. An encoder or decoder may be configured using only one of the MLP layers, convolutional layers, and deformable convolutional layers, or a combination of two or more thereof. Here, in the deformable convolutional layer, the convolutional kernel is transformed into various forms by adding a 2D offset to the grid of the convolutional kernel.

[0131] The tensor-generating model accumulates temporal information channel-wise on features passed through the encoder, and then passes the accumulated temporal information to the decoder. The temporal information represents the coordinate values ​​of a corresponding time in a time-normalized video sequence. To achieve channel-wise accumulation, the temporal information is generated as a matrix with the same height and width as the features. Using the aforementioned method, a continuous B-volume tensor can be generated for time.

[0132] The tensor-generating model accumulates viewpoint information channel-wise on the features passed through the encoder, and then passes the accumulated viewpoint information to the decoder. Viewpoint information represents the coordinate values ​​of that viewpoint in a normalized video sequence. To accumulate viewpoint information channel-wise, the viewpoint information is generated as a matrix with the same height and width as the features. Using the aforementioned method, a continuous B-volume tensor for each viewpoint can be generated.

[0133] A tensor-generating model can generate flows for input tensors using features passed through the encoder. By warping the input tensor or the output features (of the encoder) using the generated flows, the tensor-generating model can capture motion information between input tensors.

[0134] Figures 10 and 11 illustrate the tensor generation model described above. In Figure 10, the tensor generation model uses vectors as input, and in Figure 11, tensor planes as input. In Figures 10 and 11, the temporal coordinate is a matrix representing time information or point-in-time information. In Figure 11, the tensor generation model generates a flow for the input tensor.

[0135] As an example, as a tensor generation model, a single model can be used commonly for the B volumes within the aforementioned GOVs. As another example, as a tensor generation model, a single model can be used commonly for multiple GOVs.

[0136] The volume encoder can train the structure of various tensor generation models and transmit the structure that provides optimal performance to the volume decoder.

[0137] FIGS. 12A to 12C are exemplary diagrams showing the creation of a B volume according to one embodiment of the present disclosure.

[0138] The B volume generator uses, as described above, fixed tensors for one or more axes of the I volume, and generated tensors for axes orthogonal to the fixed tensors. The tensor generation model generates a generated tensor for a target time using one or more tensors of adjacent I volumes, as described above, wherein the tensors of the adjacent I volumes are orthogonal to the fixed tensors. As an example, in FIG. 12a, the B volume is reconstructed using fixed XY, YZ, XZ planes, and generated Z, X, Y vectors orthogonal to the fixed planes.

[0139] When the VM decomposition technique is used as a tensor decomposition method, a plane or a vector can be fixed. When the CP decomposition technique is used, one or more vectors corresponding to each axis can be fixed. In FIGS. 12A to 12C, the B volume is reconstructed based on the generated tensor. In FIGS. 12A and 12B, the B volume is reconstructed based on the VM decomposition technique, and in FIG. 12C, the B volume is reconstructed based on the CP decomposition technique. In FIG. 12A, the B volume for the target time is reconstructed based on the cross product between the fixed tensor planes (XY, YZ, XZ) obtained from the I volume and the generated tensor vectors. In FIG. 12B, the B volume for the target time is reconstructed based on the cross product between the fixed tensor vectors (X, Y, Z) obtained from the I volume and the generated tensor planes. In FIG. 12C, the B volume for the target time is reconstructed based on the cross product between the fixed tensor vectors obtained from the I volume and the generated tensor vectors.

[0140] As described above, the multi-view model calculates colors and densities of the B volume at voxels corresponding to the target viewpoint based on the reconstructed B volume, and calculates the color of the target pixel by integrating the calculated colors and densities (i.e., renders a video or image corresponding to the target viewpoint).

[0141] The volume decoding device illustrated in FIG. 2 decodes features and parameters of a multi-view model from a bitstream using a feature decompressor (210). Here, the features of the multi-view model may represent decomposed tensor features of the I volume. The parameters of the multi-view model may represent parameters of a tensor generation model in relation to a tensor generation model that constitutes a generation tensor of the B volume.

[0142] The volume decoding device decodes information about the structure of the multi-view model, quantization parameters, camera parameters, and motion compensation as additional information. For example, the volume decoding device can decode information related to the structure of the GOV as additional information related to motion compensation.

[0143] The multi-view video decoding unit (220) generates a video from multiple views using the reconstructed multi-view model. As in the volume encoding device, by utilizing the GOV including the I volume and the B volume, the multi-view model can process the time of dynamic 3D video. After acquiring an arbitrary viewpoint j, the volume decoding device can input the arbitrary viewpoint into the multi-view model to generate a 2D image corresponding to the arbitrary viewpoint. By connecting the 2D images corresponding to the I volume and the B volume, the volume decoding device can generate a 2D video corresponding to the arbitrary viewpoint.

[0144] At this time, any point in time must exist within the input range. For example, if any point in time does not exist within the input range, the volume decoding device may not provide an image corresponding to the point in time. As described above, based on the camera parameters, a transformation between "(x, y, z), (θ, φ)" and a pixel in the image plane can be defined. "(x, y, z), (θ, φ)" represents the input of the multi-view model excluding t.

[0145] In terms of decoding, the multi-view model includes an I volume renderer and a B volume generator, as illustrated in Fig. 2.

[0146] The I-Volume renderer utilizes the reconstructed I-Volume tensor features (e.g., tensors representing color and density). Here, the reconstructed I-Volume tensor features may be decomposed tensor features. Based on the reconstructed I-Volume tensor features, the I-Volume renderer computes colors and densities at voxels corresponding to any viewpoint, and integrates the computed colors and densities to compute the color of the target pixel (i.e., renders a video or image corresponding to the target viewpoint). At this time, the density can be used without an additional step. To enhance the global characteristics, an implicit representation can be added to the color by inputting the color to a small-sized MLP. In addition, the viewpoint direction can be reflected in the color by inputting the output of the MLP and the viewpoint direction to the shader. The shader can be implemented as an MLP or a SH function.

[0147] In terms of decoding, the operation of the B volume generator is similar to that of the B volume generator in the encoding aspect described above. The B volume generator generates the B volume based on bidirectional prediction by referencing the two nearest reconstructed I volumes at the current time. The B volume generator constructs a generation tensor using the reconstructed tensor generation model, and reconstructs tensors representing the B volume (tensors for color and volume density) using the generation tensor and the fixed tensor. Here, the reconstructed B volume tensor features may be decomposed tensor features. Based on the reconstructed B volume tensor features, the B volume generator calculates colors and densities at voxels corresponding to any time point, and integrates the calculated colors and densities to calculate the color of the target pixel (i.e., renders the video or image corresponding to the target time point). At this time, the density of the B volume can be used without any additional steps. To enhance the global characteristics, an implicit representation can be added to the color by inputting the color of the B volume into a small-sized MLP. Additionally, by inputting the output of the MLP and the viewpoint direction into the shader, the viewpoint direction can be reflected in the color. The shader can be implemented as an MLP or SH function.

[0148] The information transmitted between the volume encoding device and the volume decoding device may include, in relation to the multi-view model, tensor features (color and volume density) constituting the I volume, parameters (i.e., weights) of the tensor generation model in the B volume generator, and weights of the MLP utilized for rendering. The volume encoding device may transmit additional information to the volume encoding device to support the decoding operation.

[0149] Below, we describe a method for compressing the tensor features that constitute the I volume.

[0150] As an example, tensor features can be compressed using existing 2D video codecs.

[0151] The encoding process for the I volume is tensor feature packing. The encoding process spatially packs planar feature maps of dimension 1×C×H×W into a plane of dimension 1×1×pH×pW. C represents the number of channels of the planar feature map, and H and W represent the resolution of the grid. pH and pW represent the height and width after the plane packing. As an example, if the color and density planes of the XY, YZ, and ZX planes are generated as spaces of dimension [48, 12, 12] and [16, 4, 4], respectively, the encoding process can set pW and pH as follows. When C is 48, pW = 8×W, pH = 6×H can be set. When C is 16, pW = 4×W, pH = 4×H can be set. When C is 12, pW = 4×W, pH = 3×H can be set. If C is 4, pW = 2×W, ​​pH = 2×H can be set.

[0152] The encoding process converts the packed features into YUV format by quantizing them into m bits (where m is an arbitrary positive integer). At this time, the minimum and maximum values ​​can be signaled to the decoding process.

[0153] The encoding process can generate a bitstream by encoding YUV converted into planar features in an All intra, Random access, or Low delay configuration according to the existing image encoding method.

[0154] The decoding process for the I volume decodes the bitstream. The decoding process converts the YUV format into a tensor type, and dequantizes the converted tensor features into 32-bit floating point (float32) using the passed minimum and maximum values. The decoding process unpacks the spatially packed 1×1×pH×pW dimensional plane into 1×C×H×W dimensional planar feature maps.

[0155] As another example, tensor features can be compressed using a Learning Image Compression (LIC) model.

[0156] The encoding process performs feature packing in the same way as the tensor feature packing for the 2D video codec described above. The dimensions of the color and density plane feature maps are R N×C×H×W Here, N is 3 and represents the XY, YZ, and ZX planes. C represents the number of channels of the planar feature map, and H and W represent the resolution of the grid.

[0157] The encoding process compresses planar feature maps channel by channel using a trained image compression model (e.g., the encoder of the compression model). Existing LIC models are trained based on three-channel RGB images, so they cannot directly compress planar feature maps. Existing LIC models can be used after copying or expanding the planar feature maps of each channel. Alternatively, a single-channel LIC model that encodes one channel of planar features can be used.

[0158] For a 1-channel LIC model, as an example, a method of applying LIC compression of a 1-channel feature map multiple times can be used. R N×C×H×W After separating the color and density plane feature maps of the dimension into 1 channel each, the LIC model can be applied to each channel.

[0159] Alternatively, a method of applying LIC compression once to a 1-channel feature map according to the spatial packing described above can be used. In the same way as the above-described method, R N×C×H×W Spatially packing the color and density planar feature maps of the dimension R N×1×pH×pW After converting to a 1-channel feature map of dimensionality, LIC compression is applied to the converted feature map.

[0160] The decoding process uses a learned image compression model (e.g., a decoder of the compression model) to convert the compressed feature map to the original size R.N×C×H×W Restore the I volume feature map of the dimension.

[0161] As another example, tensor features can be compressed using Neural Network Compression (NNC). NNC can be used to reduce the size and computational complexity of tensor features while maximizing the performance of the tensor features in the I-volume, and to compress tensor features into a form that is interoperable across various deep learning frameworks.

[0162] The encoding process of NNC can be performed by the encoding pipeline illustrated in Fig. 13. The encoding pipeline can include all or part of a model pruner (1320), a model quantizer (1330), and a weight encoder (1340).

[0163] The model pruner (1320) simplifies the tensor features of the I volume using pruning. For example, pruning sets weights (e.g., tensor feature values) smaller than a preset threshold to zero. The model quantization unit (1330) quantizes the pruned weights (e.g., tensor feature values ​​greater than or equal to the threshold). The weight encoding unit (1340) applies entropy encoding to the quantized weights to generate a bitstream of the weights.

[0164] To improve feature compression performance, post-processing techniques (e.g., updating features by fine-tuning bias values ​​while keeping the features fixed) can be applied to features to which weight encoding (or entropy encoding) has been applied.

[0165] The decoding process can be performed by a decoding pipeline as illustrated in Fig. 14. The decoding pipeline can include all or part of a weight decoder (1410), a model dequantizer (1420), and a model reconstructor (1430).

[0166] The weight decoding unit (1410) applies entropy encoding to the bitstream of weights (e.g., tensor feature values ​​of the I volume) to generate quantized weights. The model dequantization unit (1420) dequantizes the quantized weights to generate pruned weights. The model restoration unit (1430) restores the I volume based on the pruned weights.

[0167] Meanwhile, the aforementioned neural network compression method can be applied to compress the parameters of a tensor generation model with respect to the generated tensor of volume B. Alternatively, the aforementioned neural network compression method can also be applied to compress the parameters of other neural networks (e.g., MLPs used in the rendering process) within a volume encoding device.

[0168] Additional information about the optimal network structure and compression method for compressing features can be conveyed using the bitstream format / unit defined in the NNC standard (e.g., NNR (Neural Network compression and Representation) unit).

[0169] As an example, a dynamic volume IPS (INR Parameter Set) is defined that includes the following syntax elements to convey the aforementioned additional information.

[0170] Information related to the neural network structure includes syntax defining the sizes of MLPs and CNNs that constitute the model (e.g., tensor generation models), syntax defining the structure of edges that constitute the model, syntax defining the format and precision of parameters that constitute the model, syntax defining upper and lower bounds on the sizes and number of MLPs and CNNs that constitute the model, etc.

[0171] Information related to camera parameters includes camera internal parameters and camera external parameters as described above.

[0172] Information related to quantization parameters includes quantization parameters of a 2D video codec, minimum and maximum values ​​of features before quantization in 2D video codec / LIC-based compression, syntax defining feature packing based on 2D video codec / LIC, quantization parameters of NNC, syntax defining upper and lower limits of the available sizes of quantization parameters of NNC, etc.

[0173] Information about motion compensation includes syntax that defines the structure of GOV.

[0174] Information related to the input range includes syntax defining the upper and lower bounds of the 4D spatiotemporal coordinates and the viewpoint direction. As mentioned above, the 4D spatiotemporal coordinates represent (x, y, z, t).

[0175] Meanwhile, IPS can be transmitted after lossless compression.

[0176] Hereinafter, a method for encoding or restoring a dynamic volume, i.e., a dynamic multi-view video, is described using the illustrations of FIGS. 15 and 16.

[0177] FIG. 15 is a flowchart illustrating a method for a volume encoding device to encode dynamic multi-view video according to one embodiment of the present disclosure.

[0178] The volume encoding device obtains the current time and current view point (S1500).

[0179] A volume encoding device can obtain information about the structure of a GOV (Group of Volume) including an I volume and a B volume. By utilizing the GOV, the volume encoding device can process the time of a dynamic 3D video.

[0180] The volume encoding device obtains the I volume when the current time corresponds to the encoding of the I volume, and decomposes the I volume to generate low-dimensional tensors expressing the tensor features of the I volume (S1502).

[0181] The volume encoding device can decompose the I volume based on vector matrix decomposition or CP (Canonical Polyadic) decomposition, as shown in FIG. 9a or FIG. 9b.

[0182] The volume encoding device reconstructs the tensor features of the B volume based on the two closest I volumes using the tensor generation model when the current time corresponds to the encoding of the B volume (S1504).

[0183] The volume encoder can determine two nearest I-volumes based on information about the structure of the GOV. The volume encoder can obtain tensors having the same axis from the two nearest I-volumes. The volume encoder can input the tensors having the same axis into a tensor generation model to construct a generated tensor, as in the example of FIG. 10 or FIG. 11.

[0184] As an example, a tensor generation model may include an encoder and a decoder, as illustrated in FIG. 10. The tensor generation model may input tensors with the same axis to the encoder to generate intermediate features and accumulate the current time in the intermediate features. The tensor generation model may input the intermediate features with the accumulated current time to the decoder to construct a generated tensor.

[0185] The volume encoder can obtain a fixed tensor orthogonal to the generated tensor from an adjacent I volume. The volume encoder can reconstruct the decomposed tensors of the B volume based on the fixed tensor and the generated tensor, as in the examples of FIGS. 12A to 12C.

[0186] The volume encoding device encodes tensor features of the I volume (S1506).

[0187] As an example, a volume encoding device can pack tensor features of an I volume into a video format. The volume encoding device can apply a video encoding method to the tensor features of the packed I volume to generate a bitstream, and transmit the generated bitstream to a volume decoding device.

[0188] As another example, a volume encoder can pack tensor features of an I volume. The volume encoder can apply a deep learning-based compression model to the tensor features of the packed I volume to generate a bitstream, and transmit the generated bitstream to a volume decoder.

[0189] As another example, the volume encoder can compress tensor features of the I volume using a neural network compression method based on an encoding pipeline.

[0190] The volume encoding device encodes the weights of the tensor generation model (S1508).

[0191] As an example, a volume encoding device may encode the weights of a tensor-generating model using a neural network compression method based on an encoding pipeline. The volume encoding device may prune the weights of the tensor-generating model. The volume encoding device may quantize the pruned weights using quantization parameters and encode the quantized weights.

[0192] Meanwhile, the volume encoding device can generate an image of the current time based on the camera parameters and the tensor features of the I volume, if the current time corresponds to the encoding of the I volume. As described above, a transformation between a "coordinate / viewpoint on the volume" and a pixel on the image plane can be defined based on the camera parameters.

[0193] The volume encoding device can calculate colors and volume densities in voxels corresponding to the current time point based on the decomposed tensors of the I volume. The volume encoding device can generate an image of the current time point based on the colors and volume densities.

[0194] Additionally, the volume encoding device can generate an image of the current time based on the camera parameters and the tensor features of the reconstructed B volume, if the current time corresponds to the encoding of the B volume.

[0195] The volume encoding device can calculate colors and volume densities in voxels corresponding to the current time point based on the decomposed tensors of the B volume. The volume encoding device can generate an image of the current time point based on the colors and volume densities.

[0196] A volume encoder can train tensor features and a tensor generation model of an I volume based on an image at the current point in time and a ground truth (GT). Here, the GT can include original images corresponding to the I volume or the B volume.

[0197] The volume encoder can calculate a loss function based on the difference between the current image and the original image. The volume encoder can update the tensor features of the I volume and update the weights of the tensor generation model in a direction that reduces the loss function.

[0198] As an example, a volume encoding device can encode information about the structure of a GOV as additional information.

[0199] Additionally, the volume encoding device can encode additional information including information related to quantization parameters, information related to the structure of the tensor generation model, and camera parameters.

[0200] FIG. 16 is a flowchart illustrating a method for a volume decoding device to restore dynamic multi-view video according to one embodiment of the present disclosure.

[0201] The volume decoding device decodes tensor features representing a volume, weights of a tensor generation model, and additional information from a bitstream (S1600).

[0202] Here, the tensor features of the I-volume are represented by decomposed tensors. The decomposed tensors of the I-volume can be low-dimensional tensors generated by vector-matrix decomposition or canonical polyadic (CP) decomposition.

[0203] The tensor generative model can be used to reconstruct the B volume.

[0204] Additional information may include information related to quantization parameters, information about the structure of the tensor generation model, and camera parameters.

[0205] The volume decoding device can decode the tensor features of the I volume as follows.

[0206] As an example, a volume decoding device can decode tensor features of an I volume packed in a video format. The volume decoding device can apply a video decoding method to the tensor features of the I volume in the video format to generate tensor features of the unpacked I volume.

[0207] As another example, a volume decoding device can decode tensor features of an I-volume compressed by an encoder of a deep learning-based compression model. The volume decoding device can apply a decoder of the deep learning-based compression model to the tensor features of the compressed I-volume to generate tensor features of a decompressed I-volume.

[0208] As another example, the volume decoding device can decode tensor features of the I volume using a neural network compression method based on a decoding pipeline.

[0209] Additionally, the volume decoding device can decode information regarding the structure of a GOV (Group of Volume) including an I volume and a B volume as additional information. By utilizing the GOV, the volume decoding device can process the time of dynamic 3D video.

[0210] The volume decoding device uses information about the weights and quantization parameters of the tensor generation model to restore the tensor generation model so that it matches information about the structure of the tensor generation model (S1602).

[0211] As an example, a volume decoding device can decode quantized weights of a tensor-generating model using a neural network compression method based on a decoding pipeline. The volume decoding device can dequantize the quantized weights to generate pruned weights. The volume decoding device can generate a tensor-generating model using the dequantized weights.

[0212] The volume decoding device reconstructs the tensor features of the B volume based on the two nearest decoded I volumes using a tensor generation model (S1604).

[0213] As an example, the volume decoding device can determine the locations of two closest I volumes based on information about the structure of the GOV.

[0214] A volume decoding device can obtain tensors having the same axis from two adjacent I volumes. The volume decoding device can input the tensors having the same axis into a tensor generation model to construct a generated tensor, as in the example of FIG. 10 or FIG. 11.

[0215] As an example, a tensor generation model may include an encoder and a decoder, as illustrated in FIG. 10. The tensor generation model may input tensors with the same axis to the encoder to generate intermediate features and accumulate the current time in the intermediate features. The tensor generation model may input the intermediate features with the accumulated current time to the decoder to construct a generated tensor.

[0216] The volume decoding device can obtain a fixed tensor orthogonal to the generated tensor from an adjacent I volume. The volume decoding device can reconstruct the decomposed tensors of the B volume based on the fixed tensor and the generated tensor, as shown in the examples of FIGS. 12A to 12C.

[0217] The volume decryption device obtains the current time and current view point (S1606).

[0218] The volume decoding device generates an image of the current time based on the camera parameters and the tensor features of the I volume when the current time corresponds to the decoding of the I volume (S1608).

[0219] The volume decoding device can calculate colors and volume densities in voxels corresponding to the current time point based on the decomposed tensors of the I volume. The volume decoding device can generate an image of the current time point based on the colors and volume densities.

[0220] As described above, a transformation between a “coordinate / viewpoint on the volume” and a pixel on the image plane can be defined based on the camera parameters.

[0221] The volume decoding device generates an image of the current time based on the camera parameters and the tensor features of the reconstructed B volume when the current time corresponds to the decoding of the B volume (S1610).

[0222] The volume decoding device can calculate colors and volume densities in voxels corresponding to the current time point based on the decomposed tensors of the B volume. The volume decoding device can generate an image of the current time point based on the colors and volume densities.

[0223] The volume decoding device can generate a 2D video corresponding to an arbitrary point in time by concatenating 2D images corresponding to the I volume and the B volume.

[0224] Although the flowchart / timing diagram of this specification describes each process as being executed sequentially, this is merely an illustrative description of the technical idea of ​​one embodiment of the present disclosure. In other words, a person of ordinary skill in the art to which one embodiment of the present disclosure belongs may modify and apply various modifications and variations by changing the order described in the flowchart / timing diagram without departing from the essential characteristics of one embodiment of the present disclosure, or by executing one or more of the processes in parallel. Therefore, the flowchart / timing diagram is not limited to a chronological order.

[0225] It should be understood that the exemplary embodiments described above can be implemented in many different ways. The functions or methods described in one or more examples can be implemented in hardware, software, firmware, or any combination thereof. It should be understood that the functional components described herein are labeled as "units" to further emphasize their implementation independence.

[0226] Meanwhile, the various functions or methods described in this embodiment may be implemented as instructions stored on a non-transitory storage medium that can be read and executed by one or more processors. Non-transitory storage media include, for example, all types of storage devices that store data in a form readable by a computer system. For example, non-transitory storage media include storage media such as erasable programmable read-only memory (EPROM), flash drives, optical drives, magnetic hard drives, and solid-state drives (SSDs).

[0227] The above description is merely an example of the technical idea of ​​the present embodiment, and those skilled in the art will appreciate that various modifications and variations can be made without departing from the essential characteristics of the present embodiment. Therefore, the present embodiments are not intended to limit the technical idea of ​​the present embodiment, but rather to explain it, and the scope of the technical idea of ​​the present embodiment is not limited by these embodiments. The scope of protection of the present embodiment should be interpreted by the claims below, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of rights of the present embodiment.

[0228]

[0229]

[0230] CROSS-REFERENCE TO RELATED APPLICATION

[0231] This patent application claims priority to Korean patent application No. 10-2023-0181372, filed in Korea on December 14, 2023, and Korean patent application No. 10-2024-0165811, filed in Korea on November 20, 2024, the entire contents of which are incorporated herein by reference.

Claims

1. A method for restoring multi-view video performed by a volume decoding device, A step of decoding tensor features representing a volume I from a bitstream, weights of a tensor generation model, and side information, wherein the tensor features of the I volume are represented by decomposed tensors, the tensor generation model is used for reconstructing a volume B, and the side information includes information related to quantization parameters and information about a structure of the tensor generation model; A step of restoring the tensor generation model so as to match information about the structure of the tensor generation model by using information related to the weights of the tensor generation model and the quantization parameters; A step of reconstructing tensor features of the B volume based on two adjacent decoded I volumes using the above tensor generation model; Step of obtaining the current time and current view point; and If the current time corresponds to the decoding of the I volume, a step of generating an image of the current time based on the tensor features of the I volume. If the current time corresponds to the decoding of the B volume, a step of generating an image of the current time based on the tensor features of the reconstructed B volume. A method comprising:

2. In paragraph 1, The step of decoding the tensor features of the above I volume is, A step of decoding tensor features of an I volume packed in a video format; and A step of generating tensor features of an unpacked I volume by applying a video decoding method to the tensor features of the I volume of the above video format. A method comprising:

3. In paragraph 1, The step of generating an image of the current time based on the tensor features of the above I volume is as follows. A step of calculating colors and volume densities in voxels corresponding to the current point in time based on the decomposed tensors of the above I volume; and A step of generating an image of the current point in time based on the above colors and the above volume densities. A method comprising:

4. In paragraph 1, As the above additional information, it further includes a step of decrypting information about the structure of a GOV (Group of Volume) including the I volume and the B volume, The step of reconstructing the tensor features of the above B volume is: A method for determining the locations of two adjacent I volumes based on information about the structure of the above GOV (Group of Volume).

5. In paragraph 1, The step of reconstructing the tensor features of the above B volume is: A step of obtaining tensors having the same axis from the two adjacent I volumes; and A step of configuring a generated tensor by inputting tensors having the same axis into the tensor generation model. A method comprising:

6. In paragraph 5, The step of reconstructing the tensor features of the above B volume is: A step of obtaining a fixed tensor orthogonal to the generated tensor from an adjacent I volume; and A step of reconstructing the decomposed tensors of the B volume based on the fixed tensor and the generated tensor. A method comprising:

7. In paragraph 5, The above tensor generation model includes an encoder and a decoder, The steps of constructing the above generated tensor are: A step of generating intermediate features by inputting tensors having the same axis into the encoder; a step of accumulating the current time in the above intermediate feature; and A step of configuring the generated tensor by inputting the intermediate features accumulated at the current time into the decoder. A method comprising:

8. A method for encoding multi-view video performed by a volume encoding device, Step of obtaining the current time and current view point; If the current time corresponds to the encoding of the I volume, a step of obtaining the I volume and decomposing the I volume to generate low-dimensional tensors expressing tensor features of the I volume; A step of reconstructing tensor features of the B volume based on two nearest neighboring I-volumes using a tensor generation model, when the current time corresponds to the encoding of the B volume; A step of encoding tensor features of the above I volume; and Step of encoding the weights of the above tensor generation model A method comprising:

9. In paragraph 8, A step of generating an image of the current time based on tensor features of the I volume or tensor features of the reconstructed B volume, depending on whether the current time corresponds to encoding of the I volume or the B volume. A method further comprising:

10. In paragraph 9, A step of calculating a loss function based on the difference between the image at the current point in time and the original image; A step of updating the tensor features of the I volume in a direction that reduces the loss function; and A step of updating the weights of the tensor generation model in a direction that reduces the loss function. A method further comprising:

11. In paragraph 8, A step of obtaining information about the structure of a GOV (Group of Volume) including the above I volume and the above B volume; A step of determining the positions of the two most adjacent I volumes based on information about the structure of the above GOV; and As additional information, a step of encoding information about the structure of the GOV A method further comprising:

12. In paragraph 8, The step of reconstructing the tensor features of the above B volume is: A step of obtaining tensors having the same axis from the two adjacent I volumes; and A step of configuring a generated tensor by inputting tensors having the same axis into the tensor generation model. A method comprising:

13. In paragraph 12, The step of reconstructing the tensor features of the above B volume is: A step of obtaining a fixed tensor orthogonal to the generated tensor from an adjacent I volume; and A step of reconstructing the decomposed tensors of the B volume based on the fixed tensor and the generated tensor. A method comprising:

14. In paragraph 8, The step of encoding the tensor features of the above I volume is, A step of packing the tensor features of the above I volume into a video format; and A step of generating a bitstream by applying a video encoding method to the tensor features of the packed I volume. A method including:

15. In paragraph 8, The step of encoding the weights of the above tensor generation model is: A step of pruning the weights of the above tensor generation model; A step of quantizing the pruned weights using information related to quantization parameters; and Step of encoding quantized weights A method comprising:

16. In Article 16, A method further comprising the step of encoding additional information including information related to the quantization parameters and information regarding the structure of the tensor generation model.

17. A method for providing video data to a video decoding device, A step of encoding the above video data into a bitstream; and A step of transmitting the above bitstream to the image decoding device Including, The step of encoding the above video data is: Step of obtaining the current time and current view point; If the current time corresponds to the encoding of the I volume, a step of obtaining the I volume and decomposing the I volume to generate low-dimensional tensors expressing tensor features of the I volume; A step of reconstructing tensor features of the B volume based on two nearest neighboring I-volumes using a tensor generation model, when the current time corresponds to the encoding of the B volume; A step of encoding tensor features of the above I volume; and Step of encoding the weights of the above tensor generation model A method, characterized by including:

Citation Information

Patent Citations

  • Positive electrode active material for lithium secondary battery, method of preparing the same, and the lithium secondary battery

    KR1020240100228A

  • Insulation monitoring system

    KR102158595B1

  • Method, device, and medium for generating super-resolution video

    US20230057261A1