Method and apparatus for video coding based on implicit neural video representation

The proposed video coding method and device address the inefficiencies in existing INVR models by compressing low-frequency components and generating high-frequency components as needed, enhancing video encoding efficiency and quality.

WO2025121685A1PCT designated stage expired Publication Date: 2025-06-12HYUNDAI MOTOR CO LTD +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/017295
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-04
Filing Date
2024-11-05
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing implicit neural video representation (INVR) models face challenges in efficiently encoding and decoding videos due to increased model size and learning time when learning high-frequency components, which affects video encoding efficiency and quality.

Method used

A method and device for video coding that utilize an INVR model to represent videos in the frequency domain, compress low-frequency components, and generate high-frequency components as needed on the decoder side, thereby optimizing video encoding efficiency and quality.

Benefits of technology

This approach improves video encoding efficiency and quality by reducing the overhead required for embedding high-frequency components and allowing for flexible generation of high-frequency components on the decoder side.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024017295_12062025_PF_FP_ABST
    Figure KR2024017295_12062025_PF_FP_ABST
Patent Text Reader

Abstract

The present embodiment discloses a video coding method and apparatus based on an implicit neural video representation (INVR). In the present embodiment, an image decoding apparatus decodes decoder parameters and embedding of a 2D INVR model. Here, the embedding is a compressed representation of a low frequency component of the current 2D frame, and is generated and provided by an encoder of the 2D INVR model. The image decoding apparatus reconstructs a decoder of the 2D INVR model by using the decoder parameters. The image decoding apparatus reconstructs the current 2D frame by inputting the embedding into the decoder of the 2D INVR model.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for video coding based on implicit neural network video representation

[0001] The present disclosure relates to a video coding method and device based on an implicit neural network video representation.

[0002] The content described below merely provides background information related to the present invention and does not constitute prior art.

[0003] Recently, there has been active research on implicit neural network representation models, which represent various data, including images, using neural network structures. Conventional video representation methods use explicit representations, which represent RGB pixel values ​​for each pixel location. To replace these explicit representations, implicit neural representations are being introduced, which represent functions that generate (r,g,b) values ​​from (x,y) coordinates (pixel locations) in an image using neural networks. Compared to explicit representations, implicit neural representations can be used regardless of image resolution, and the (r,g,b) values ​​are instantly restored when the spatial and temporal coordinates of an image pixel are input, offering practical advantages over existing decoders.

[0004] Meanwhile, video compression technology based on implicit neural video representation (INVR, or Neural Representation for Videos, NeRV) generates an implicit representation of the video in the pixel domain. According to the Neural Tangent Kernel (NTK) theory related to neural network learning, neural network learning first learns low-frequency components and then high-frequency components. Since existing INVR also performs learning according to the NTK theory, the model size may increase and the learning time may be long in order to learn high-frequency components. Therefore, in order to improve video encoding efficiency and video quality, a method is needed to construct a video representation model based on implicit neural network representation that considers frequency components.

[0005] The present disclosure provides a video coding method and device for generating a compressed embedding by applying an implicit neural video representation (INVR) model to a video in the frequency domain, transmitting model parameters and the embedding, and restoring the video based on the transmitted model parameters and the embedding.

[0006] In addition, the present disclosure aims to provide a video coding method and device that compresses low-frequency components of a video based on an INVR model and generates high-frequency components of the video as needed on the decoder side.

[0007] According to an embodiment of the present disclosure, a method for restoring a current 2D frame, performed by a video decoding device, is provided, comprising: decoding decoder parameters and an embedding of a 2D (Dimensional) INVR (Implicit Neural Video Representation) model from a bitstream, wherein the embedding is a compressed representation of a low-frequency component of the current 2D frame and is generated and provided by an encoder of the 2D INVR model; restoring a decoder of the 2D INVR model using the decoder parameters; and inputting the embedding to a decoder of the 2D INVR model to restore the current 2D frame, wherein the restoring the current 2D frame comprises: applying convolution-based upsampling to the embedding to restore low-frequency components of multiple resolutions; and fusion of the low-frequency components of the multi-resolutions to restore the low-frequency component of the current 2D frame.

[0008] According to another embodiment of the present disclosure, a method for encoding a current 2D frame, performed by a video encoding device, is provided, comprising: inputting the current 2D frame into an encoder of a 2D (Dimensional) INVR (Implicit Neural Video Representation) model to generate a compressed embedding; inputting the embedding into a decoder of the 2D INVR model to reconstruct the current 2D frame; and encoding decoder parameters of the embedding and the 2D INVR model, wherein the step of generating the embedding includes: transforming the current 2D frame based on a preset transformation method to generate a low-frequency component of the current 2D frame; and applying convolution-based downsampling to the low-frequency component to generate the embedding.

[0009] According to another embodiment of the present disclosure, a method for providing video data to a video decoding device is provided, comprising: encoding the video data into a bitstream; and transmitting the bitstream to the video decoding device, wherein the encoding the video data comprises: inputting a current 2D frame into an encoder of a 2D (Dimensional) INVR (Implicit Neural Video Representation) model to generate a compressed embedding; inputting the embedding into a decoder of the 2D INVR model to reconstruct the current 2D frame; and encoding the embedding and decoder parameters of the 2D INVR model, wherein the generating the embedding comprises: converting the current 2D frame based on a preset conversion method to generate a low-frequency component of the current 2D frame; and applying convolution-based downsampling to the low-frequency component to generate the embedding.

[0010] As described above, according to the present embodiment, there is provided a video coding method and device that applies an INVR model to a video converted to a frequency domain to generate a compressed embedding, transmits model parameters and the embedding, and restores the video based on the transmitted model parameters and the embedding, thereby making it possible to improve video encoding efficiency and video quality.

[0011] In addition, according to the present embodiment, by providing a video coding method and device that compresses low-frequency components of a video based on an INVR model and generates high-frequency components of the video as needed on the decoder side, it is possible to reduce the overhead required for the INVR model to embed high-frequency components.

[0012] FIG. 1 is a block diagram illustrating an INVR (Implicit Neural Video Representation)-based video encoding device according to one embodiment of the present disclosure.

[0013] FIG. 2 is a block diagram illustrating an INVR-based video decoding device according to one embodiment of the present disclosure.

[0014] FIG. 3 is an exemplary diagram showing the operation of a convolutional layer according to one embodiment of the present disclosure.

[0015] Figure 4 is an example diagram showing a SISR (Single Image Super Resolution) network.

[0016] Figure 5 is an example diagram showing a residual block used in SISR.

[0017] Figure 6 is an example diagram showing the NeRV (Neural Representations for Videos) model.

[0018] FIG. 7A and FIG. 7B are exemplary diagrams illustrating the concept of the NeRV model according to one embodiment of the present disclosure.

[0019] FIG. 8 is an exemplary diagram illustrating an encoding pipeline of a NeRV model according to one embodiment of the present disclosure.

[0020] FIG. 9 is an exemplary diagram illustrating a decryption pipeline of a NeRV model according to one embodiment of the present disclosure.

[0021] FIG. 10A and FIG. 10B are exemplary diagrams showing a spatial information-based INVR model according to one embodiment of the present disclosure.

[0022] FIG. 11 is an exemplary diagram illustrating a spatial information-based INVR model according to one embodiment of the present disclosure.

[0023] FIG. 12 is an exemplary diagram illustrating an encoding pipeline of an INVR model according to one embodiment of the present disclosure.

[0024] FIG. 13 is an exemplary diagram illustrating a decryption pipeline of an INVR model according to one embodiment of the present disclosure.

[0025] FIG. 14 is a flowchart illustrating a method for an image encoding device to encode 2D video according to one embodiment of the present disclosure.

[0026] FIG. 15 is a flowchart illustrating a method for a video decoding device to restore a 2D video according to one embodiment of the present disclosure.

[0027] Hereinafter, embodiments of the present invention will be described in detail with reference to exemplary drawings. When designating components in each drawing, it should be noted that, where possible, identical components are given the same reference numerals, even if they appear in different drawings. Furthermore, in describing the present embodiments, detailed descriptions of related known structures or functions will be omitted if they are deemed to obscure the gist of the present embodiments.

[0028] The present embodiment relates to encoding and decoding of images (video). More specifically, the present invention provides a video coding method and device that applies an implicit neural video representation (INVR) model to a 2D video converted to the frequency domain to generate a compressed embedding, transmits the model parameters and the embedding, and restores the video based on the transmitted model parameters and the embedding.

[0029] FIG. 1 is a block diagram illustrating a 2D INVR-based image encoding device according to one embodiment of the present disclosure.

[0030] In the example of FIG. 1, a 2D (Dimensional) INVR-based video encoding device (hereinafter, referred to as the 'video encoding device') generates (i.e., trains) a 2D INVR model to learn to generate an embedding of a 2D video in the frequency domain, compresses the decoder and embedding of the learned 2D INVR model to generate a bitstream, and transmits the generated bitstream. The video encoding device includes an encoder (110), a decoder (120), and a neural network compressor (130) corresponding to an encoder and a decoder of the 2D INVR model. Here, the components included in the video encoding device according to the present disclosure are not necessarily limited thereto. The video encoding device may additionally include a training unit (not shown) for training a 2D INVR model, or may be implemented in a form linked with an external training unit.

[0031] Each component of the video encoding device may be implemented in hardware, software, or a combination of hardware and software. Furthermore, the functions of each component may be implemented in software, with a microprocessor executing the software functions corresponding to each component.

[0032] The video encoding device can store the bitstream of encoded video data on a non-transitory recording medium or transmit it to a 2D INVR-based video decoding device using a communication network.

[0033] FIG. 2 is a block diagram illustrating a 2D INVR-based image decoding device according to one embodiment of the present disclosure.

[0034] In the example of FIG. 2, a 2D INVR-based video decoding device (hereinafter, referred to as the "video decoding device") decompresses a bitstream to restore a 2D INVR model and embedding, and generates a 2D video from the embedding using the restored 2D INVR model. The video decoding device includes a neural network decompressor (210), and a decoder (220) including a decoder of the restored 2D INVR model.

[0035] Similar to the video encoding device illustrated in Fig. 1, each component of the video decoding device may be implemented in hardware, software, or a combination of hardware and software. Furthermore, the functions of each component may be implemented in software, with a microprocessor configured to execute the software functions corresponding to each component.

[0036] Conventional INVR generates an implicit representation of a video in the pixel domain, but a video encoding device according to the present disclosure converts a video into the frequency domain, inputs the video in the frequency domain into a 2D INVR model, and performs video compression. When generating an implicit neural network for a video in the frequency domain, the video encoding device separates low-frequency signals and high-frequency signals of the video, and then trains the implicit neural network. Since high-frequency signals are more difficult to learn than low-frequency signals, the video encoding device skips the encoding process of high-frequency signals according to a predetermined threshold of the video, and preferentially learns low-frequency signals and embeds them in the implicit neural network. To embed high-frequency signals, the size of the neural network may increase, and the transmission bit rate may increase. Therefore, the video encoding device can reduce the overhead required to embed high-frequency components by generating high-frequency components in a decoder as needed. In addition, the video decoding device can improve the quality of an output video based on a fixed bitstream size.

[0037] Below, before describing the components illustrated in FIGS. 1 and 2, the deep learning-related element technologies used in the present disclosure are described.

[0038] I-1. Multilayer Perceptron (MLP)

[0039] A multilayer perceptron (MLP) consists of edges connecting multiple neurons. These neurons are organized into layers. In addition to an input layer and an output layer, an MLP may include one or more hidden layers. A hidden layer may contain one or more hidden nodes. Each edge may be assigned a different weight. An activation function may be used as the output propagates from one layer to the next. Examples of activation functions include the sigmoid function, the tangent hyperbolic function, or the Relu (Rectified Linear Unit) function.

[0040] To learn the target behavior, the process of updating the weights that make up an MLP is called training. Training typically utilizes the Stochastic Gradient Descent (SGD) method, which is based on the backpropagation algorithm. The process of calculating the feed-forward output of the MLP using the weights generated through training is called inference or testing.

[0041] Hereinafter, multilayer neural networks, MLPs, or forward networks can be used interchangeably.

[0042] I-2. CNN(Convolutional Neural Network)

[0043] CNN, a neural network comprised of multiple convolutional layers and pooling layers, is a deep learning technique known to be ideal for image processing. Convolutional layers extract feature maps (also known as "features" or "characteristics") using multiple kernels or filters. The kernel coefficients that make up the filters are parameters determined during the learning process.

[0044] Among the convolutional layers of CNN, the front layer closer to the input extracts feature maps that respond to simple, low-level image features such as lines, points, or surfaces, while the back layer closer to the output extracts feature maps that respond to higher-level features such as textures and object parts.

[0045] FIG. 3 is an exemplary diagram showing the operation of a convolutional layer according to one embodiment of the present disclosure.

[0046] A convolutional layer generates a feature map from an input image using a convolution operation. The example in Fig. 3 illustrates a kernel (or filter) with a kernel size of 3×3. The kernel size is also referred to as the kernel size or filter size. The kernel has kernel parameters, also called weights. The kernel illustrated in Fig. 3 has a total of nine kernel parameters. The kernel parameters are initially set to arbitrary values, and their values ​​can be updated based on learning.

[0047] The convolution layer performs convolution operations using blocks of the same size as the kernel size in the input image. At this time, the blocks of the same size as the kernel size in the input image are referred to as windows.

[0048] When filtering an input image in raster scan order, the window movement size is called the stride. In the example of Fig. 3, the stride is 1. If the stride is set to 2, the convolution operation is performed by spacing the window by 2 samples, and as a result, the width and height of the feature map become half the width and height of the input image.

[0049] As mentioned above, a single convolutional layer can include multiple filters. The number of filters or kernels is called a channel. In other words, the number of channels is equal to the number of filters. Furthermore, the number of filters determines the dimensionality of the feature map.

[0050] Padding refers to a method of expanding input data by filling the area around it with a specific value before performing a convolution operation. Padding is primarily used to adjust the spatial size of the output data. The padding value can be determined by hyperparameters, but zero-padding is commonly used. Without padding, the spatial size of the output data decreases with each convolutional layer, potentially causing boundary information to disappear. Therefore, padding is used to prevent this problem. Specifically, padding can be used to match the spatial sizes of the output data from the convolutional layer to the input data.

[0051] The deconvolution layer performs the opposite operation to the convolution layer. It generates the desired data image from the input feature map as output.

[0052] The pooling layer performs pooling, a process of subsampling the feature map generated by the convolutional layer. The pooling layer uses a 2×2 window to select samples so that the output is half the width and height of the input. In other words, the pooling layer is used to reduce the size of the input image or input feature map by condensing a 2×2 region into a single sample.

[0053] The opposite concept of a pooling layer is defined as an unpooling layer. Unpooling layers, in contrast to pooling layers, function to expand the dimensionality and are primarily used after deconvolution layers.

[0054] The convolutional encoder-decoder architecture is a network structure composed of pairs of convolutional layers and deconvolutional layers. The convolutional encoder consists of a convolutional layer and a pooling layer, and outputs a feature map (or feature vector) from the input image. The final output vector of the convolutional encoder is also referred to as a latent vector. The convolutional decoder consists of a deconvolutional layer and an unpooling layer, and generates an output image from the feature map or latent vector.

[0055] The inputs and outputs of a convolutional encoder-decoder can be configured in various ways depending on the application and network purpose. For example, the inputs and outputs can be optical flow maps, saliency maps, image frames, etc.

[0056] Figure 4 is an example diagram showing an SISR network.

[0057] One example of CNN application is Single Image Super Resolution (SISR). The SISR network generates a high-resolution image from a low-resolution input image. As illustrated in Figure 4, the SISR network may include multiple convolutional layers. Each convolutional layer includes an activation function, such as the Rectified Linear Unit (ReLU). The parameters of the SISR network can be trained so that the resulting SR (Super Resolution) image approximates the Ground Truth (GT).

[0058] SR methods using CNN can improve SR performance by increasing the depth (e.g., increasing the number of convolutional layers). To overcome the overfitting problem that may occur in learning due to the increase in depth, a residual block that can perform skip connection and residual learning can be used in the SISR network. The residual block, as illustrated in Fig. 5, is a block that stores input features x l In addition to the path that applies the convolution operation, it includes a skip path. In addition, the residual block outputs x l+1 When generating, a path or skip path for applying convolution operations can also be selected based on learning efficiency. In the example of Fig. 5, the residual block includes a BN (Batch Normalization) layer.

[0059] For example, Enhanced Deep Residual Networks (EDSR) improves network performance by continuously connecting residual blocks to increase their depth. Another example is Accurate Image Super-Resolution Using Very Deep Convolutional Networks (VDSR), a CNN model based on the Visual Geometry Group (VGG) network. It uses residual learning, a method that adds residual frames to the final output. VDSR adds residual signals to the very end of the network, thereby adding them to the input signal.

[0060] I-3. Implicit Neural Representation Model

[0061] Recently, research on implicit neural network representation models, which represent various data, including images, using neural network structures, has been active. Conventional video representation methods use an explicit representation method, which assigns RGB pixel values ​​to each pixel location. To replace this explicit representation method, implicit neural network representation models are being introduced, which represent functions that generate (r,g,b) values ​​from (x,y) coordinates (pixel locations) in an image using neural networks. Compared to explicit representation methods, implicit neural network representation models can be used regardless of image resolution.

[0062] As an example of 2D video coding using an implicit neural network representation model, there is Implicit Neural Video Representation for Coding (INVRC). INVRC technology generates (r,g,b) values ​​for the (x,y) coordinates, which are pixel locations in an image. As an example of INVRC technology, there is the NeRV (Neural Representation for Video) model. As mentioned above, the explicit representation method expresses pixel values ​​on the (x,y,t) grid using the pixel locations (x,y) in the image and the time index t of the frame. The implicit neural network representation model outputs RGB pixel values ​​from the input representing the location (x,y,t). However, training the neural network at all locations (x,y,t) can significantly increase the computational complexity, so, as shown in the example of FIG. 6, the NeRV model outputs the RGB image of the entire frame at the time index t using a neural network structure that only takes the time index t as input.

[0063] FIG. 7A and FIG. 7B are exemplary diagrams illustrating the concept of the NeRV model according to one embodiment of the present disclosure.

[0064] To easily provide image output for a time index t input, it may be more effective to use a convolutional layer than to use an MLP network of an existing implicit representation model. The NeRV model can generate output using a structure in which NeRV blocks composed of multiple convolutional layers are stacked. As shown in the example of Fig. 7a, an MLP-based output and a NeRV block-based output can be compared. The NeRV block includes convolutional layers that perform convolution, a pixel shuffle layer, and an activation layer, as shown in the example of Fig. 7b. In the example of Fig. 7b, C represents the number of channels, W and H represent the width and height of the input of the NeRV block, and S represents an upscaling factor.

[0065] Additionally, the NeRV model embeds the time index t in a high-dimensional space as in Equation 1 and then uses the embedded index as input.

[0066]

[0067] As illustrated in Equation 3, the temporal index t can be mapped to a 2L-dimensional vector γ(t). By using the embedded temporal index t to train the NeRV model, the NeRV model can better predict video data containing high-frequency variations.

[0068] As the loss function of the NeRV model, a combination of L1 loss and SSIM (Structural Similarity Index) loss can be used. During training of the NeRV model, the aforementioned loss can be calculated at all pixel locations of the image estimated by the NeRV model based on the input and the ground truth (GT) image (i.e., the original image). The L1 loss is calculated using the absolute value of the difference between pixels in the estimated image and the GT image. The SSIM loss is calculated based on the mean, standard deviation, and correlation of pixels in the estimated image and the GT image. Thereafter, the weights (or parameters) that constitute the NeRV model during the training process can be updated based on the calculated loss.

[0069] FIG. 8 is an exemplary diagram illustrating an encoding pipeline of a NeRV model according to one embodiment of the present disclosure.

[0070] Since the original video can be approximately represented using the NeRV model, an end-to-end compression technology using a neural network can be implemented by compressing and transmitting the NeRV model. An encoding pipeline for encoding the NeRV model may include all or part of a video overfitter (910), a model pruner (920), a model quantizer (930), and a weight encoder (940), as shown in the example of Fig. 8.

[0071] The video overfitter (910) expresses the input video frame as a NeRV model. At this time, the video encoding device trains the NeRV model using the aforementioned loss function, thereby allowing the NeRV model to express the video frame. The model pruner (920) simplifies the NeRV model structure, which has an MLP or a mixed form of MLP and convolutional layers, using pruning. For example, pruning sets weights smaller than a preset threshold to zero. The model quantization unit (930) quantizes the weights to which pruning has been applied (e.g., weights greater than the threshold). The weight encoding unit (940) applies entropy encoding to the quantized weights to generate a bitstream of the weights.

[0072] Meanwhile, as a reverse order of the encoding pipeline illustrated in FIG. 8, a decoding pipeline for decoding a NeRV model can be illustrated as in FIG. 9. The decoding pipeline can include all or part of a weight decoder (1010), a model dequantizer (1020), a model reconstructor (1030), and a video generator (1040), as illustrated in FIG. 9.

[0073] The weight decoding unit (1010) applies entropy encoding to the bitstream of the weights of the NeRV model to generate quantized weights. The model dequantization unit (1020) dequantizes the quantized weights to generate pruned weights. The model restoration unit (1030) restores the NeRV model based on the pruned weights. The video generation unit (1040) uses the restored NeRV model to generate restored video frames according to the time index.

[0074] As mentioned above, the NeRV model uses the time index t of the frame as input. Therefore, when a video having pixels of size t×h×w is input, the NeRV model samples the input video only t times, not t×h×w times, so a large gain in both encoding speed and decoding speed can be expected.

[0075] Meanwhile, the encoding pipeline and decoding pipeline as described above can also be used for compression and restoration of the INVR model.

[0076] Hereinafter, the implicit neural network video representation model, INVR model, and 2D INVR model are used interchangeably.

[0077] II. Embodiments according to the present disclosure

[0078] The video encoding device illustrated in FIG. 1 trains a 2D INVR model including an encoder (110) and a decoder (120) using a 2D video, thereby enabling the 2D INVR model to generate a 2D video. At this time, a video frame at time t is used as the input and GT (Ground Truth) of the 2D INVR model. That is, the encoder (110) and the decoder (120) constitute an autoencoder and can be trained end-to-end. The encoder (110) generates an embedding in which low-frequency components of a video frame are compressed, and the decoder (120) can be trained to generate a video frame based on the embedding. The training unit updates the parameters (weights) of the 2D INVR model using the difference between the output of the 2D INVR model and the GT, thereby enabling the INVR model to learn to generate a 2D video. A 2D INVR model that has completed learning can include 2D planar information contained in a 2D video based on parameters within the model.

[0079] The neural network compressor (130) compresses the parameters and embedding of the trained 2D INVR model to generate a bitstream. For example, the 2D INVR model can be compressed using the encoding pipeline illustrated in FIG. 8. The image encoding device transmits the generated bitstream to the image decoding device.

[0080] As an example, an image encoding device may encode quantization parameters related to the compression of an INVR model as additional information. The image encoding device may encode information regarding the structure of the INVR model as additional information.

[0081] The image decoding device illustrated in FIG. 2 decodes parameters and embedding of a 2D INVR model from a bitstream using a neural network decompressor (210). For example, a decoder of a 2D INVR model can be restored using the decoding pipeline illustrated in FIG. 9.

[0082] As an example, an image decoding device can decode quantization parameters related to the compression of an INVR model as additional information. The image decoding device can decode information regarding the structure of the INVR model as additional information.

[0083] A decoder (220) generates a 2D video from the embedding using a decoder of the restored 2D INVR model. The video decoding device decodes the embedding and then inputs the decoded embedding into the 2D INVR model to generate a restored 2D video.

[0084] The decoder (120) within the video encoding device and the decoder (220) within the video decoding device have the same structure. However, the decoder (120) within the video encoding device is trained end-to-end together with the encoder (110) by the training unit, and the decoder (220) within the video decoding device operates based on the trained parameters. Hereinafter, the encoder (110) and the decoder (120) within the video encoding device will be described. The structure and operation of the decoder (220) within the video decoding device may be replaced with the description of the decoder (120) within the video encoding device.

[0085] As an example, an image encoding device using a spatial information-based INVR model is described.

[0086] FIG. 10A and FIG. 10B are exemplary diagrams showing a spatial information-based INVR model according to one embodiment of the present disclosure.

[0087] The INVR model segments the input video into low-frequency and high-frequency components using a 2D discrete wavelet transform (DWT) to improve learning of implicit representations of the video by mitigating spectral bias. The encoder (110) of the INVR model forms a compact embedding using low-frequency components with small variations across frames. The embedding contains spatial information about the input video frame. The decoder (120) of the INVR model reconstructs the low-frequency components based on the embedding delivered by the encoder (110) and reconstructs the high-frequency components using the reconstructed low-frequency components.

[0088] Video V = [V1, V2, ..., V T ] is a video frame V at any time t. t is provided as an input to the encoder (110). At this time, 1 ≤ t ≤ T, and V t ∈V. The encoder (110) is a video V tTransform the signal w in the frequency domain t and can perform INVR-based compression. As in the example of Fig. 10a, the encoder (110) can include a 2D DWT module and a CNN-based downsampling block (Downsampling Block, DB).

[0089] To utilize spatial information even after transformation, the encoder (110) can apply 2D DWT to the video frame to transform it into multiple sub-band wavelet components. For example, by applying Haar Wavelet transform, the encoder (110) can transform the video frame V t to w t =[w t (LL), w t (LH), w t (HL), w t Convert to [(HH)]. Here, w t (LL), w t (LH), w t (HL), w t (HH) is a frequency domain signal generated by using the Haar Wavelet transform with low-frequency and high-frequency basis functions in the vertical and horizontal directions. For example, w t (LL) is a low-frequency signal generated by applying vertical and horizontal low-frequency basis functions.

[0090] To improve the image quality information of the INVR model, the encoder (110) w t Only the low frequency components of the encoder (110) can be embedded. t Medium w t By inputting only (LL) into the DB, a spatially adaptive embedding based on downsampling can be generated. The encoder (110) transmits the generated embedding to the decoder (120).

[0091] The decoder (120) decodes the frequency domain signal w based on the embedding t Restore and restore the w t_hat Restore frame V by inverse transformingt_hat As in the example of FIG. 10a, the decoder (120) may include all or part of a CNN-based upsampling block (UB), a multi-resolution fusion unit (MFU), a high frequency restorer (HFR), and a 2D inverse discrete wavelet transform (Inverse DWT, IDWT) module.

[0092] The decoder (120) restores low-frequency components from the embedding using multiple UBs. The multiple UBs sequentially upsample the spatial size of the embedding to the original size based on 2D convolution and pixel shuffling, thereby restoring the low-frequency components to the original dimension.

[0093] Since upsampling alone has limitations in improving the restoration quality, the decoder (120) improves the quality of the restored low-frequency components by using the MFU with a coarse-to-fine structure. By utilizing the low-frequency components of various resolutions generated by the UB (e.g., S1, S2, S3 in Fig. 10a), the MFU can improve the quality of the restored low-frequency components. The restored low-frequency components with improved quality are w t_hat It is expressed as (LL).

[0094] As shown in the example of Fig. 10b, the MFU includes at least one multi-resolution fusion block (MFB), each MFB including a transposed convolution (CT) block and residual blocks (RB). The MFB applies a transposed convolution to a low-frequency component of a first resolution (e.g., S1 in Fig. 10b) to generate a first low-frequency component of a second resolution. The second resolution of the first low-frequency component may be the same as the resolution of the second low-frequency component (e.g., S2 in Fig. 10b). The MFB concatenates the first low-frequency component to which the transposed convolution is applied and the second low-frequency component, and passes the concatenated low-frequency component to the residual block. RB includes two 2D convolution (Conv2D) blocks and a Leaky Rectified Linear Unit (LReLU) activation block, as shown in the example in Figure 10b. Additionally, RB includes a skip path between input and output. RB generates a residual of combined low-frequency components and adds the combined low-frequency components to the generated residual. By repeatedly applying MFB, MFU can improve the quality of the reconstructed low-frequency components based on the fusion of low-frequency components of various resolutions.

[0095] HFR is the restored low frequency component w t_hat Using (LL), the remaining band components w t (LH), w t (HL), w t By restoring (HH), w t_hat can be configured. HFR includes two Conv2D blocks and an LReLU activation block, as shown in the example of Fig. 10b. HFR can restore the remaining frequency components by applying 2D convolution to the low-frequency components.

[0096] The decoder (120) applies the restored frequency components to the 2D IDWT module to perform inverse discrete wavelet transform, thereby restoring the restored video frame V. t_hat Instead of restoring high frequency components based on learning of high frequency components, the number of parameters used in the training process can be reduced by having the decoder (120) synthesize high frequency components based on HFR.

[0097] As described above, the training unit can train the deep learning-based components (DB, UB, MFU, HFR, etc.) included in the encoder (110) and decoder (120) end-to-end. For example, the loss function can be defined based on the difference between the restored frame of the decoder (120) and the GT. The training unit can update the parameters of the encoder (110) and decoder (120) included in the INVR model in a direction that reduces the loss function.

[0098] As described above, the decoder (220) within the image decoding device has the same structure as the decoder (120) within the image encoding device, and operates in the same manner as the decoder (120) within the image encoding device based on trained parameters. Therefore, a description of the structure and operation of the decoder (220) within the image decoding device is omitted.

[0099] The total bits generated by INVR-based compression are the sizes of the embedding vector and the network information of the decoder (120). The size of the embedding vector is w t =[w t (LL), w t (LH), w t (HL), w t (HH)] depends on determining the low frequency band to be transmitted. In addition, the size of the network information of the decoder (120) includes network parameters and network structure information, and may be proportional to the resilience of the network. Therefore, the original video V t Wow restoration video Vt_hat D, a distortion of the liver t and full bit R t In terms of rate distortion optimization based on the combination of , a low frequency band to be transmitted can be selected. For example, J shown in Equation 2 t A low frequency band that minimizes the noise can be selected.

[0100]

[0101] In Equation 2, λ is the Lagrangian multiplier.

[0102] As mentioned above, the spatial information-based INVR model trains each video frame separately. As another example of incorporating temporal information, a video encoding device utilizing a spatiotemporal information-based INVR model is described.

[0103] FIG. 11 is an exemplary diagram illustrating a spatiotemporal information-based INVR model according to one embodiment of the present disclosure.

[0104] The spatiotemporal information-based INVR model utilizes not only the embedding of information about the current frame, but also additionally inputs information about differences between adjacent frames. To utilize this information, the spatiotemporal information-based INVR model includes the same components as the spatiotemporal information-based INVR model, plus 2D DWT modules and 1D DWT modules.

[0105] The encoder (110) performs a 3D discrete wavelet transform using the t-th frame and the adjacent t-1-th and t+1-th frames as follows. The encoder (110) decomposes low-frequency and high-frequency components on the spatial axis using the 2D discrete wavelet transform, and decomposes low-frequency and high-frequency components on the temporal axis using an additional 1D discrete wavelet transform on the 2D transform result. In addition to the embedding for the t-th frame (embedding (t) in FIG. 11), the encoder (110) can generate two additional embeddings (embedding (t-1) and embedding (t+1) in FIG. 11) by downsampling the generated low-frequency components for space and time using a DB. The encoder (110) transfers the embeddings to the decoder (120). The decoder (120) restores the frequency domain signal from the embeddings, and as described above, the restored w t_hat Restore frame V at time t by inverse transforming t_hat Creates.

[0106] Since the 2D INVR model can represent the image to be restored, transmitting the model has the same effect as transmitting the image. That is, by compressing and transmitting the 2D INVR model, an end-to-end compression technique using a neural network can be implemented. Below, a method for compressing an image by encoding the parameters and output embedding of the decoder (120) that generates the image is described.

[0107] FIG. 12 is an exemplary diagram illustrating an encoding pipeline of a 2D INVR model according to one embodiment of the present disclosure.

[0108] The encoding pipeline of the 2D INVR model may include all or part of a video overfitter (910), a model pruner (920), a model quantizer (930), an embedding quantizer (1210), and a parameter encoder (1220), as shown in the example of FIG. 12.

[0109] The video overfitter (910) represents a restored image based on the parameters of the 2D INVR model. The video encoding device trains the INVR model using the aforementioned loss function, thereby allowing the INVR model to represent a video frame. As the loss function, a combination of L1 loss and SSIM loss can be used based on the difference between the restored frame and the GT. In addition to the loss based on the difference between the restored frame and the GT, by additionally calculating the loss associated with the restored 2D wavelet coefficients, high-frequency components can be better represented.

[0110] The model pruner (920) reduces the weight of the INVR model by removing low-importance parameters using pruning. For example, a global unstructured pruning technique can be used, which removes the smallest connections across the entire model. Alternatively, pruning can set parameters smaller than a preset threshold to zero.

[0111] The model quantization unit (930) quantizes parameters to which pruning has been applied (e.g., parameters above a threshold). The embedding quantization unit (1210) quantizes the embedding generated by the encoder (110). The final parameters and embedding of the lightweight INVR model can be quantized into m bits and n bits, respectively (m and n are arbitrary positive integers).

[0112] The parameter encoding unit (1220) can reduce the size of the bitstream by utilizing, for example, Huffman coding or entropy coding for lossless compression of quantized parameters and embeddings. As more parameters are pruned, parameter encoding can be used to effectively compress model parameters by using fewer bits.

[0113] As described above, the image encoding device can reduce the size of the parameters required by the image decoding device by quantizing the embedding and reducing the input spatial size based on the DWT.

[0114] For example, a video encoding device may encode quantization parameters related to the compression of an INVR model as additional information. The video encoding device may also encode information regarding the structure of the INVR model as additional information. As an example, an INVR model parameter set (IPS) may be utilized to convey the additional information. The video encoding device may include the IPS in the bitstream.

[0115] Meanwhile, as a reverse order of the encoding pipeline illustrated in FIG. 12, a decoding pipeline for decoding an INVR model can be illustrated as in FIG. 13. The decoding pipeline can include all or part of a parameter decoder (1310), an embedding dequantizer (1320), a model dequantizer (1020), a model restoration unit (1030), and a video generation unit (1040), as in the example of FIG. 10.

[0116] The parameter decoding unit (1310) generates quantized weights and quantized embedding from the bitstream using, for example, Huffman coding or entropy coding. The embedding dequantization unit (1320) dequantizes the quantized embedding to generate a restored embedding. The model dequantization unit (1020) dequantizes the quantized weights to generate parameters to which pruning has been applied. The model restoration unit (1030) restores the INVR model based on the parameters to which pruning has been applied. The video generation unit (1040) generates a restored video frame according to the restored embedding using the restored INVR model.

[0117] As an example, an image decoding device can decode quantization parameters related to the compression of an INVR model as additional information. The image decoding device can decode information regarding the structure of the INVR model as additional information.

[0118] As an example, in addition to using 2D DWT, other 1D / 2D transforms (e.g., discrete cosine transform (DCT), discrete sine transform (DST), non-separable transform, etc.) can also be utilized. For example, an image encoding device generates transform coefficients based on any transform, and then calculates a specific coordinate value (x) within the transformed block / picture. th , y th ) can be used to separate the transformation coefficients into low-frequency and high-frequency components. The coordinate values ​​of the transformation coefficients are 0 ≤ x ≤ x th , 0 ≤ y ≤ y th In this case, the image encoding device can distinguish the corresponding transform coefficients as low-frequency components and distinguish the remaining transform coefficients as high-frequency components. Here, a specific coordinate value x th , y this a positive integer greater than or equal to 0, and may be preset according to an agreement between the video encoding device and the video decoding device. Alternatively, specific coordinate values ​​may be determined by the video encoding device in terms of rate distortion optimization and signaled to the video decoding device.

[0119] As another example, when a scan is performed on the generated transform coefficients, the image encoding device can distinguish between low-frequency and high-frequency components based on a preset threshold. When scanning starts with the high-frequency component, if an index according to the scan is greater than or equal to the preset threshold, the image encoding device can distinguish the corresponding transform coefficients as low-frequency components and distinguish the remaining transform coefficients as high-frequency components. When scanning starts with the low-frequency component, if an index according to the scan is less than or equal to the preset threshold, the image encoding device can distinguish the corresponding transform coefficients as low-frequency components and distinguish the remaining transform coefficients as high-frequency components. Here, the preset threshold is a positive integer greater than or equal to 0, and can be preset according to an agreement between the image encoding device and the image decoding device. Alternatively, the preset threshold can be determined by the image encoding device in terms of rate distortion optimization and signaled to the image decoding device.

[0120] Hereinafter, a method for encoding or restoring 2D video based on a 2D INVR model is described using the cities of FIGS. 14 and 15.

[0121] FIG. 14 is a flowchart illustrating a method for an image encoding device to encode 2D video according to one embodiment of the present disclosure.

[0122] The video encoding device inputs the current 2D frame into the encoder of the 2D INVR model to generate a compressed embedding (S1400).

[0123] The video encoding device can generate the embedding as follows.

[0124] The video encoding device transforms the current 2D frame based on a preset transformation method to generate low-frequency components and remaining frequency components of the current 2D frame (S1410). Here, the preset transformation method may include a 2D discrete wavelet transform, DCT, DST, or non-separable transformation.

[0125] The image encoding device generates an embedding by applying convolution-based downsampling to low-frequency components (S1412).

[0126] The video encoding device inputs the embedding into the decoder of the 2D INVR model to restore the current 2D frame (S1402).

[0127] The video encoding device can restore the current 2D frame as follows.

[0128] The image encoding device upsamples the embedding to restore low-frequency components of multiple resolutions (S1420).

[0129] The image encoding device restores the low-frequency components of the current 2D frame by fusion of low-frequency components of multiple resolutions (S1422).

[0130] The image encoding device applies 2D convolution to the low-frequency components of the current 2D frame to restore the remaining frequency components of the current 2D frame (S1424).

[0131] The video encoding device inversely transforms the low-frequency component and the remaining frequency components of the current 2D frame based on a preset transformation method (S1426).

[0132] The video encoding device can train the encoder of the 2D INVR model and the decoder of the 2D INVR model end-to-end based on a loss function. Here, the loss function can be defined based on the difference between the input frame of the encoder of the 2D INVR model and the output frame of the decoder of the 2D INVR model.

[0133] The video encoding device encodes decoder parameters of the embedding and 2D INVR model (S1404).

[0134] Additionally, the video encoding device can quantize decoder parameters and embedding based on the quantization parameter. The video encoding device can encode the quantization parameter as additional information.

[0135] FIG. 15 is a flowchart illustrating a method for a video decoding device to restore a 2D video according to one embodiment of the present disclosure.

[0136] The video decoding device decodes decoder parameters and an embedding of a 2D INVR model from a bitstream (S1500). Here, the embedding is a compressed representation of the low-frequency components of the current 2D frame, and is generated and provided by the encoder of the 2D INVR model.

[0137] Additionally, the video decoding device decodes quantization parameters as additional information from the bitstream. The video decoding device can dequantize decoder parameters and embedding using the quantization parameters.

[0138] The video decoding device restores the decoder of the 2D INVR model using decoder parameters (S1502).

[0139] The video decoding device inputs the embedding into the decoder of the 2D INVR model to restore the current 2D frame (S1504).

[0140] The video decoding device can restore the current 2D frame as follows.

[0141] The image decoding device applies convolution-based upsampling to the embedding to restore low-frequency components of multiple resolutions (S1514).

[0142] The image decoding device restores the low-frequency components of the current 2D frame by fusion of low-frequency components of multiple resolutions (S1512).

[0143] The image decoding device applies 2D convolution to the low-frequency components of the current 2D frame to restore the remaining frequency components of the current 2D frame (S1516).

[0144] The image decoding device inversely transforms the low-frequency components and remaining frequency components of the current 2D frame based on a preset transformation method (S1518). Here, the preset transformation method may include a 2D discrete wavelet transform, DCT, DST, or non-separable transformation.

[0145] The encoder of a 2D INVR model and the decoder of a 2D INVR model can be pre-trained end-to-end based on a loss function. The loss function can be defined based on the difference between the input frame of the 2D INVR model's encoder and the output frame of the 2D INVR model's decoder.

[0146] Although the flowchart / timing diagram of this specification describes each process as being executed sequentially, this is merely an illustrative description of the technical idea of ​​one embodiment of the present disclosure. In other words, a person of ordinary skill in the art to which one embodiment of the present disclosure belongs may modify and apply various modifications and variations by changing the order described in the flowchart / timing diagram without departing from the essential characteristics of one embodiment of the present disclosure, or by executing one or more of the processes in parallel. Therefore, the flowchart / timing diagram is not limited to a chronological order.

[0147] It should be understood that the exemplary embodiments described above can be implemented in many different ways. The functions or methods described in one or more examples can be implemented in hardware, software, firmware, or any combination thereof. It should be understood that the functional components described herein are labeled as "units" to further emphasize their implementation independence.

[0148] Meanwhile, the various functions or methods described in this embodiment may also be implemented as instructions stored on a non-transitory storage medium that can be read and executed by one or more processors. Non-transitory storage media include, for example, all types of storage devices that store data in a form readable by a computer system. For example, non-transitory storage media include storage media such as erasable programmable read-only memory (EPROM), flash drives, optical drives, magnetic hard drives, and solid-state drives (SSDs).

[0149] The above description is merely an example of the technical idea of ​​the present embodiment, and those skilled in the art will appreciate that various modifications and variations can be made without departing from the essential characteristics of the present embodiment. Therefore, the present embodiments are not intended to limit the technical idea of ​​the present embodiment, but rather to explain it, and the scope of the technical idea of ​​the present embodiment is not limited by these embodiments. The scope of protection of the present embodiment should be interpreted by the claims below, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of rights of the present embodiment.

[0150]

[0151]

[0152] CROSS-REFERENCE TO RELATED APPLICATION

[0153] This patent application claims priority to Korean patent application No. 10-2023-0176421, filed in Korea on December 7, 2023, and Korean patent application No. 10-2024-0154232, filed in Korea on November 4, 2024, the entire contents of which are incorporated herein by reference.

Claims

1. A method for restoring a current 2D frame performed by an image decoding device, A step of decoding decoder parameters and embedding of a 2D (Dimensional) INVR (Implicit Neural Video Representation) model from a bitstream, wherein the embedding is a compressed representation of low frequency components of the current 2D frame and is generated and provided by an encoder of the 2D INVR model; A step of restoring the decoder of the 2D INVR model using the above decoder parameters; and A step of restoring the current 2D frame by inputting the above embedding into the decoder of the 2D INVR model. Including, The steps for restoring the current 2D frame are: A step of restoring low-frequency components of multi-resolution by applying convolution-based upsampling to the above embedding; and A step of restoring the low-frequency components of the current 2D frame by fusing the low-frequency components of the above multi-resolutions. A method comprising:

2. In paragraph 1, The steps for restoring the current 2D frame are: A step of applying 2D convolution to the low frequency components of the current 2D frame to restore the remaining frequency components of the current 2D frame. A method further comprising:

3. In paragraph 2, The steps for restoring the current 2D frame are: A step of inversely transforming the low frequency component and the remaining frequency components of the current 2D frame based on a preset transformation method. A method further comprising:

4. In paragraph 1, A method wherein the encoder of the 2D INVR model and the decoder of the 2D INVR model are pre-trained end-to-end according to a loss function, and the loss function is defined based on the difference between the input frame of the encoder of the 2D INVR model and the output frame of the decoder of the 2D INVR model.

5. In paragraph 1, A step of decoding quantization parameters as additional information from the above bitstream; A step of dequantizing the decoder parameters and the embedding using the above quantization parameters. A method further comprising:

6. A method for encoding a current 2D frame performed by a video encoding device, A step of inputting the current 2D frame to an encoder of a 2D (Dimensional) INVR (Implicit Neural Video Representation) model to generate a compressed embedding; A step of restoring the current 2D frame by inputting the above embedding into the decoder of the 2D INVR model; and A step of encoding the above embedding and decoder parameters of the above 2D INVR model. Including, The steps for generating the above embedding are: A step of generating a low-frequency component of the current 2D frame by converting the current 2D frame based on a preset conversion method; and A step of generating the embedding by applying convolution-based downsampling to the above low-frequency components. A method comprising:

7. In paragraph 6, The steps for restoring the current 2D frame are: A step of restoring low-frequency components of multi-resolution by applying convolution-based upsampling to the above embedding; and A step of restoring the low-frequency components of the current 2D frame by fusing the low-frequency components of the above multi-resolutions. A method comprising:

8. In paragraph 6, The steps for restoring the current 2D frame are: A step of applying 2D convolution to the low frequency components of the current 2D frame to restore the remaining frequency components of the current 2D frame. A method further comprising:

9. In paragraph 8, The steps for restoring the current 2D frame are: A step of inversely transforming the low frequency component and the remaining frequency components of the current 2D frame based on the above preset conversion method. A method further comprising:

10. In paragraph 6, Further comprising a step of training the encoder of the above 2D INVR model and the decoder of the above 2D INVR model end-to-end based on a loss function, A method wherein the loss function is defined based on the difference between the input frame of the encoder of the 2D INVR model and the output frame of the decoder of the 2D INVR model.

11. In paragraph 6, A step of quantizing the decoder parameters and the embedding based on the quantization parameters; and A step of encoding the above quantization parameters as additional information. A method further comprising:

12. A method for providing video data to a video decoding device, A step of encoding the above video data into a bitstream; and A step of transmitting the above bitstream to the image decoding device Including, The step of encoding the above video data is: A step of inputting the current 2D frame into the encoder of a 2D (Dimensional) INVR (Implicit Neural Video Representation) model to generate a compressed embedding; A step of restoring the current 2D frame by inputting the above embedding into the decoder of the 2D INVR model; and A step of encoding the above embedding and decoder parameters of the above 2D INVR model. Including, The steps for generating the above embedding are: A step of generating a low-frequency component of the current 2D frame by converting the current 2D frame based on a preset conversion method; and A step of generating the embedding by applying convolution-based downsampling to the above low-frequency components. A method, characterized by including:

Citation Information

Patent Citations

  • Training device, estimation device, model generation method, and neural network model

    JP2021182204A

  • Optical atomizer, spraying method using the same, and manufacturing method of the same

    KR1020250026530A

  • Method and device for neural network-based processing in video coding

    KR102124714B1

  • Method for generating high quality video with professional filming techniques through deep learning technology based 3D space modeling and point-of-view synthesis and apparatus for same

    KR102593135B1