Image and video compression using implicit neural characterized learning dictionaries

By partitioning tiles in video images and learning local implicit neural representations, the problem of insufficient elimination of space-time redundancy in the prior art is solved, and more efficient image and video compression effects are achieved.

CN120035844APending Publication Date: 2025-05-23INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072638.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-11
Filing Date
2023-09-28
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing implicit neural characterization (INR) methods have insufficient performance in image and video compression, and it is difficult to effectively eliminate spatiotemporal and spatial redundancy, resulting in insufficient rate distortion.

Method used

通过在视频图像中分区图块,学习局部隐式神经表征,并通过重复使用表征的较大部分来消除时空冗余。具体方法包括确定头部网络和尾部网络的权重,并将尾部层的权重编码为视频数据。

Benefits of technology

The performance of the implicit neural characterization method is improved, making it more suitable in terms of rate distortion, effectively eliminating spatiotemporal redundancy, and improving the efficiency of image and video compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120035844A_ABST
    Figure CN120035844A_ABST
Patent Text Reader

Abstract

Compression and decompression of video images are accomplished by an improved implicit neural representation (INR) method for compression. Specifically, for a smaller image region, the characterization is learned locally and spatio-temporal redundancy is eliminated by reusing a larger portion of the characterization. In one embodiment, a local INR corresponding to a partition of an image / video is calculated for each portion of the partition. A larger portion of the characterization in terms of the weights of the neural network is reused by decomposing each characterization into two portions: a head (preferably including the majority of the weights) and a tail (a lightweight network designed to accommodate the head to a particular image tile). In another embodiment, the larger portion of the characterization is learned over a large number of image tiles to obtain dictionaries known both on the coding side and on the decoding side in order to reduce the amount of data transmitted.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of European application serial number 22306530.1 filed on October 11, 2022, which is incorporated herein by reference in its entirety. Technical Field

[0003] At least one of the present embodiments generally relates to a method or apparatus for compressing images and videos using a neural network (ANN) based tool. Background Art

[0004] Machine learning (ML) has emerged as a new tool to revolutionize compression compared to classical approaches used in standardization. The main idea is to learn the entire compression chain, including content description, quantization, entropy coding, and descriptor decompression. Summary of the invention

[0005] At least one of the present embodiments generally relates to a method or apparatus in the context of using a tool based on a novel neural network (ANN) to compress images and videos. In particular, one purpose of the described embodiments is to improve the implicit neural representation (INR) method for compression. Specifically, the representation is learned locally (for smaller image regions), and spatiotemporal redundancy is eliminated by reusing larger parts of the representation. Local INR will be more suitable in terms of rate distortion.

[0006] According to a first aspect, a method is provided. The method comprises the following steps: partitioning at least a portion of a video image into tiles; determining at least one head network based on global information of at least a portion of the video image; determining a plurality of tail networks based on the tiles of at least a portion of the video image; determining weights corresponding to the head networks; optimizing weights of the tail layers by minimizing an implicit neural representation of the tiles by learning weights of at least one head network; and encoding the weights of the tail layers into video data.

[0007] According to a second aspect, a method is provided. The method comprises the following steps: parsing video data to obtain weights of a tail network for a tile of the video data; and reconstructing the tile of the video data using the weights and an optimized head network.

[0008] According to another aspect, a device is provided. The device includes a processor. The device can be configured to implement the general aspects by executing any of the described methods.

[0009] According to another general aspect of at least one embodiment, there is provided an apparatus comprising: a device according to any one of the decoding embodiments; and at least one of the following: (i) an antenna configured to receive a signal, the signal including a video block, (ii) a band limiter configured to limit the received signal to a band including the video block, or (iii) a display configured to display an output representative of the video block.

[0010] According to another general aspect of at least one embodiment, there is provided a non-transitory computer-readable medium containing data content generated according to any one of the described encoding embodiments or variations.

[0011] According to another general aspect of at least one embodiment, there is provided a signal including video data generated according to any one of the described encoding embodiments or variations.

[0012] According to another general aspect of at least one embodiment, a bitstream is formatted to include data content generated according to any one of the described encoding embodiments or variations.

[0013] According to another general aspect of at least one embodiment, there is provided a computer program product including instructions that, when executed by a computer, cause the computer to perform any one of the described decoding embodiments or variations.

[0014] These and other aspects, features, and advantages of the general aspects will become apparent from the following detailed description of exemplary embodiments read in conjunction with the accompanying drawings.

[0015] According to another general aspect of at least one embodiment, there is provided a non-transitory computer-readable medium containing data content including instructions for performing any one of an encoding method or a decoding method. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 An example image parameterization is shown.

[0017] Figure 2 An example of an implicit neural representation is shown.

[0018] Figure 3 An example of an image with a SIREN network and residual fitting is shown.

[0019] Figure 4 An example of multi-tile learning is shown.

[0020] Figure 5Shows examples of a) ground truth patches, b) patches fitted by a shared network, and c) patches fine-tuned one by one.

[0021] Figure 6 Shows another example of a) ground truth patches, b) patches fitted by a shared network, and c) patches fine-tuned one by one.

[0022] Figure 7 Shows another example of a) ground truth patches, b) patches fitted by a shared network, and c) patches fine-tuned one by one.

[0023] Figure 8 Shows another example of a) ground truth patches, b) patches fitted by a shared network, and c) patches fine-tuned one by one.

[0024] Fig. 9 Shows another example of a) ground truth patches, b) patches fitted by a shared network, and c) patches fine-tuned one by one.

[0025] Fig.10 Shows another example of a) ground truth patches, b) patches fitted by a shared network, and c) patches fine-tuned one by one.

[0026] Fig.11 Shows another example of a) ground truth patches, b) patches fitted by a shared network, and c) patches fine-tuned one by one.

[0027] Fig.12 Shows another example of a) ground truth patches, b) patches fitted by a shared network, and c) patches fine-tuned one by one.

[0028] Fig.13 Shows another example of a) ground truth patches, b) patches fitted by a shared network, and c) patches fine-tuned one by one.

[0029] Fig.14 Shows another example of a) ground truth patches, b) patches fitted by a shared network, and c) patches fine-tuned one by one.

[0030] Fig.15 Shows an embodiment of a method for encoding a video using the described embodiments.

[0031] Fig.16 Shows an embodiment of a method for decoding a video using the described embodiments.

[0032] Fig.17 Shows an embodiment of a device for encoding or decoding using the described embodiments.

[0033] Fig.18 A standard general purpose video compression scheme is shown.

[0034] Fig.19 A standard general purpose video decompression scheme is shown.

[0035] Fig. 20 A processor-based system for encoding / decoding according to the generally described aspects is shown. DETAILED DESCRIPTION

[0036] The context of the described embodiments is compression of image / video content.

[0037] Among the many tools, a new proposal based on Implicit Neural Representation (INR) has recently emerged. The main idea is to represent an image function. For a 2D image, this function maps 2D coordinates to signal RGB values. Once this function is overfitted using a neural network, transmitting the image is equivalent to transmitting the weights of this neural network. INR has been used for various signals, including audio and 3D point clouds.

[0038] Today, even the latest methods, the performance of INR for compression barely reaches the performance of SOTA methods. The described embodiments provide a contribution to improving this performance.

[0039] The described embodiments are directed to improving state-of-the-art INR methods for compression. Specifically, the representation is learned locally (for smaller image regions), and spatiotemporal redundancy is eliminated by reusing larger portions of the representation. Although we have mentioned partitioning video images into "tiles," the current aspects and embodiments are not limited to tiles and may be applicable to any irregular partitioning or segmentation of an image.

[0040] The described embodiments involve finding local INRs that will be more suitable in terms of rate distortion. This amounts to partitioning the image / video to be compressed (regularly or irregularly) and calculating the INR for each part of the partition. In order to exploit spatiotemporal redundancy, a larger part of the representation (in terms of the weights of the neural network) should be reused. More specifically, each representation is decomposed into two parts: a head (preferably containing most of the weights) and a tail (a lightweight network designed to adapt the head to a specific image tile). The described embodiments propose two approaches:

[0041] -Use larger portions of the representation on different parts of the image. We call this multi-patch learning. For a given image, one or more head networks can be used. This collection of head networks can also be used ad hoc.

[0042] -Or, learning this larger part of the representation over a large number of image patches, we get a dictionary for the head network. Since this dictionary is learned, it is known on both the encoder and decoder side, so it is no longer a transmission cost.

[0043] Some background and known technology of INR is first reviewed which is necessary for understanding the aspects and embodiments described herein.

[0044] Implicit representation or coordinate-based representation parameterizes the signal as a continuous function. For a 2D image, a pixel has coordinates (x, y) that describe its position in the image. The goal of the continuous function is to obtain the RGB value of the pixel based on the coordinates of the pixel. Figure 1 A coordinate-based characterization is shown.

[0045] like Figure 1 As shown, based on the coordinate representation, each pixel has coordinates (x, y), such as x∈[0, 1] and y∈[0, 1], and the function f is parameterized by a neural network to obtain RGB values. The implicit representation takes the pixel coordinates as input and parameterizes the function f by a neural network to obtain the corresponding RGB values.

[0046] The exact mathematical formula of this function is unknown, and a neural network is used to approximate the function, which is called an implicit neural representation. Figure 2 Implicit neural representations are shown. Figure 2 An illustration of an implicit neural representation is shown: The neural network takes pixel coordinates as input and gives an RGB image as output. During training, the network will learn θ which represents the weights of the neural network. Basically, the neural network is learning a function that matches pixel coordinates to their RGB values.

[0047] INRs have some benefits, first they are no longer coupled to spatial resolution. This means (the image is now independent of the number of pixels). INRs have "infinite resolution" because they can be sampled at any spatial resolution. Second, since INRs are independent of spatial resolution, the memory requirements scale with the complexity of the underlying signal.

[0048] A major problem with the classic multilayer perceptron (MPL) is that the obtained function is smooth. Therefore, it is difficult to approximate the high frequencies of the signal. To solve this problem, methods using sin activation functions or Fourier embedding have been proposed.

[0049] Consider I as an image fitted with a SIREN (Sinusoidal Representation Network) network. The image can be represented by the following notation I[x,y], where (x,y) represents each pixel coordinate. The SIREN network will return the RGB value at each pixel location (x,y). The SIREN function can be represented by the parameter θ Thus, the pixel position is mapped to the RGB value in the image I[x, y], that is, f θ (x, y) = (r, g, b).

[0050] The goal is to transform f according to some distortion measure θ Fitting to I[x, y], we use the mean squared error, which yields the following optimization problem

[0051]

[0052] where the sum is the sum of all pixels, f 0 (x, y) is the predicted image, and I[x, y] is the ground truth image. Function f θ It is parameterized by a neural network consisting of a linear projection and a sin activation function, and is classically learned by backpropagation of the reconstruction error.

[0053] Figure 3 An example of a reconstruction that can be achieved using Siren INR is provided.

[0054] Figure 3 SIREN fitted on image 1 according to the kodak dataset is shown. The left is the original image. The middle is the reconstructed image obtained by SIREN. The residual image on the right comes from subtracting the reconstructed image from the original image. For a model size of 0.6bpp, the PSNR of the reconstructed image is 22.45dB.

[0055] The Fourier method differs only in the fact that a layer is added after the image coordinates, thereby mapping to a higher-dimensional space with the known Fourier function.

[0056] Multi-tile learning

[0057] INR uses a single function f parameterized by a neural network θ The loss function in Equation (1) is optimized over the entire image. It implicitly captures the redundancy in the image, however, if it is operated on local patches, the redundancy can be better reduced. Therefore, it is proposed to combine the two functions The multi-tile lNR is a neural network that operates on local tiles, where is marked as a head and The head captures global information and its representation is shared among all tiles (or images, or frames in a video), while the tail is adapted to a specific tile (or image, or frame). Let I be an image and it is divided into M local regions and each local region is labeled as a tile. Thus, an image is represented as a collection of tiles, such that I = [P 1 , P 2 , ..., PM ].

[0058] The same decomposition is also applied to a group of frames to take into account temporal redundancy. In the following, reference is made to image decomposition, but it should be understood that the method is applicable to a group of frames without loss of generality.

[0059] Image specific multi-tile INR

[0060] set up is a network consisting of a head network and a tail network. For each image, there is a head and several tails. The weight of the head θ g is shared among all tiles, and each local tile has its own tail function to adapt to the local content. The loss function for minimizing multi-tile INR can be written as:

[0061]

[0062] where i varies over the set index of the tile and (x, y) is the local image coordinate. Figure 4 The network architecture is depicted in .

[0063] Figure 4 Multi-patch learning is shown. The input image is split into tiles, here 6 tiles. All tiles have the same shared network or shared head, and each tile has its own network. The tail network consists of tile-specific networks. The shared head will learn a global representation for all tiles, and the tail network will adapt this global representation to each tile. The MPL network is trained in parallel on all tiles, and the tiles are merged to reconstruct the image.

[0064] The minimization of equation (2) is performed in two steps:

[0065] (1) Jointly learn the weights of the head layer and all tail layers.

[0066] (2) The weights of the head layer are frozen, and only the weights of the tail layer are optimized to adapt to the specific local content of each tile.

[0067] In this way, the first step ensures that the global information present in the image is captured by the head layer and the second step ensures adaptation to the local content of each tile. The weights of the loss function are optimized using gradient descent or stochastic gradient descent methods. Once the weights are optimized, they are further encoded in the bitstream using any entropy encoder technique.

[0068] Image-independent multi-tile INR

[0069] In the previous section, each image had its own specific head, whereas here the head is independent and common to all images. The weights are the weights of the head learned over a large number of patches. Let [P 1 , P 2 , ..., P N ] is a large number of different N tiles. The head layer is optimized with respect to the following loss function, where Composed into a single function (network)

[0070]

[0071] Once the weights of the head network are trained, for any given image to be encoded, the image is divided into local tiles and the tail network is added in addition to the optimized head network (frozen) to adapt to the local content of the tile and optimize the weights of the tail network. The weights of the tail network can also be adjusted by adding L to the loss 1 Regularization is forced to be sparse, as in Equation (4). Since the head network is independent of the image, which is known to both the encoding and decoding sides, only the weights of the tail network need to be transmitted.

[0072]

[0073] Without loss of generality, this applies to temporal encoding of video sequences, where a head is learned on the first frame of the sequence and used for the remaining frames.

[0074] Dictionary learning

[0075] Here, the goal is to learn a set of head layers. Let D = [d 1 , d 2 ,...d K ] is the dictionary to be learned, and each element d in the dictionary k are the weights of the head layer. Once we have this dictionary, for a given image, we select the best head layer from D.

[0076] The dictionary D is learned from a large number of image patches in two ways

[0077] The tiles are clustered into K clusters, and for each cluster, the weights of the head layer are learned. Thus, each cluster can represent a different information content of the image. For clustering, any existing clustering technique can be used. Clustering can be performed on the image domain (pixels) or can be computed on features extracted from off-the-shelf deep neural networks.

[0078] According to the set of head layers in the image-specific multi-tile INR description, they are clustered to learn K elements of the dictionary.

[0079] During encoding, the selection of the header (dictionary) can be performed in several ways:

[0080] In the first case, the selection of the head layer may be performed by calculating the distance between the image patch and the cluster center, and the cluster center with the smallest distance is selected as the head layer. The distance may be the mean squared error.

[0081] A head layer having the smallest loss between the reconstructed image and the original image may be selected as the head layer.

[0082] We train a classifier that takes as input a tile and outputs the index of the head to be selected.

[0083] In the bitstream, in addition to the weights, the index to the dictionary used as the header layer is also encoded.

[0084] Encoding the weights

[0085] Once the weights are optimized as described in the section on multi-tile INR, they are encoded in the bitstream. To encode the weights, this can be done in the following way:

[0086] (1) The weights can be encoded with half precision (16 bits).

[0087] (2) The weights can be encoded using maximum absolute value normalization and 8-bit quantization.

[0088] (3) The weights can be encoded using an explicit probability distribution by estimating the parameters of the distribution from the weights. For example, for a Gaussian distribution, the mean and variance are estimated from the weights.

[0089] (4) The weights can be encoded with 8-bit quantization using maximum absolute value normalization and boundary-aware entropy model.

[0090] (5) Equations (2) and (4) can also include an entropy model, so the weights are minimized using the entropy model. In this case, equation (2) becomes

[0091]

[0092] And equation (4) becomes

[0093]

[0094] During training, uniform noise is added to the weights to approximate the quantization error during test (encoding) time. During encoding, the weights are quantized to the nearest integer and encoded into the bitstream via the learned CDF of p(.).

[0095] result

[0096] Multi-tile INR experiment on Kodak dataset images

[0097] A multi-tile INR experiment is performed using the FOUREN network. The goal here is to reconstruct an image with a small tail and send the weights of the tail to the receiver side. The encoder and decoder already know the header. The goal is to find the best network architecture to reconstruct an image with a tail at 0.68bpp.

[0098] Multi-tile INR with FOUREN is implemented in PyTorch and all experiments are performed on a single Tesla M60 GPU with 8GB RAM. There are two steps to train the model. First, train our network on all tiles. This will give a global representation over all tiles. Then, fine-tune each tile one by one to get a better representation for each tile.

[0099] The technique used for this training is called freezing. In the context of neural networks, freezing a layer is about controlling the way updates to weights are made. Freezing a layer means that its weights cannot be modified further. In this case, during training, the layers of the head are frozen. The tail has the same number of subnetworks as the number of tiles. For training a specific tile, the other subnetworks are also frozen. This means that only the weights of the current tile are updated.

[0100] In the following experiments, the goal is to find the best architecture that sends less information to the decoder. We will work with a tail size at 0.68bpp and put all the network complexity in the head.

[0101] Multi-tile INR applied to a collection of tiles

[0102] Select 10 tiles from some images of the Kodak dataset. Use the shared network and fine-tune the network one by one to reconstruct these tiles. To train our network, some parameters are introduced: tile size 64×64 pixels, map size 256. The network has been trained with 2 -4 The learning rate was trained for 15K iterations.

[0103] Figure 5 Ground truth patches are shown: 10 patches of size 64*64 pixels retrieved from some kodak dataset images. (a): All original patches. (b): Patches fitted by the shared network. Our network is trained with 20 head layers and 1 tail layer, where the layer width is 20 nodes and the learning rate is 2. -4 , and perform 15K iterations. (c): Fine-tuning the tiles one by one, with 0.5*2 for each tile -4 We retrain our network with a learning rate of 0.5*15K iterations.

[0104] exist Figure 5 In , it can be seen that fitting the tiles with the same shared network introduces some artifacts in the tiles, since the network is learning to reconstruct the tiles simultaneously with the same representation. We then fine-tune on a tile-by-tile basis to get better reconstructions of the tiles, since the weights of each tail are updated again.

[0105] Table 1 shows more experimental results based on the network architecture. Each network has a tail layer at 0.68bpp. By changing the number of head layers, some changes in the network performance can be seen. When the head is too small or too large, the network performance is low. We found that a head size of about 40bpp to 45bpp gives a good PSNR average for 10 tiles. The next step will be to apply this approach to the entire image from the kodak dataset.

[0106] Table 1: Network performance evaluated on 10 previous tiles for different network architectures. NL_head: number of head layers, NL_Tail: number of tail layers, num_Layers: total number of layers (head+tail), Head_node: width of each head layer, tail_node: width of each tail layer, Mapping_size: Fourier feature map size, mean_PSNR: average PSNR of the shared network over all tiles, mean_PSNR_fine_tuned: average PSNR over all tiles after tile-by-tile fine-tuning, bpp_head: bpp of the head, bpp_tail: bpp of the tail.

[0107]

[0108] Multi-tile INR applied to the complete image

[0109] Following the same idea of ​​applying MPL to a collection of tiles, the same approach will be applied to images. The images are divided into tiles of equal size. The experiments are performed on three different tile sizes 64*64, 128*128, and 256*256. The kodak dataset images have the same size of 768*512 pixels, which gives 96 tiles for tile size 64, 24 tiles for tile size 128, and 6 tiles for tile size 256. In theory, the network will learn a global representation that matches all tiles. It is desirable to keep the tail bbp as small as possible, and 0.68bpp seems to be a good value for the tail. In these experiments, the images are reconstructed by varying the head size to find the best architecture for the image representation.

[0110] Multi-tile INR at low head size

[0111] Our first MPL network is built based on a tile size of 64*64. The model has 10 head layers with a layer width of 64 and 1 tail layer with a layer width of 28. The head size is 5.53bpp and the tail size is 0.68bpp. Fourier feature map size = 256, with 2 -4 Our network is trained for 10k iterations with a learning rate of . The network is given the reconstructed image from the shared network and fine-tuned on the reconstructed image tile by tile. Figure 6 , Figure 7 and Figure 8 The results in show that this network architecture performs poorly because the network learns on 96 tiles, which is too small to handle. PSNR values ​​are below 30dB and the head size needs to be increased, which means a larger tile size is needed.

[0112] Figure 6 The MPL fitted on image 1 according to the kodak dataset is shown. The head size of the network is at 5.53bpp and the tail size is at 0.68bpp. On the left is the original image. In the middle is the reconstructed image obtained by sharing the network. On the right is the reconstructed image obtained by fine-tuning the network tile by tile. The shared network reconstructs the image with a PSNR of 21.09dB. The network is fine-tuned tile by tile and a PSNR of 21.10dB is obtained.

[0113] Figure 7 The MPL fitted on image 15 according to the kodak dataset is shown. The head size of the network is at 5.53bpp and the tail size is at 0.68bpp. On the left is the original image. In the middle is the reconstructed image obtained by sharing the network. On the right is the reconstructed image obtained by fine-tuning the network tile by tile. The shared network reconstructs the image with a PSNR of 26.30dB. The network is fine-tuned tile by tile and a PSNR of 26.33dB is obtained.

[0114] Figure 8 The MPL fitted on image 24 according to the kodak dataset is shown. The head size of the network is at 5.53bpp and the tail size is at 0.68bpp. On the left is the original image. In the middle is the reconstructed image obtained by sharing the network. On the right is the reconstructed image obtained by fine-tuning the network tile by tile. The shared network reconstructs the image with a PSNR of 21.74dB. The network is fine-tuned tile by tile and a PSNR of 21.75dB is obtained.

[0115] Multi-tile INR at medium head size

[0116] Our second MPL network is built based on a tile size of 128*128. The model has 8 head layers with a layer width of 160 and 1 tail layer with a layer width of 110. Now, the head size is 20.70bpp and the tail size is 0.65bpp. Fourier feature map size = 256, with 2 -4 We train our network for 10k iterations with a learning rate of . From Figures 6.9, 6.10, and 6.11, we can see that the PSNR value is about 30dB, which is an acceptable quality reconstruction, as we only have 24 tiles and the head size is large enough.

[0117] Fig. 9 The MPL fitted on image 1 according to the kodak dataset is shown. The head size of the network is at 20.70bpp and the tail size is at 0.65bpp. On the left is the original image. In the middle is the reconstructed image obtained by sharing the network. On the right is the reconstructed image obtained by fine-tuning the network tile by tile. The shared network reconstructs the image with a PSNR of 29.14dB. The network is fine-tuned tile by tile and a PSNR of 29.22dB is obtained.

[0118] Fig.10 The MPL fitted on image 15 according to the kodak dataset is shown. The head size of the network is at 20.70bpp and the tail size is at 0.65bpp. On the left is the original image. In the middle is the reconstructed image obtained by sharing the network. On the right is the reconstructed image obtained by fine-tuning the network tile by tile. The shared network reconstructs the image with a PSNR of 35.44dB. The network is fine-tuned tile by tile and a PSNR of 35.59dB is obtained.

[0119] Fig.11 The MPL fitted on image 24 according to the kodak dataset is shown. The head size of the network is at 20.70bpp and the tail size is at 0.65bpp. On the left is the original image. In the middle is the reconstructed image obtained by sharing the network. On the right is the reconstructed image obtained by fine-tuning the network tile by tile. The shared network reconstructs the image with a PSNR of 30.44dB. The network is fine-tuned tile by tile and a PSNR of 30.61dB is obtained.

[0120] Multi-tile INR at high head size

[0121] We want to get a quasi-perfect reconstruction of the image with a high head size. The MPL network is built based on a tile size of 256*256. The model has 5 head layers with a layer width of 512 and 1 tail layer with a layer width of 450. Now, the head size is 104.29bpp and the tail size is 0.66bpp. Fourier feature map size = 256, with 2-4 We train our network for 10k iterations with a learning rate of Figure 6 As can be seen in .12, 6.13 and 6.14, the high header size gives a PSNR value of about 40dB, which is a perfect reconstruction quality of the image.

[0122] Fig.12 The MPL fitted on image 1 according to the kodak dataset is shown. The head size of the network is at 104.29bpp and the tail size is at 0.66bpp. On the left is the original image. In the middle is the reconstructed image obtained by sharing the network. On the right is the reconstructed image obtained by fine-tuning the network tile by tile. The shared network reconstructs the image with a PSNR of 40.12dB. The network is fine-tuned tile by tile and a PSNR of 40.21dB is obtained.

[0123] Fig.13 The MPL fitted on image 15 according to the kodak dataset is shown. The head size of the network is at 104.29bpp and the tail size is at 0.66bpp. On the left is the original image. In the middle is the reconstructed image obtained by sharing the network. On the right is the reconstructed image obtained by fine-tuning the network tile by tile. The shared network reconstructs the image with a PSNR of 40.95dB. The network is fine-tuned tile by tile and a PSNR of 41.14dB is obtained.

[0124] Fig.14 The MPL fitted on image 24 according to the kodak dataset is shown. The head size of the network is at 104.29bpp and the tail size is at 0.66bpp. On the left is the original image. In the middle is the reconstructed image obtained by sharing the network. On the right is the reconstructed image obtained by fine-tuning the network tile by tile. The shared network reconstructs the image with a PSNR of 39.73dB. The network is fine-tuned tile by tile and a PSNR of 39.99dB is obtained.

[0125] exist Fig.15An embodiment of a method 1500 for encoding video data is shown in FIG. The method starts at start block 1501 and continues to block 1510 for partitioning at least a portion of a video image into tiles. Control continues from block 1510 to block 1520 for determining at least one head network based on global information of the at least a portion of the video image. Control continues from block 1520 to block 1530 for determining a plurality of tail networks based on the tiles of the at least a portion of the video image. Control continues from block 1530 to block 1540 for determining weights corresponding to the head networks. Control continues from block 1540 to block 1550 for optimizing weights of a tail layer by minimizing an implicit neural representation of the tile by learning weights of at least one head network. Control continues from block 1550 to block 1560 for encoding the weights of the tail layer into video data.

[0126] exist Fig.16 One embodiment of a method 1600 for decoding video data is shown in FIG. The method starts at start block 1601 and continues to block 1610 to parse the video data to obtain weights of a tail network for a tile of the video data. Control continues from block 1610 to block 1620 for reconstructing a tile of the video data using the weights and the optimized head network.

[0127] Fig.17 An embodiment of a device 1700 for compressing, encoding or decoding a video using the above method is shown. The device includes a processor 1710 and can be interconnected with a memory 1720 through at least one port. Both the processor 1710 and the memory 1720 can also have one or more additional interconnections to connect to the outside.

[0128] The processor 1710 is also configured to insert or receive information in a bit stream and compress, encode or decode using the above-mentioned methods.

[0129] Embodiments described herein include various aspects, including tools, features, embodiments, models, methods, etc. Many aspects in these aspects are described as having specificity, and at least in order to show each feature, are usually described in a manner that may sound restrictive. However, this is for the purpose of describing clarity, and does not limit the application or scope of these aspects. In fact, all different aspects can be combined and interchanged to provide other aspects. In addition, these aspects can also be combined and interchanged with the aspects described in previous files.

[0130] The aspects described and contemplated in this application can be implemented in many different forms. Fig.18 , Fig.19 and Fig. 20Some embodiments are provided, but other embodiments are contemplated and are not intended to be construed as Fig.18 , Fig.19 and Fig. 20 The discussion does not limit the breadth of implementations. At least one aspect generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated bitstream or an encoded bitstream. These and other aspects can be implemented as methods, devices, computer-readable storage media having stored thereon instructions for encoding or decoding video data according to any of the methods described, and / or computer-readable storage media having stored thereon bitstreams generated according to any of the methods described.

[0131] In this application, the terms "reconstruction" and "decoding" can be used interchangeably, the terms "pixel" and "sample" can be used interchangeably, and the terms "image", "picture" and "frame" can be used interchangeably. Usually, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" or "reconstruction" is used on the decoder side.

[0132] Various methods are described herein, and each method in these methods includes one or more steps or actions for realizing the method.Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined.Additionally, in various embodiments, terms such as "first", "second" can be used to modify elements, parts, steps, operations, etc., such as, for example, "first decoding" and "second decoding".Unless specifically required, the use of such terms does not imply that the modified operations are sorted.Therefore, in this example, the first decoding does not need to be performed before the second decoding, but can occur in a time period overlapping with the second decoding, such as before, during, or during the second decoding.

[0133] The various methods and other aspects described in this application can be used to modify Fig.18 and Fig.19 Modules of the video encoder 100 and decoder 200 shown, for example, intra prediction modules, entropy coding modules and / or decoding modules (160, 360, 145, 330). In addition, the present aspects are not limited to VVC or HEVC, and can be applied to, for example, other standards and recommendations (whether previously existing or developed in the future), as well as extensions of any such standards and recommendations (including VVC and HEVC). Unless otherwise specified or technically excluded, the aspects described in the present application can be used alone or in combination.

[0134] Various numerical values ​​are used in this application. The specific values ​​are for illustrative purposes, and the described aspects are not limited to these specific values.

[0135] Fig.18An encoder 100 is shown. Variations of this encoder 100 are contemplated, but for clarity, the encoder 100 is described below without describing all contemplated variations.

[0136] Before being encoded, the video sequence may undergo a pre-encoding process (101), for example, applying a color transform to an input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of input picture components in order to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and appended to the bitstream.

[0137] In encoder 100, a picture is encoded by encoder elements as described below. The picture to be encoded is partitioned (102) and processed in units such as CUs. Each unit is encoded using, for example, intra mode or inter mode. When the unit is encoded in intra mode, intra prediction (160) is performed. In inter mode, motion estimation (175) and compensation (170) are performed. The encoder decides (105) which mode, intra mode or inter mode, to use to encode the unit, and indicates the intra / inter decision by, for example, a prediction mode flag. For example, a prediction residual is calculated by subtracting (110) the predicted block from the original image block.

[0138] The prediction residual is then transformed (125) and quantized (130). The quantized transform coefficients, along with motion vectors and other syntax elements, are entropy encoded (145) to output a bitstream. The encoder may skip the transform and apply quantization directly to the untransformed residual signal. The encoder may bypass both the transform and quantization, i.e., directly encode the residual without applying the transform or quantization process.

[0139] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (140) and inverse transformed (150) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (155) to reconstruct the image block. An in-loop filter (165) is applied to the reconstructed picture to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (180).

[0140] Fig.19 2 shows a block diagram of a video decoder 200. In the decoder 200, the bitstream is decoded by the decoder elements as described below. The video decoder 200 generally performs the same operations as described above. Fig.18 The encoding process shown is the reverse of the decoding process. Encoder 100 also typically performs video decoding as part of encoding the video data.

[0141] In particular, the input to the decoder includes a video bitstream, which may be generated by the video encoder 100. The bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors and other encoding information. Picture partition information indicates how the picture is partitioned. Thus, the decoder can divide (235) the picture according to the decoded picture partition information. The transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (255) to reconstruct the image block. The prediction block can be obtained (270) from intra-frame prediction (260) or motion compensated prediction (i.e., inter-frame prediction) (275). An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (280).

[0142] The decoded picture may also undergo post-decoding processing (285), such as an inverse color transform (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping, which performs the inverse of the remapping process performed in the pre-encoding process (101). The post-decoding processing may use metadata derived in the pre-encoding process and signaled in the bitstream.

[0143] Fig. 20 A block diagram of an example of a system in which various aspects and embodiments are implemented is shown. System 1000 may be embodied as a device including various components described below, and is configured to perform one or more aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, and servers. The elements of system 1000 may be embodied in a single integrated circuit (IC), multiple ICs, and / or discrete components, either individually or in combination. For example, in at least one embodiment, the processing elements and encoder / decoder elements of system 1000 are distributed over multiple ICs and / or discrete components. In various embodiments, system 1000 is communicatively connected to one or more other systems or other electronic devices via, for example, a communication bus or by dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more aspects described in this document.

[0144] The system 1000 includes at least one processor 1010 configured to execute instructions loaded therein to implement, for example, various aspects described in this document. The processor 1010 may include embedded memory, input and output interfaces, and various other circuits known in the art. The system 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). The system 1000 includes a storage device 1040, which may include a non-volatile memory and / or a volatile memory, including but not limited to an electrically erasable programmable read-only memory (EEPROM), a read-only memory (ROM), a programmable read-only memory (PROM), a random access memory (RAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), a flash memory, a magnetic disk drive, and / or an optical disk drive. As a non-limiting example, the storage device 1040 may include an internal storage device, an attached storage device (including a removable and non-removable storage device), and / or a network accessible storage device.

[0145] The system 1000 includes an encoder / decoder module 1030, which is configured to process data, for example, to provide encoded video or decoded video, and the encoder / decoder module 1030 may include its own processor and memory. The encoder / decoder module 1030 represents a module that can be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of the encoding and decoding modules. Additionally, the encoder / decoder module 1030 may be implemented as a separate element of the system 1000, or may be combined within the processor 1010 as a combination of hardware and software known to those skilled in the art.

[0146] Program code to be loaded onto the processor 1010 or the encoder / decoder 1030 to perform various aspects described in this document may be stored in the storage device 1040 and subsequently loaded onto the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 may store one or more of the various items during the execution of the processes described in this document. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from equations, formulas, operations, and operational logic processing.

[0147] In some embodiments, memory internal to the processor 1010 and / or encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory may be the memory 1020 and / or the storage device 1040, such as a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store, for example, an operating system for the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG refers to Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, i.e., a new standard developed by JVET (Joint Video Experts Group)).

[0148] Input to the elements of system 1000 may be provided through various input devices, as indicated in block 1130. Such input devices include, but are not limited to, (i) a radio frequency (RF) portion that receives an RF signal transmitted over the air, for example, by a broadcaster, (ii) a component (COMP) input terminal (or a set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Fig. 20 Other examples not shown include composite video.

[0149] In various embodiments, the input device of block 1130 has associated corresponding input processing elements known in the art. For example, the RF portion may be associated with elements suitable for the following operations: (i) selecting a desired frequency (also referred to as selecting a signal, or band limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower frequency band to select a signal frequency band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired packet stream. The RF portion of various embodiments includes one or more elements that perform these functions, such as a frequency selector, a signal selector, a frequency band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF portion may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a near-baseband frequency) or baseband. In a set-top box embodiment, the RF part and its associated input processing element receive the RF signal transmitted by wired (for example, cable) medium, and filter to the desired frequency band again by filtering, down-conversion and perform frequency selection. Various embodiments rearrange the order of above-mentioned (and other) elements, remove some and / or add other elements of similar or different functions in these elements. Adding element can include inserting element between existing elements, for example, such as inserting amplifier and analog-to-digital converter. In various embodiments, the RF part includes antenna.

[0150] Additionally, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 1000 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, within a separate input processing IC or within the processor 1010 as desired. Similarly, aspects of USB or HDMI interface processing may be implemented, for example, within a separate interface IC or within the processor 1010 as desired. The demodulated, error corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 1010 and an encoder / decoder 1030 operating in conjunction with memory and storage elements, to process the data streams as desired for presentation on an output device.

[0151] The various components of system 1000 may be disposed in an integrated housing in which the various components may be interconnected and transmit data between them using suitable connection means (e.g., internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards).

[0152] The system 1000 includes a communication interface 1050 that enables communication with other devices via a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to transmit and receive data through the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or a network card, and the communication channel 1060 may be implemented, for example, within a wired and / or wireless medium.

[0153] In various embodiments, a wireless network (such as a Wi-Fi network, for example, IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)) is used to stream or otherwise provide data to the system 1000. The Wi-Fi signal of these embodiments is received by a communication channel 1060 and a communication interface 1050 suitable for Wi-Fi communication. The communication channel 1060 of these embodiments is usually connected to an access point or router, which provides access to external networks including the Internet to allow streaming applications and other excessive communications. Other embodiments use a set-top box to provide streaming data to the system 1000, which delivers data through the HDMI connection of the input block 1130. Other other embodiments use the RF connection of the input block 1130 to provide streaming data to the system 1000. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use a wireless network other than Wi-Fi, for example, a cellular network or a Bluetooth network.

[0154] The system 1000 can provide output signals to various output devices, including a display 1100, a speaker 1110, and other peripheral devices 1120. The display 1100 of various embodiments includes, for example, one or more of a touch screen display, an organic light emitting diode (OLED) display, a curved display, and / or a foldable display. The display 1100 can be used for a television, a tablet computer, a laptop computer, a mobile phone, or another device. The display 1100 can also be integrated with other components (for example, as in a smartphone), or be separate (for example, an external monitor of a laptop computer). In various examples of embodiments, other peripheral devices 1120 include one or more of a stand-alone digital video disc (or digital versatile disc) (DVR, for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 1120 that provide functions based on the output of the system 1000. For example, a disc player performs the function of playing the output of the system 1000.

[0155] In various embodiments, control signals are transmitted between the system 1000 and the display 1100, speaker 1110, or other peripheral device 1120 using signaling (such as AV.Link, , Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention. Output devices may be communicatively coupled to the system 1000 via dedicated connections through respective interfaces 1070, 1080, and 1090. Alternatively, output devices may be connected to the system 1000 using a communication channel 1060 via a communication interface 1050. The display 1100 and speaker 1110 may be integrated into a single unit with other components of the system 1000 in an electronic device (e.g., such as a television). In various embodiments, the display interface 1070 includes a display driver, such as, for example, a timing controller (T Con) chip.

[0156] For example, if the RF input portion 1130 is part of a separate set-top box, the display 1100 and the speaker 1110 may alternatively be separate from one or more of the other components. In various embodiments where the display 1100 and the speaker 1110 are external components, the output signals may be provided via dedicated output connections including, for example, an HDMI port, a USB port, or a COMP output.

[0157] The embodiments may be performed by computer software implemented by the processor 1010 or by hardware or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. As a non-limiting example, the memory 1020 may be of any type suitable for the technical environment and may be implemented using any appropriate data storage technology, such as an optical memory device, a magnetic memory device, a semiconductor-based memory device, a fixed memory, and a removable memory. As a non-limiting example, the processor 1010 may be of any type suitable for the technical environment and may cover one or more microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0158] Various implementations involve decoding. "Decoding" as used in this application may encompass, for example, all or part of a process performed on a received coded sequence to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also include or alternatively include processes performed by an encoder of the various implementations described in this application.

[0159] As another example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Based on the context of the specific description, it will be clear whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to a broader decoding process, and it is believed that those skilled in the art will be well understood.

[0160] Various implementations involve encoding. In a manner similar to the discussion above about "decoding", "encoding" as used in this application may include, for example, all or part of the processing performed on an input video sequence to produce an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, such as partitioning, differential encoding, transforms, quantization, and entropy encoding. In various embodiments, such processes also include or alternatively include processes performed by a decoder of the various implementations described in this application.

[0161] As another example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of entropy encoding and differential encoding. Based on the context of the specific description, it will be clear whether the phrase "encoding process" is intended to refer specifically to a subset of operations or to a broader encoding process, and it is believed that those skilled in the art will be well understood.

[0162] It should be noted that the syntactic elements used herein are descriptive terms. Therefore, they do not exclude the use of other syntactic element names.

[0163] When a figure is presented as a flow chart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow chart of the corresponding method / device.

[0164] Various embodiments may relate to parameter models and rate-distortion optimization. In particular, during the encoding process, a balance or trade-off between rate and distortion is generally considered, which is generally in view of constraints on computational complexity. It can be measured by a rate-distortion optimization (RDO) metric or by a least mean square (LMS), mean absolute error (MAE) or other such measurements. Rate-distortion optimization is generally expressed as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. There are different approaches to solving the rate-distortion optimization problem. For example, these methods can be based on extensive testing of all coding options, including all considered modes or coding parameter values, where the coding cost and the associated distortion of the reconstructed signal are fully evaluated after encoding and decoding. Faster methods can also be used to save coding complexity, particularly to calculate approximate distortion based on prediction or prediction residual signals rather than reconstructed signals. A mixture of these two methods can also be used, such as by using approximate distortion only for some possible coding options and using full distortion for other coding options. Other methods only evaluate a subset of possible coding options. More generally, many approaches employ any of a variety of techniques to perform optimization, but the optimization is not necessarily a comprehensive assessment of the coding cost and associated distortion.

[0165] The implementations and aspects described herein may be implemented in, for example, a method or process, a device, a software program, a data stream, or a signal. Even if discussed only in the context of a single implementation (e.g., discussed only as a method), the implementation of the features discussed may also be implemented in other forms (e.g., a device or program). The device may be implemented, for example, with appropriate hardware, software, and firmware. The method may be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, for example, such as a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate information communication between end users.

[0166] Reference to "one embodiment" or "an embodiment" or "one implementation" or "an implementation" and other variations thereof means that a particular feature, structure, characteristic, etc. described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation" and any other variations appearing in various places throughout this application are not necessarily all referring to the same embodiment.

[0167] Additionally, the present application may refer to “determining” various information. Determining information may include one or more of the following: for example, estimating information, calculating information, predicting information, or retrieving information from a memory.

[0168] Furthermore, the present application may refer to "accessing" various information. Accessing information may include one or more of: for example, receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0169] Additionally, the present application may involve "receiving" various information. Like "accessing," receiving is intended to be a broad term. Receiving information may include one or more of: for example, accessing information or retrieving information (e.g., retrieving from a memory). Furthermore, "receiving" is often involved in one way or another during an operation, such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0170] It should be understood that, for example, in the case of "A / B", "A and / or B", and "at least one of A and B", the use of any of the following " / ", "and / or", and "at least one" is intended to cover selecting only the first listed option (A), or selecting only the second listed option (B), or selecting both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to include selecting only the first listed option (A), or selecting only the second listed option (B), or selecting only the third listed option (C), or selecting only the first and second listed options (A and B), or selecting only the first and third listed options (A and C), or selecting only the second and third listed options (B and C), or selecting all three options (A and B and C). As will be apparent to those of ordinary skill in this and related arts, this can be extended to many of the listed items.

[0171] In addition, as used herein, the word "signaling (signal)" is especially directed to the corresponding decoder to indicate something. For example, in some embodiments, the encoder signals a specific one of multiple transformations, coding modes or flags. In this way, in an embodiment, the same transformation, parameter or mode is used on both the encoder side and the decoder side. Therefore, for example, the encoder can transmit (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. On the contrary, if the decoder already has specific parameters and other parameters, signaling can be used without transmission (implicit signaling) to simply allow the decoder to know and select specific parameters. By avoiding the transmission of any actual function, bit saving is achieved in various embodiments. It should be understood that signaling can be implemented in a variety of ways. For example, in various embodiments, one or more syntactic elements, flags, etc. are used to send information to the corresponding decoder. Although the verb form of the word "signaling" is involved above, the word "signal (signal)" can also be used as a noun in this article.

[0172] It will be apparent to one of ordinary skill in the art that implementations may generate a variety of signals formatted to carry information that may be stored or transmitted, for example. The information may include, for example, instructions for executing a method, or data generated by one of the described implementations. For example, a signal may be formatted to carry a bitstream of the described embodiments. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using a radio frequency portion of a spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog information or digital information. As is well known, a signal may be transmitted over a variety of different wired or wireless links. The signal may be stored on a processor readable medium.

[0173] The previous section describes many embodiments across various claim categories and types. The features of these embodiments may be provided individually or in any combination. In addition, embodiments may include one or more of the following features, devices, or aspects, individually or in any combination, across various claim categories and types:

[0174] At least one embodiment includes encoding and decoding video information using a neural network.

[0175] At least one embodiment includes determining, for a video image, weights corresponding to a tail network.

[0176] At least one embodiment includes using weights in conjunction with a head network for compression / decompression of video images.

[0177] At least one embodiment includes building a dictionary for a head layer of a neural network.

[0178] At least one embodiment includes performing clustering of patches in an image domain or on features extracted from a deep neural network to learn head layer weights for each cluster.

[0179] At least one embodiment comprises a bitstream or signal including one or more of the described syntax elements or variations thereof.

[0180] At least one embodiment includes a bitstream or signal including syntax conveying information generated according to any of the described embodiments.

[0181] At least one embodiment comprises creating and / or transmitting and / or receiving and / or decoding according to any of the described embodiments.

[0182] At least one embodiment includes parsing video data or a bitstream to determine an operating point of a codec.

[0183] At least one embodiment includes a method, process, apparatus, medium storing instructions, medium storing data, or signal according to any of the described embodiments.

[0184] At least one embodiment comprises inserting a syntax element in the signaling that enables a decoder to determine decoding information in a manner corresponding to that used by an encoder.

[0185] At least one embodiment includes creating and / or transmitting and / or receiving and / or decoding a bitstream or signal including one or more of the described syntax elements or variations thereof.

[0186] At least one embodiment includes a TV, set-top box, mobile phone, tablet or other electronic device that performs a transformation method according to any of the described embodiments.

[0187] At least one embodiment includes a TV, set-top box, cell phone, tablet or other electronic device that performs a transformation method determination according to any of the described embodiments and displays (eg, using a monitor, screen or other type of display) the resulting image.

[0188] At least one embodiment includes: a TV, set-top box, mobile phone, tablet or other electronic device that selects, band limits or tunes a channel (e.g., using a tuner) to receive a signal including an encoded image and performs a transformation method according to any of the described embodiments.

[0189] At least one embodiment includes a TV, set-top box, cell phone, tablet or other electronic device that receives over the air (eg, using an antenna) a signal including an encoded image and performs the transformation method.

Claims

1. A method, include: partitioning at least a portion of the video image into tiles; determining at least one head network based on the global information of the at least a portion of the video image; determining a plurality of tail networks based on the tiles of the at least a portion of the video image; determining weights corresponding to the head network; optimizing weights of a tail layer by minimizing an implicit neural representation of the patch via learning weights of the at least one head network; as well as The weights of the tail layer are encoded as video data.

2. A device, include: Memory, and A processor configured to execute: partitioning at least a portion of the video image into tiles; determining at least one head network based on the global information of the at least a portion of the video image; determining a plurality of tail networks based on the tiles of the at least a portion of the video image; determining weights corresponding to the head network; optimizing weights of a tail layer by minimizing an implicit neural representation of the patch via learning weights of the at least one head network; as well as The weights of the tail layer are encoded as video data.

3. A method, include: Parsing the video data to obtain weights of a tail network for a tile of the video data; A tile of the video data is reconstructed using the weights and the optimized head network.

4. A device, include: Memory, and A processor configured to execute: Parsing the video data to obtain weights of a tail network for a tile of the video data; A tile of the video data is reconstructed using the weights and the optimized head network.

5. A method according to any one of claims 1 or 3, or an apparatus according to any one of claims 2 or 4, wherein the head network and the tail network are part of a neural network.

6. A method according to any one of claims 1, 3 or 5, or an apparatus according to any one of claims 2, 4 or 5, wherein minimization of a loss function of an implicit neural representation is used.

7. A method according to any one of claims 1 or 3 or 5 to 6, or an apparatus according to any one of claims 2 or 4 or 5 to 6, wherein the portion of video data is part of a sequence of frames.

8. A method according to any one of claims 1 or 5 to 7, or an apparatus according to any one of claims 2 or 5 to 7, wherein the head network is common to all images.

9. A method according to any one of claims 1 or 5 to 8, or an apparatus according to any one of claims 2 or 5 to 8, wherein the head network is independent of the video image.

10. A method according to any one of claims 1 or 5 to 9, or an apparatus according to any one of claims 2 or 5 to 9, wherein the head network is learned on a first frame of a sequence of video data and used for subsequent video data.

11. A device, include: An apparatus according to any one of claims 4 or 5 to 10; as well as At least one of: (i) an antenna configured to receive a signal, the signal including the video block, (ii) a band limiter configured to limit the received signal to a frequency band including the video block, and (iii) a display configured to display an output representing the video block.

12. A non-transitory computer readable medium containing data content generated by the method according to any one of claims 1 or 5 to 10, or generated by the apparatus according to any one of claims 2 or 5 to 10 or 11, for playback using a processor.

13. A signal comprising video data generated by a method according to any one of claims 1 or 5 to 10, or by an apparatus according to any one of claims 2 or 5 to 11, for playback using a processor.

14. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to perform the method according to any one of claims 1 or 3 or 5 to 10.

15. A non-transitory computer-readable medium containing data content including instructions for executing the method according to any one of claims 1 or 3 and 5 to 10.