Method, apparatus, and medium for video processing

Neural representation-based video compression with parameter reuse improves coding efficiency by enhancing network depth and width, addressing complexity and inter-picture redundancy in existing standards.

WO2026039436A1PCT designated stage Publication Date: 2026-02-19BYTEDANCE INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/041658
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-16
Filing Date
2025-08-12
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Existing video compression technologies, such as MPEG, ITU-T H.264, and VVC, require improvements in coding efficiency, and neural network-based methods face challenges in complexity and inter-picture redundancy, especially in video coding standards like HEVC and VVC.

Method used

Implement neural representation-based video compression using parameter reuse to enhance network expressive power without increasing network parameters, leveraging convolutional neural networks and auto-encoders for improved rate-distortion performance.

Benefits of technology

Enhances neural representation-based video compression performance by increasing network depth and width, achieving comparable or better results than current standards like VVC without increasing computational resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025041658_19022026_PF_FP_ABST
    Figure US2025041658_19022026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is proposed. The method comprises: performing, for a conversion between a video unit of a video and a bitstream of the video, a compression on the video unit based on a neural representation, wherein one or more parameters are reused in the neural representation; and performing the conversion based on the compressed video unit.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD, APPARATUS, AND MEDIUM FOR VIDEO PROCESSINGFIEEDS[OOOlJEmbodiments of the present disclosure relates generally to video processing techniques, and more particularly, to parameter reuse for neural video compression.BACKGROUND

[0002] In nowadays, digital video capabilities are being applied in various aspects of peoples’ lives. Multiple types of video compression technologies, such as motion picture expert group (MPEG)-2, MPEG-4, international telecommunication union - telecommunication standardization sector (ITU-T) H.263, ITU-T H.264 / MPEG-4 Part 10 advanced video coding (AVC), ITU-T H.265 high efficiency video coding (HEVC) standard, versatile video coding (VVC) standard, have been proposed for video encoding / decoding. However, coding efficiency of video coding techniques is generally expected to be further improved.SUMMARY

[0003] Embodiments of the present disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is proposed. The method comprises: performing, for a conversion between a video unit of a video and a bitstream of the video, a compression on the video unit based on a neural representation, wherein one or more parameters are reused in the neural representation; and performing the conversion based on the compressed video unit. In this way, using parameter reuse technology can improve the network's expressive power without increasing the number of network parameters and further enhance neural representation based video compression performance.

[0005] In a second aspect, an apparatus for video processing is proposed. The apparatus comprises a processor and a non-transitory memory with instructions thereon. The instructions upon execution by the processor, cause the processor to perform a method in accordance with the first aspect of the present disclosure.

[0006] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non- transitory computer-readable storage medium stores instructions that cause a processor to perform a method in accordance with the first aspect of the present disclosure.

[0007] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream of a video which is generated by a method performed by an apparatus for video processing. The method comprises: performing a compression on a video unit of the video based on a neural representation, wherein one or more parameters are reused in the neural representation; and generating the bitstream based on the compressed video unit.

[0008] In a fifth aspect, a method for storing a bitstream of a video is proposed. The method comprises: performing a compression on a video unit of the video based on a neural representation, wherein one or more parameters are reused in the neural representation; generating the bitstream based on the compressed video unit; and storing the bitstream in anon-transitory computer-readable recording medium.1 F1254283PCT

[0009] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.BRIEF DESCRIPTION OF THE DRAWINGS[OOlOJThrough the following detailed description with reference to the accompanying drawings, the above and other objectives, features, and advantages of example embodiments of the present disclosure will become more apparent. In the example embodiments of the present disclosure, the same reference numerals usually refer to the same components.[OOllJFig. 1 illustrates a block diagram of an example video coding system in accordance with some embodiments of the present disclosure;

[0012] Fig. 2 illustrates a block diagram of an example video encoder in accordance with some embodiments of the present disclosure;

[0013] Fig. 3 illustrates a block diagram of an example video decoder in accordance with some embodiments of the present disclosure;

[0014] Fig. 4 illustrates an illustration of a typical transform coding scheme;

[0015] Fig. 5A to Fig. 5D illustrate an illustration of the example video neural representation;

[0016] Fig. 6A to Fig. 6D illustrate illustrations of examples of structure and encoding pipeline; and

[0017] Fig. 7 illustrates a flowchart of a method for video processing in accordance with some embodiments of the present disclosure;

[0018] Fig. 8 illustrates a block diagram of a computing device in which various embodiments of the present disclosure can be implemented.

[0019] Throughout the drawings, the same or similar reference numerals usually refer to the same or similar elements.DETAILED DESCRIPTION

[0020] Principle of the present disclosure will now be described with reference to some embodiments. It is to be understood that these embodiments are described only for the purpose of illustration and help those skilled in the art to understand and implement the present disclosure, without suggesting any limitation as to the scope of the disclosure. The disclosure described herein can be implemented in various manners other than the ones described below.

[0021] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.

[0022] References in the present disclosure to “one embodiment,” “an embodiment,” “an example embodiment,” and the like indicate that the embodiment described may include a particular feature, structure, or characteristic, but it is not necessary that every embodiment includes the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example2 F1254283PCTembodiment, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.

[0023] It shall be understood that although the terms “first” and “second” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and similarly, a second element could be termed a first element, without departing from the scope of example embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0024] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments. As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises”, “comprising”, “has”, “having”, “includes” and / or “including”, when used herein, specify the presence of stated features, elements, and / or components etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.Example Environment

[0025] Fig. 1 is a block diagram that illustrates an example video coding system 100 that may utilize the techniques of this disclosure. As shown, the video coding system 100 may include a source device 110 and a destination device 120. The source device 110 can be also referred to as a video encoding device, and the destination device 120 can be also referred to as a video decoding device. In operation, the source device 110 can be configured to generate encoded video data and the destination device 120 can be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0026] The video source 112 may include a source such as a video capture device. Examples of the video capture device include, but are not limited to, an interface to receive video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.

[0027] The video data may comprise one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded picture is a coded representation of a picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The encoded video data may be transmitted directly to destination device 120 via the I / O interface 116 through the network 130A. The encoded video data may also be stored onto a storage medium / server 130B for access by destination device 120.

[0028] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video3 F1254283PCTdata to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120 which is configured to interface with an external display device.

[0029] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard and other current and / or further standards.

[0030] Fig. 2 is a block diagram illustrating an example of a video encoder 200, which may be an example of the video encoder 114 in the system 100 illustrated in Fig. 1, in accordance with some embodiments of the present disclosure.

[0031] The video encoder 200 may be configured to implement any or all of the techniques of this disclosure. In the example of Fig. 2, the video encoder 200 includes a plurality of functional components. The techniques described in this disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.

[0032] In some embodiments, the video encoder 200 may include a partition unit 201, a prediction unit 202 which may include a mode select unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.

[0033] In other examples, the video encoder 200 may include more, fewer, or different functional components. In an example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is a picture where the current video block is located.

[0034] Furthermore, although some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be integrated, but are represented in the example of Fig. 2 separately for purposes of explanation.

[0035] The partition unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0036] The mode select unit 203 may select one of the coding modes, intra or inter, e.g., based on error results, and provide the resulting intra-coded or inter-coded block to a residual generation unit 207 to generate residual block data and to a reconstruction unit 212 to reconstruct the encoded block for use as a reference picture. In some examples, the mode select unit 203 may select a combined inter and intra prediction (CIIP) mode in which the prediction is based on an inter prediction signal and an intra prediction signal. The mode select unit 203 may also select a resolution for a motion vector (e.g., a subpixel or integer pixel precision) for the block in the case of inter-prediction.

[0037] To perform inter prediction on a current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 to the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the buffer 213 other than the picture associated with the current video block.4 F1254283PCT

[0038] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations for a current video block, for example, depending on whether the current video block is in an I-slice, a P-slice, or a B-slice. As used herein, an “I-slice” may refer to a portion of a picture composed of macroblocks, all of which are based upon macroblocks within the same picture. Further, as used herein, in some aspects, “P-slices” and “B-slices” may refer to portions of a picture composed of macroblocks that are not dependent on macroblocks in the same picture.

[0039] In some examples, the motion estimation unit 204 may perform uni-directional prediction for the current video block, and the motion estimation unit 204 may search reference pictures of list 0 or list 1 for a reference video block for the current video block. The motion estimation unit 204 may then generate a reference index that indicates the reference picture in list 0 or list 1 that contains the reference video block and a motion vector that indicates a spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, a prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate the predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.

[0040] Alternatively, in other examples, the motion estimation unit 204 may perform bi-directional prediction for the current video block. The motion estimation unit 204 may search the reference pictures in list 0 for a reference video block for the current video block and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 204 may then generate reference indexes that indicate the reference pictures in list 0 and list 1 containing the reference video blocks and motion vectors that indicate spatial displacements between the reference video blocks and the current video block. The motion estimation unit 204 may output the reference indexes and the motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate the predicted video block of the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0041] In some examples, the motion estimation unit 204 may output a full set of motion information for decoding processing of a decoder. Alternatively, in some embodiments, the motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.

[0042] In one example, the motion estimation unit 204 may indicate, in a syntax structure associated with the current video block, a value that indicates to the video decoder 300 that the current video block has the same motion information as the another video block.

[0043] In another example, the motion estimation unit 204 may identify, in a syntax structure associated with the current video block, another video block and a motion vector difference (MVD). The motion vector difference indicates a difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current5 F1254283PCTvideo block.

[0044] As discussed above, video encoder 200 may predictively signal the motion vector. Two examples of predictive signaling techniques that may be implemented by video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.

[0045] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0046] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by the minus sign) the predicted video block (s) of the current video block from the current video block. The residual data of the current video block may include residual video blocks that correspond to different sample components of the samples in the current video block.

[0047] In other examples, there may be no residual data for the current video block, for example in a skip mode, and the residual generation unit 207 may not perform the subtracting operation.

[0048] The transform unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to a residual video block associated with the current video block.

[0049] After the transform unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0050] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transforms to the transform coefficient video block, respectively, to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in the buffer 213.

[0051] After the reconstruction unit 212 reconstructs the video block, loop filtering operation may be performed to reduce video blocking artifacts in the video block.

[0052] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.

[0053] Fig. 3 is a block diagram illustrating an example of a video decoder 300, which may be an example of the video decoder 124 in the system 100 illustrated in Fig. 1, in accordance with some embodiments of the present disclosure.

[0054] The video decoder 300 may be configured to perform any or all of the techniques of this disclosure. In the example of Fig. 3, the video decoder 300 includes a plurality of functional components. The6 F1254283PCTtechniques described in this disclosure may be shared among the various components of the video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.

[0055] In the example of Fig. 3, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306 and a buffer 307. The video decoder 300 may, in some examples, perform a decoding pass generally reciprocal to the encoding pass described with respect to video encoder 200.

[0056] The entropy decoding unit 301 may retrieve an encoded bitstream. The encoded bitstream may include entropy coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy coded video data, and from the entropy decoded video data, the motion compensation unit 302 may determine motion information including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 may, for example, determine such information by performing the AMVP and merge mode. AMVP is used, including derivation of several most probable candidates based on data from adjacent PBs and the reference picture. Motion information typically includes the horizontal and vertical motion vector displacement values, one or two reference picture indices, and, in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, a “merge mode” may refer to deriving the motion information from spatially or temporally neighboring blocks.

[0057] The motion compensation unit 302 may produce motion compensated blocks, possibly performing interpolation based on interpolation filters. Identifiers for interpolation filters to be used with sub-pixel precision may be included in the syntax elements.

[0058] The motion compensation unit 302 may use the interpolation filters as used by the video encoder 200 during encoding of the video block to calculate interpolated values for sub-integer pixels of a reference block. The motion compensation unit 302 may determine the interpolation filters used by the video encoder 200 according to the received syntax information and use the interpolation filters to produce predictive blocks.

[0059] The motion compensation unit 302 may use at least part of the syntax information to determine sizes of blocks used to encode frame(s) and / or slice(s) of the encoded video sequence, partition information that describes how each macroblock of a picture of the encoded video sequence is partitioned, modes indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-encoded block, and other information to decode the encoded video sequence. As used herein, in some aspects, a “slice” may refer to a data structure that can be decoded independently from other slices of the same picture, in terms of entropy coding, signal prediction, and residual signal reconstruction. A slice can either be an entire picture or a region of a picture.

[0060] The intra prediction unit 303 may use intra prediction modes for example received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 304 inverse quantizes, i.e., de-quantizes, the quantized video block coefficients provided in the bitstream and decoded7 F1254283PCTby entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.

[0061] The reconstruction unit 306 may obtain the decoded blocks, e.g., by summing the residual blocks with the corresponding prediction blocks generated by the motion compensation unit 302 or intraprediction unit 303. If desired, a deblocking filter may also be applied to filter the decoded blocks in order to remove blockiness artifacts. The decoded video blocks are then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.

[0062] Some example embodiments of the present disclosure will be described in detailed hereinafter. It should be understood that section headings are used in the present document to facilitate ease of understanding and do not limit the embodiments disclosed in a section to only that section. Furthermore, while certain embodiments are described with reference to Versatile Video Coding or other specific video codecs, the disclosed techniques are applicable to other video coding technologies also. Furthermore, while some embodiments describe video coding steps in detail, it will be understood that corresponding steps decoding that undo the coding will be implemented by a decoder. Furthermore, the term video processing encompasses video coding or compression, video decoding or decompression and video transcoding in which video pixels are represented from one compressed format into another compressed format or at a different compressed bitrate.1. Brief Summary

[0063] This application relates to neural video compression, and more specifically, to video compression using neural representations as compression information tools. The described technology can involve improved algorithms and methods for enhancing the optimal rate-distortion (R-D) performance of neural representation based video compression. Typically, the described techniques rely on parameter reuse methods based on neural representation network structures. Without increasing network parameters, the network layers, modules, and other network components are stacked multiple times in the network to increase its depth or width, thereby enhancing the network's expressive power and improving the R-D performance of video compression based on neural representation. For example, a video neural representation is composed of multiple layers of convolutional neural networks, which increase the network depth by adding convolutional network layers with the same parameters after each convolutional network layer, and then train to fit the target video.2. Introduction

[0064] The past decade has witnessed the rapid development of deep learning in a variety of areas, especially in computer vision and image processing. Inspired from the great success of deep learning technology to computer vision areas, many researchers have shifted their attention from conventional image / video compression techniques to neural image / video compression technologies. Neural network was invented originally with the interdisciplinary research of neuroscience and mathematics. It has shown strong capabilities in the context of non-linear transform and classification. Neural network-based image / video compression technology has gained significant progress during the past half decade. It is reported that the latest neural network -based image compression algorithm achieves comparable R-D performance with Versatile Video Coding (VVC), the latest video coding standard developed by Joint8 F1254283PCTVideo Experts Team (JVET) with experts from MPEG and VCEG. With the performance of neural image compression continually being improved, neural network-based video compression has become an actively developing research area. However, neural network-based video coding still remains in its infancy due to the inherent difficulty of the problem.2.1. I magc / Vidco compression

[0065] Image / video compression usually refers to the computing technology that compresses image / video into binary code to facilitate storage and transmission. The binary codes may or may not support losslessly reconstructing the original image / video, termed lossless compression and lossy compression. Most of the efforts are devoted to lossy compression since lossless reconstruction is not necessary in most scenarios. Usually the performance of image / video compression algorithms is evaluated from two aspects, i.e. compression ratio and reconstruction quality. Compression ratio is directly related to the number of binary codes, the less the better; Reconstruction quality is measured by comparing the reconstructed image / video with the original image / video, the higher the better.

[0066] Image / video compression techniques can be divided into two branches, the classical video coding methods and the neural-network-based video compression methods. Classical video coding schemes adopt transform-based solutions, in which researchers have exploited statistical dependency in the latent variables (e.g., DCT or wavelet coefficients) by carefully hand-engineering entropy codes modeling the dependencies in the quantized regime. Neural network-based video compression is in two flavors, neural network-based coding tools and end-to-end neural network -based video compression. The former is embedded into existing classical video codecs as coding tools and only serves as part of the framework, while the latter is a separate framework developed based on neural networks without depending on classical video codecs.

[0067] In the last three decades, a series of classical video coding standards have been developed to accommodate the increasing visual content. The international standardization organizations ISO / IEC has two expert groups namely Joint Photographic Experts Group (JPEG) and Moving Picture Experts Group (MPEG), and ITU-T also has its own Video Coding Experts Group (VCEG) which is for standardization of image / video coding technology. The influential video coding standards published by these organizations include JPEG, JPEG 2000, H.262, H.264 / AVC and H.265 / HEVC. After H.265 / HEVC, the Joint Video Experts Team (JVET) formed by MPEG and VCEG has been working on a new video coding standard Versatile Video Coding (VVC). The first version of VVC was released in July 2020. An average of 50% bitrate reduction is reported by VVC under the same visual quality compared with HEVC.

[0068] Neural network-based image / video compression is not a new solution since there were a number of researchers working on neural network-based image coding. But the network architectures were relatively shallow, and the performance was not satisfactory. Benefit from the abundance of data and the support of powerful computing resources, neural network-based methods are better exploited in a variety of applications. At present, neural network-based image / video compression has shown promising improvements, confirmed its feasibility. Nevertheless, this technology is still far from mature and a lot of challenges need to be addressed.9 F1254283PCT2.2. Neural networks

[0069] Neural networks, also known as artificial neural networks (ANN), are the computational models used in machine learning technology which are usually composed of multiple processing layers and each layer is composed of multiple simple but non-linear basic computational units. One benefit of such deep networks is believed to be the capacity for processing data with multiple levels of abstraction and converting data into different kinds of representations. Note that these representations are not manually designed; instead, the deep network including the processing layers is learned from massive data using a general machine learning procedure. Deep learning eliminates the necessity of handcrafted representations, and thus is regarded useful especially for processing natively unstructured data, such as acoustic and visual signal, whilst processing such data has been a longstanding difficulty in the artificial intelligence field.2.3. Neural networks for image compression

[0070] Existing neural networks for image compression methods can be classified in two categories, i.e., pixel probability modeling and auto-encoder. The former one belongs to the predictive coding strategy, while the latter one is the transform -based solution. Sometimes, these two methods are combined together in literature.2.3.1. Pixel Probability Modeling

[0071] According to Shannon’s information theory, the optimal method for lossless coding can reach the minimal coding rate — log2p(x) where p(x) is the probability of symbol x. A number of lossless coding methods were developed in literature and among them arithmetic coding is believed to be among the optimal ones. Given a probability distribution p(x), arithmetic coding ensures that the coding rate to be as close as possible to its theoretical limit — log2p(x) without considering the rounding error. Therefore, the remaining problem is to how to determine the probability, which is however very challenging for natural image / video due to the curse of dimensionality.

[0072] Following the predictive coding strategy, one way to model p(x) is to predict pixel probabilities one by one in a raster scan order based on previous observations, where x is an image.where m and n are the height and width of the image, respectively. The previous observation is also known as the context of the current pixel. When the image is large, it can be difficult to estimate the conditional probability, thereby a simplified method is to limit the range of its context. p(x) = p(x1)p(x2|x1) ...p(Xi|Xi_fe, ...' Xi-i) ... p(xmxn|xmxn-k, ...,xmxn-1) (2) where k is a pre-defined constant controlling the range of the context.

[0073] It should be noted that the condition may also take the sample values of other color components into consideration. For example, when coding the RGB color component, R sample is dependent on previously coded pixels (including R / G / B samples), the current G sample may be coded according to previously coded pixels and the current R sample, while for coding the current B sample, the previously coded pixels and the current R and G samples may also be taken into consideration.

[0074] Neural networks were originally introduced for computer vision tasks and have been proven to be effective in regression and classification problems. Therefore, it has been proposed using neural networks10 F1254283PCTto estimate the probability of p(x;) given its context x1,x2, ..., xi-1. In some solutions, the pixel probability is proposed for binary images, i.e., xtG {— 1, +1}. The neural autoregressive distribution estimator (NADE) is designed for pixel probability modeling, where is a feed-forward network with a single hidden layer. A similar work is presented, where the feed-forward network also has connections skipping the hidden layer, and the parameters are also shared. Some solutions perform experiments on the binarized MNIST dataset. In an example solution, NADE is extended to a real-valued model RNADE, where the probability p xi\x1, .... x^^ is derived with a mixture of Gaussians. Their feed-forward network also has a single hidden layer, but the hidden layer is with rescaling to avoid saturation and uses rectified linear unit (ReLU) instead of sigmoid. In an example solution, NADE and RNADE are improved by using reorganizing the order of the pixels and with deeper neural networks.

[0075] Designing advanced neural networks plays an important role in improving pixel probability modeling. In an example solution, multi-dimensional long short-term memory (LSTM) is proposed, which is working together with mixtures of conditional Gaussian scale mixtures for probability modeling. LSTM is a special kind of recurrent neural networks (RNNs) and is proven to be good at modeling sequential data. The spatial variant of LSTM is used for images later in an example solution. Several different neural networks are studied, including RNNs and CNNs namely PixelRNN and PixelCNN, respectively. In PixelRNN, two variants of LSTM, called row LSTM and diagonal BiLSTM are proposed, where the latter is specifically designed for images. PixelRNN incorporates residual connections to help train deep neural networks with up to 12 layers. In PixelCNN, masked convolutions are used to suit for the shape of the context. Comparing with previous works, PixelRNN and PixelCNN are more dedicated to natural images: they consider pixels as discrete values (e.g., 0, 1, ... , 255) and predict a multinomial distribution over the discrete values; they deal with color images in RGB color space; they work well on large-scale image dataset ImageNet. In an example solution, Gated PixelCNN is proposed to improve the PixelCNN, and achieves comparable performance with PixelRNN but with much less complexity. In an example solution, PixelCNN++ is proposed with the following improvements upon PixelCNN: a discretized logistic mixture likelihood is used rather than a 256-way multinomial distribution; downsampling is used to capture structures at multiple resolutions; additional short-cut connections are introduced to speed up training; dropout is adopted for regularization; RGB is combined for one pixel. In an example solution, PixelSNAIL is proposed, in which casual convolutions are combined with selfattention.

[0076] Most of the above methods directly model the probability distribution in the pixel domain. Some researchers also attempt to model the probability distribution as a conditional one upon explicit or latent representations. That being said, it may estimatewhere h is the additional condition and p (x) = p (h)p (x| ft), meaning the modeling is split into an unconditional one and a conditional one. The additional condition can be image label information or high-level representations.2.3.2. Auto-encoder

[0077] Auto-encoder originates from the well-known work proposed. The method is trained for dimensionality reduction and consists of two parts: encoding and decoding. The encoding part converts11 F1254283PCTthe high-dimension input signal to low-dimension representations, typically with reduced spatial size but a greater number of channels. The decoding part attempts to recover the high-dimension input from the low-dimension representation. Auto-encoder enables automated learning of representations and eliminates the need of hand-crafted features, which is also believed to be one of the most important advantages of neural networks.

[0078] Fig. 4 illustrates an illustration of a typical transform coding scheme. The original image x is transformed by the analysis network gato achieve the latent representation y. The latent representation y is quantized and compressed into bits. The number of bits R is used to measure the coding rate. The quantized latent representation y is then inversely transformed by a synthesis network gsto obtain the reconstructed image x. The distortion is calculated in a perceptual space by transforming x and x with the function gv.

[0079] It is intuitive to apply auto-encoder network to lossy image compression. It only needs to encode the learned latent representation from the well-trained neural networks. However, it is not trivial to adapt auto-encoder to image compression since the original auto-encoder is not optimized for compression thereby not efficient by directly using a trained auto-encoder. In addition, there exist other major challenges: First, the low-dimension representation should be quantized before being encoded, but the quantization is not differentiable, which is required in backpropagation while training the neural networks. Second, the objective under compression scenario is different since both the distortion and the rate need to be take into consideration. Estimating the rate is challenging. Third, a practical image coding scheme needs to support variable rate, scalability, encoding / decoding speed, interoperability. In response to these challenges, a number of researchers have been actively contributing to this area.

[0080] The prototype auto-encoder for image compression is in Fig. 4, which can be regarded as a transform coding strategy. The original image x is transformed with the analysis network y = ga(xf where y is the latent representation which will be quantized and coded. The synthesis network will inversely transform the quantized latent representation y back to obtain the reconstructed image x = 5s(y)- The framework is trained with the rate-distortion loss function, i.e., £ = D + AR, where D is the distortion between x and x, R is the rate calculated or estimated from the quantized representation y, and A is the Lagrange multiplier. It should be noted that D can be calculated in either pixel domain or perceptual domain. All existing research works follow this prototype and the difference might only be the network structure or loss function.

[0081] In terms of network structure, RNNs and CNNs are the most widely used architectures. In the RNNs relevant category, it proposes a general framework for variable rate image compression using RNN. They use binary quantization to generate codes and do not consider rate during training. The framework indeed provides a scalable coding functionality, where RNN with convolutional and deconvolution layers is reported to perform decently. Some solutions then proposed an improved version by upgrading the encoder with a neural network similar to PixelRNN to compress the binary codes. The performance is reportedly better than JPEG on Kodak image dataset using MS-SSIM evaluation metric. Some solution further improve the RNN-based solution by introducing hidden-state priming. In addition, an SSIM- weighted loss function is also designed, and spatially adaptive bitrates mechanism is enabled. They12 F1254283PCTachieve better results than BPG on Kodak image dataset using MS-SSIM as evaluation metric. Some solutions support spatially adaptive bitrates by training stop-code tolerant RNNs.

[0082] An example solution proposes a general framework for rate-distortion optimized image compression. The use multiary quantization to generate integer codes and consider the rate during training, i.e. the loss is the joint rate-distortion cost, which can be MSE or others. They add random noise to stimulate the quantization during training and use the differential entropy of the noisy codes as a proxy for the rate. They use generalized divisive normalization (GDN) as the network structure, which consists of a linear mapping followed by a nonlinear parametric normalization. The effectiveness of GDN on image coding is verified an example solution. The example solution then proposes an improved version, where they use 3 convolutional layers each followed by a down-sampling layer and a GDN layer as the forward transform. Accordingly, they use 3 layers of inverse GDN each followed by an up-sampling layer and convolution layer to stimulate the inverse transform. In addition, an arithmetic coding method is devised to compress the integer codes. The performance is reportedly better than JPEG and JPEG 2000 on Kodak dataset in terms of MSE. Furthermore, some solutions improve the method by devising a scale hyper-prior into the auto-encoder. They transform the latent representation y with a subnet hato z = ha(y) and z will be quantized and transmitted as side information. Accordingly, the inverse transform is implemented with a subnet hsattempting to decode from the quantized side information z to the standard deviation of the quantized y, which will be further used during the arithmetic coding of y. On the Kodak image set, their method is slightly worse than BGP in terms of PSNR. Some solutions further exploit the structures in the residue space by introducing an autoregressive model to estimate both the standard deviation and the mean. In the latest work, it uses Gaussian mixture model to further remove redundancy in the residue. The reported performance is on par with VVC on the Kodak image set using PSNR as evaluation metric.2.4. Neural networks for video compression

[0083] Similar to conventional video coding technologies, neural image compression serves as the foundation of intra compression in neural network-based video compression, thus development of neural network -based video compression technology comes later than neural network-based image compression but needs far more efforts to solve the challenges due to its complexity. Starting from 2017, a few researchers have been working on neural network-based video compression schemes. Compared with image compression, video compression needs efficient methods to remove inter-picture redundancy. Inter-picture prediction is then a crucial step in these works. Motion estimation and compensation is widely adopted but is not implemented by trained neural networks until recently.

[0084] Studies on neural network-based video compression can be divided into two categories according to the targeted scenarios: random access and the low-latency. In random access case, it requires the decoding can be started from any point of the sequence, typically divides the entire sequence into multiple individual segments and each segment can be decoded independently. In low-latency case, it aims at reducing decoding time thereby usually merely temporally previous frames can be used as reference frames to decode subsequent frames.13 F1254283PCT2.4.1. Low-latency

[0085] Some solutions are the first to propose a video compression scheme with trained neural networks. They first split the video sequence frames into blocks and each block will choose one from two available modes, either intra coding or inter coding. If intra coding is selected, there is an associated auto-encoder to compress the block. If inter coding is selected, motion estimation and compensation are performed with tradition methods and a trained neural network will be used for residue compression. The outputs of auto-encoders are directly quantized and coded by the Huffman method.

[0086] Some solutions propose another neural network-based video coding scheme with PixelMotionCNN. The frames are compressed in the temporal order, and each frame is split into blocks which are compressed in the raster scan order. Each frame will firstly be extrapolated with the preceding two reconstructed frames. When a block is to be compressed, the extrapolated frame along with the context of the current block are fed into the PixelMotionCNN to derive a latent representation. Then the residues are compressed by the variable rate image scheme. This scheme performs on par with H.264.

[0087] Some solutions propose the real-sense end-to-end neural network-based video compression framework, in which all the modules are implemented with neural networks. The scheme accepts current frame and the prior reconstructed frame as inputs and optical flow will be derived with a pre-trained neural network as the motion information. The motion information will be warped with the reference frame followed by a neural network generating the motion compensated frame. The residues and the motion information are compressed with two separate neural auto -encoders. The whole framework is trained with a single rate-distortion loss function. It achieves better performance than H.264.

[0088] Some solutions propose an advanced neural network-based video compression scheme. It inherits and extends traditional video coding schemes with neural networks with the following major features: 1) using only one auto-encoder to compress motion information and residues; 2) motion compensation with multiple frames and multiple optical flows; 3) an on-line state is learned and propagated through the following frames over time. This scheme achieves better performance in MS-SSIM than HEVC reference software.

[0089] Some solutions propose an extended end-to-end neural network-based video compression framework. In this solution, multiple frames are used as references. It is thereby able to provide more accurate prediction of current frame by using multiple reference frames and associated motion information. In addition, motion field prediction is deployed to remove motion redundancy along temporal channel. Postprocessing networks are also introduced in this work to remove reconstruction artifacts from previous processes. The performance is better and H.265 by a noticeable margin in terms of both PSNR and MS-SSIM.

[0090] Some solutions propose scale-space flow to replace commonly used optical flow by adding a scale parameter based on framework of a conventional solution. It is reportedly achieving better performance than H.264.

[0091] Some solutions propose a multi-resolution representation for optical flows based on a conventional solution. Concretely, the motion estimation network produces multiple optical flows with different resolutions and let the network to learn which one to choose under the loss function. The performance is14 F1254283PCTslightly improved and beter than H.265.2.4.2.Random access

[0092] Some solutions propose a neural network-based video compression scheme with frame interpolation. The key frames are first compressed with a neural image compressor and the remaining frames are compressed in a hierarchical order. They perform motion compensation in the perceptual domain, i.e. deriving the feature maps at multiple spatial scales of the original frame and using motion to warp the feature maps, which will be used for the image compressor. The method is reportedly on par with H.264.

[0093] Some solutions propose a method for interpolation-based video compression, wherein the interpolation model combines motion information compression and image synthesis, and the same autoencoder is used for image and residual.

[0094] Some solutions propose a neural network-based video compression method based on variational auto-encoders with a deterministic encoder. Concretely, the model consists of an auto-encoder and an auto-regressive prior. Different from previous methods, this method accepts a group of pictures (GOP) as inputs and incorporates a 3D autoregressive prior by taking into account of the temporal correlation while coding the laten representations. It provides comparative performance as H.265.2.4.3.Neural representation based video compression

[0095] Neural representation based video compression uses a neural network to fit the video and compress the fitted network for video compression. Some solutions first proposed an neural representation based video compression pipeline consisting of video overfitting, model pruning, model quantization, and weight encoding. Some solutions propose a hybrid neural representation for videos. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (pp. 10270-10279).) further improved the network structure based on the conventional solution, introduced autoregressive embeddings to enhance the representation performance of video neural representation, thereby improving the final video compression performance. Some solutions improved the compression pipeline of NeRV. Nerv: Neural representations for videos. Advances in Neural Information Processing Systems, 34, 21557- 21568.) by introducing network parameter information entropy loss during training to achieve joint rate distortion optimization. In Proceedings of the IEEE / CVF International Conference on Computer Vision (pp. 12481-12491).) improved the performance of video neural representation fitting by introducing optical flow supervision, frequency domain supervision, and contrastive loss. Some solutions has divided autoregressive embeddings into different scales and improved the expressive power of the INR network under finite parameter conditions through parameter free upsampling. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (pp. 2556-2566)) utilized a conditional decoder with a time aware affine transformation module and sine NeRV sample blocks to enhance the compression performance of nerual representation based video compression.2.5. Preliminaries

[0096] Almost all the natural image / video is in digital format. A grayscale digital image can be represented by x G EDmxn, where ED is the set of values of a pixel, m is the image height and n is the image width. For example, ED = {0, 1, 2, ...,255} is a common setting and in this case |ED| = 256 = 28, thus the15 F1254283PCTpixel can be represented by an 8 -bit integer. An uncompressed grayscale digital image has 8 bits-per- pixel (bpp), while compressed bits are definitely less.[0097JA color image is typically represented in multiple channels to record the color information. For example, in the RGB color space an image can be denoted by x E HDmxnx3with three separate channels storing Red, Green and Blue information. Similar to the 8-bit grayscale image, an uncompressed 8-bit RGB image has 24 bpp. Digital images / videos can be represented in different color spaces. The neural network-based video compression schemes are mostly developed in RGB color space while the traditional codecs typically use YUV color space to represent the video sequences. In YUV color space, an image is decomposed into three channels, namely Y, Cb and Cr, where Y is the luminance component and Cb / Cr are the chroma components. The benefits come from that Cb and Cr are typically down sampled to achieve pre-compression since human vision system is less sensitive to chroma components.[0098JA color video sequence is composed of multiple color images, called frames, to record scenes at different timestamps. For example, in the RGB color space, a color video can be denoted by X = {x0, xlt..., xt, ..., xT-1} where T is the number of frames in this video sequence, x E EDmxn. If m = 1080, n = 1920 , |ED| = 28, and the video has 50 frames-per-second (fps), then the data rate of this uncompressed video is 1920 x 1080 x 8 x 3 x 50 = 2,488,320,000 bits-per-second (bps), about 2.32 Gbps, which needs a lot storage thereby definitely needs to be compressed before transmission over the internet.

[0099] Usually the lossless methods can achieve compression ratio of about 1.5 to 3 for natural images, which is clearly below requirement. Therefore, lossy compression is developed to achieve further compression ratio, but at the cost of incurred distortion. The distortion can be measured by calculating the average squared difference between the original image and the reconstructed image, i.e., mean- squared-error (MSE). For a grayscale image, MSE can be calculated with the following equation:[OlOOJAccordingly, the quality of the reconstructed image compared with the original image can be measured by peak signal-to-noise ratio (PSNR):PSNR = 10 x logw^ M^SE (v5)7where max(ED) is the maximal value in ED, e.g., 255 for 8-bit grayscale images. There are other quality evaluation metrics such as structural similarity (SSIM) and multi-scale SSIM (MS-SSIM).[OlOlJTo compare different lossless compression schemes, it is sufficient to compare either the compression ratio given the resulting rate or vice versa. However, to compare different lossy compression methods, it has to take into account both the rate and reconstructed quality. For example, to calculate the relative rates at several different quality levels, and then to average the rates, is a commonly adopted method; the average relative rate is known as Bjontegaard’s delta-rate (BD-rate). There are other important aspects to evaluate image / video coding schemes, including encoding / decoding complexity, scalability, robustness, and so on.

[0102] Figs. 5A to Fig. 5D illustrate an illustration of the example video neural representation. Fig. 5A shows a complete structure of a video neural representation network. The input coordinate (i,j, t) is sent16 F1254283PCTto get the autoregressive embeddings and further processed by stem layer, Hinerv blocks, and a head layer to get the output video frame. Fig. 5B shows the detail structure of the nthHinerv block. Xn_ j is upsampled by bilinear interpolation and adds the processd local embeddings from a linear layer. The mixed feature j is further processed by DnConvNeXt blocks to generate output feature Xn. Fig. 5C shows the details of example parameter reusing by stacking same ConvNeXt blocks by multiple times. Fig. 5D shows the details of example parameter reusing by expanding the linear layer by weight concation.3. Problems

[0103] The following problems remain in existing neural representation based video compression solutions:1. For neural representation based video compression, existing method uses autoregressive embeddings and hierarchical ConvNeXt blocks for building video neural representation. However, simple stacking of different ConvNeXt blocks cannot fully unleash the potential of network parameters for storing video information, and therefore cannot achieve the optimal video compression R-D performance. As shown in Fig. 5C, we reuse the same ConvNeXt block for multiple times to process intermediate network features after each ConvNeXt block, thereby using parameter reuse technology to improve the network's expressive power without increasing the number of network parameters and further enhance neural representation based video compression performance.4. Detailed solutions

[0104] The detailed solutions below should be considered as examples to explain general concepts. These solutions should not be interpreted in a narrow way. Furthermore, these solutions can be combined in any manner.

[0105] The techniques described herein provide a neural representation-based video compression method with parameter reuse. By reusing same network parameters in the video neural representation, the R-D performance of neural representation-based video compression can be improved. In summary, the present disclosure includes the following example embodiments:1. A neural representation is used for video compression. The neural representation utilizes parameter reuse to improve the R-D performance for video compression. a. In one example, the video neural representation might consist of convolutional layers and stack the same convolutional layer for multiple times. b. In one example, the video neural representation might consist of linear layers and stack the same convolutional linear for multiple times. c. In one example, the video neural representation might consist of complex blocks consisting of convolutional layers and linear layers, and stack the same block for multiple times.. d. In one example, the video neural representation might consist of convolutional layers and expand the convolutional layer by concatenating the weight of the convolutional layers for multiple times. e. In one example, the video neural representation might consist of linear layers and expand the linear layer by concatenating the weight of the linear layers for multiple times.17 F1254283PCT5. Example Embodiments5.1. Embodiment#!

[0106] The common neural representation based video compression methods usually use simple network blocks stacked into neural networks for fiting videos, and then achieve video compression through network compression. However, simple stacking of network blocks cannot fully unleash the video neural representation performance of network parameters, thus failing to achieve optimal video compression performance. It proposes to achieve video network representation performance enhancement by reusing network parameters and stacking the same network layer, network block, or other network components multiple times, thereby improving the R-D performance of video compression. An example is provided in the following subsections.5.1.1. Network Structure

[0107] Network begins by mapping the input patch coordinates (z, j, f) to a basic input feature o:where Ybase represents the basic embeddings, and Fstemis a stem convolutional layer that adjusts these embeddings to the required number of channels. This initial feature o then enters the Hinerv blocks, where it is progressively processed and upsampled into the final video patch.

[0108] Each HiNeRV block performs a series of transformations, starting with bilinear interpolation to upsample the input feature A^-i by a scale factor Sn. It then calculates a hierarchical embedding y„(z, J, t) based on the coordinates (z, j, f) and maps it to the necessary number of channels through a linear layer. The upsampled feature X„- is then added with this processed embedding. This mixed feature is fed into multiple ConvNeXt blocks, which are designed to refine the feature representation. A ConvNeXt block includes a depthwise separable convolution layer that extracts spatial information, followed by layer normalization. The features are then amplified by a linear layer, activated by a GeLU function, and reduced in channels by another linear layer to produce the final output. Finally, the output of the last HiNeRV block is processed by a head layer that maps it to (R, G, B~) video frames using a linear layer.5.1.2.Encoding and Decoding

[0109] The HiNeRV network is trained using mean-square-error (MSE) as the loss function to represent the target video. The network is trained and made represent the target video effectively. After achieving this initial trained network, we enhance the network’s performance through quantization-aware-training (QAT). QAT introduces quantization noise as a substitute for direct quantization to prevent the zerogradient problem. This approach allows us to quantize the network parameters to 6 bits without losing significant information. During QAT, the network is finetuned to maintain high performance despite the reduced precision. This step is crucial for achieving efficient compression while preserving video quality. An arithmetic coding is then applied to the quantized network parameters, transforming them into a compact bitstream suit- able for transmission and storage.[OllOJThe decoding process begins with arithmetic decoding, which retrieves the quantized network parameters. These parameters are then reloaded into the HiNeRV network. By performing forward propagation, the network reconstructs the video frames corresponding to each input coordinate.18 F1254283PCT5.1.3.Parameter reuse[OlllJDeepeningthe Network. As shown in Fig. 5C, it implements a scheme where ConvNext blocks are repeatedly stacked. Specifically, the layer structure of a ConvNeXt block is repeated n times. Each stacked ConvNeXt block uses the same set of parameters, which deepens the network without increasing the parameter count. During the forward pass, for each input X, the output Y is obtained after passing through n identical ConvNeXt blocks. The process is represented as follows:Y = ConvNeXtn(X).

[0112] This method increases the network’s depth while keeping the total number of parameters unchanged since each layer reuses the same parameters. It has also conducted more exploration and experiments on the location, number of layers, and granularity of parameter reuse deepening. The specific results are shown in the experimental section.

[0113] Wideningthe Network. As shown in Fig. 5D, the output channels of the first linear layer and the input channels of the second linear layer in ConvNeXt are expanded, ensuring that the input and output channels of each ConvNeXt block are aligned. To widen the linear layers, the weights of two linear layers in the ConvNext blocks are concatenated as shown in Fig. 5D. This means that the weight matrices of these layers are repeated and concatenated to form a wider matrix without adding new parameters.

[0114] For an input X, a new weight matrix is created by concatenating the original weight matrix. These concatenated weights process the input, thereby widening the network.Wlnew=concat(WltWltaxis = 0)W2new=concat(W2, W2, axis = 1)

[0115] Here, F i and W2 represent the original weight matrixes, andand W2neware the new concatenated weight matrixes. The output Y is obtained by applying the two wider linear operations with Wi new and W2new.Y = Conv X)WnewW^new

[0116] This concatenation mechanism increases the network’s width while maintaining the same parameter count.5.2. Embodiment #2Unleashing Parameter Potential of Neural Representation for Efficient Video Compression

[0117] For decades, video compression technology has been a prominent research area. Traditional hybrid video compression frameworks and end-to-end frameworks continue to explore various intra- and interframe reference and prediction strategies based on discrete transforms and deep learning. However, the emerging implicit neural representation (INR) technique models entire videos as basic units, automatically capturing intra-frame and inter-frame correlations. INR uses a compact neural network to store video information in network parameters, effectively eliminating spatial and temporal redundancy in the original video. But do existing INR video compression methods fully utilize the potential of these network parameters? The exploration and verification reveal that current INR video compression methods do not fully exploit their potential to preserve information. The potential of enhancing network parameter storage is investigated through parameter reuse. By deepening the network, a feasible INR parameter19 F1254283PCTreuse scheme is designed to improve compression performance. Extensive experimental results show that the method significantly enhances the rate-distortion performance of INR video compression.Introduction

[0118] With the rise of various User Generated Content (UGC) platforms, the rapid increase of video content on the Internet has led to an exponential growth in video data. Therefore, efficient video compression is crucial for storage and transmission. In the past few decades, video compression technology has made significant progress, resulting in complex codec standards like H.264 / AVC , H.265 / HEVC , and H.266 / VVC . These traditional techniques rely on methods such as discrete transformation, quantization, entropy coding, and filtering to compress video data. However, due to the intricate design and complexity of each module, traditional video codecs fail to achieve end-to-end ratedistortion optimization, limiting their potential for optimal performance. As video consumption continues to grow, the demand for more advanced and efficient video compression methods is becoming increasingly important.

[0119] Breakthroughs in deep learning have revolutionized video compression technology, opening up new possibilities for efficient video compression. End-to-end neural video codecs have achieved notable success by learning compression strategies directly from data in a fully differentiable manner . These neural codecs optimize the encoder, decoder, and entropy model together, resulting in superior ratedistortion performance compared to traditional video compression standards. Despite these advances, training such large-scale neural networks demands substantial amounts of data and computing power, which presents significant challenges for practical deployment, especially on devices with limited resources. Nevertheless, the potential of neural video codecs continues to drive innovation in this area.

[0120] The emergence of implicit neural representation (INR) offers a promising new direction for neural video codecs. INR leverages neural networks to represent data compactly and efficiently. INR-based video compression trains a simple neural network to fit the target video and then compresses the network parameters to achieve video compression. This approach significantly reduces model complexity and computational demands at the decoder compared to traditional end-to-end neural codecs. Recent advancements in INR-based video compression have brought rapid improvements in rate-distortion performance. However, these methods still fall short compared to the performance of traditional video codecs and end-to-end neural codecs. “Has the information storage potential of INR network parameters been fully utilized?” Addressing this question is crucial for assessing the effectiveness and practicality of INR-based video compression.

[0121] To maximize compression efficiency, the goal is to retain as much of the original video information as possible while using the least amount of data. For INR-based video compression, it needs to determine whether existing INR network parameters can fully utilize their representation capabilities to maximize the quality of retained videos at a given bit rate. The current INR structure is simple, usually consisting of basic modules (CNN or MLP) stacked together. According to the development history of large neural networks, introducing complex mechanisms such as attention can enable neural networks with the same parameter level to achieve better performance. Therefore, current INR-based video compression methods may not fully unleash the potential of available parameter space, leaving room for20 F1254283PCTfurther optimization.

[0122] Fully unleashing the information storage capability of INR network parameters remains an open problem. In this work, it validates the potential of current INR video compression methods to preserve more video information through a parameter reuse mechanism. Extensive exploration is conducted on the parameter reuse mechanism of INR. By strategically reusing and recombining INR parameters, it achieves more expressive and informative video representations, significantly improving compression efficiency. Specifically, it explores a series of parameter reuse schemes, including weight cascading and network deepening, to determine the optimal configuration that maximizes video quality at a given bit rate. The extensive experiments on various video datasets showed that the proposed method has superior performance compared to exisiting INR-basd video compression methods.

[0123] The main contributions of some example embodiments are:• It validates the network parameters of current INR-based video compression methods still have the potential to save more video information through parameter reuse.• It proposes a suite of effective parameter reuse techniques to unleash the potential of neural representations for efficient video compression.• It reports significant rate-distortion performance improvements on multiple video datasets, demonstrating the ability of the method to enhance the performance of INR video compression methods.Traditional Video Coding Standards

[0124] The evolution of video coding standards highlights significant advancements in video compression. ISO / IEC’s MPEG-2 introduces discrete cosine transform (DCT) and motion compensation. H.264 / AVC , developed by ITU-T and ISO / IEC, implements macroblock-adaptive frame-field coding and variable block sizes. H.265 / HEVC enhances compression performance with improved motion vector prediction and quadtree partitioning. The latest standard, H.266 / VVC , utilizes adaptive loop filtering and multiple reference pictures, improving compression efficiency by 50% over H.265.End-to-End Neural Video Coding

[0125] With continuous breakthroughs in deep learning for computer vision tasks, neural video codecs advance rapidly. DVC marks a significant milestone, achieving end-to-end neural video codec ratedistortion optimization for the first time while still using the traditional residual coding framework. Numerous neural video codecs follow this residual coding approach. For example, Djelouah et al. introduce an autoencoder framework with a predictive coding scheme, using a motion estimation network to predict the next frame and a residual compression network to encode the prediction error. M-LVC leverages multiple reference frames to predict associated motion vectors and perform motion compensation for frame reconstruction in an end-to-end learned video compression scheme. C2F enhances motion compensation through a two-stage coarse-to-fine deep video compression framework and employs hyperprior-guided mode prediction for optimized block resolution and residual coding.

[0126] Unlike residual coding, conditional coding uses the reference frame as conditional information for transform and entropy coding, significantly advancing end-to-end video encoding. Ladune et al. show that encoding a video frame xtwith its motion-compensated reference frame xcresults in a lower entropy21 F1254283PCTrate than encoding the residual signal xt— xcwithout conditions. Numerous studies explore using spatial-temporal context to implement conditional entropy models, thereby improving video compression performance. The DCVC series uses motion compensation to continuously obtain temporal context and improve encoding, decoding, and entropy coding. DCVC-HEM implements an entropy model that utilizes spatial -temporal context. DCVC-DC demonstrates rate-distortion performance exceeding the current best traditional video coding standard, VVC . The latest work, DCVC-FM , optimizes the range of variable bitrates the model can achieve.INR-based Video Coding[0127JINR uses neural networks to represent data with two main paradigms for videos: data coordinates and autoregressive embeddings. The data coordinates methods use neural networks to map pixel coordinates ( ,y, t) to corresponding pixel values ( / ?, G, B). However, the data coordinates do not contain any content information of the target data, making the expressive power of INR completely dependent on simple neural networks. The autoregressive embeddings methods introduce embeddings updated during training to capture video content information, reducing network training difficulty and improving representation performance.

[0128] The emergence of INR provides a new approach for deep learning based video coding, which represents videos through a simple neural network and then compresses the neural network. NeRV first proposed an INR video compression pipeline consisting of video overfitting, model pruning, model quantization, and weight encoding. Subsequent methods such as E-NeRV , HNeRV , FFNeRV , etc. further improved the network structure based on NeRV , introduced autoregressive embeddings, and added optical flow references to enhance the representation performance of video INR, thereby improving the final video compression performance. Gomes et al. improved the compression pipeline of NeRV by introducing network parameter information entropy loss during training to achieve joint rate-distortion optimization. Tang et al. improved the performance of INR video representation by introducing optical flow supervision, frequency domain supervision, and contrastive loss. HiNeRV has divided autoregressive embeddings into different scales and improved the expressive power of the INR network under finite parameter conditions through parameter free upsampling. Zhang et al. utilized a conditional decoder with a time aware affine transformation module and sine NeRV sample blocks to enhance the compression performance of INR videos. Fig. 6A to Fig. 6D illustrate a comparison between proposed method and prior arts.Preliminaries

[0129] It introduces the structure and encoding pipeline of INR video compression based on the recent work, HiNeRV.Network Structure

[0130] As shown in Fig. 6A. HiNeRV begins by mapping the input patch coordinates i.j. t') to a basic input feature Xo0 linear(.Ybase(i> j > )> where Ybase represents the basic grid which is an autoregressive embedding in HiNeRV, and Flinearis a22 F1254283PCTlinear layer that adjusts the grid to the required number of channels. This initial feature Xothen enters n HiNeRV blocks, where it is progressively processed and upsampled into the output video patch.

[0131] As shown in Fig. 6A to Fig. 6D, each HiNeRV block performs a series of transformations, starting with bilinear interpolation to upsample the input feature Xn-rby a scale factor Sn. It then calculates a hierarchical grid yn(L,j, tX) based on the coordinate (l,j, t) and maps it to the necessary number of channels through a linear layer. The upsampled feature Xn-1is then added with this processed grid. This mixed feature is fed into DnConvNeXt blocks. A ConvNeXt block includes a depthwise separable convolution layer that extracts spatial information, followed by layer normalization. The features are then amplified by a linear layer, activated by a GeLU function, and reduced in channels by another linear layer to produce the final output.

[0132] Finally, the output of the last HiNeRV block is processed by a head layer that maps it to ( / ?, G, B video frames using a convlutional layer.Encoding and Decoding

[0133] In some example embodiments, the HiNeRV network is trained using mean-square-error (MSE) as the loss function to represent the target video. The network is trained and made represent the target video effectively. After achieving initial trained network, the network’s performance is enhanced through quantization-aware-training (QAT). QAT introduces quantization noise as a substitute for direct quantization to prevent the zero-gradient problem. This approach allows us to quantize the network parameters to 6 bits without losing significant information. During QAT, the network is finetuned to maintain high performance despite the reduced precision. This step is crucial for achieving efficient compression while preserving video quality. It then quantizes the network parameters with 6 bit and applies arithmetic coding to the quantized network parameters, transforming them into a compact bitstream suitable for transmission and storage.

[0134] The decoding process begins with arithmetic decoding, which retrieves the quantized network parameters. These parameters are then reloaded into the HiNeRV network. By performing forward propagation, the network reconstructs the video frames corresponding to each input coordinate.INR Video Compression Potential

[0135] The primary objective of video compression is to minimize the data size while maintaining high video quality or to enhance the quality of the reconstructed video within a fixed data size. For INR video compression, this challenge becomes an optimization problem that balances the number of parameters with video quality. Despite the already minimal parameter count in current INR models and the use of network compression techniques aimed at further reducing data size, a new angle: “Has the information storage potential of INR network parameters been fully utilized?” is suggested.

[0136] To explore whether INR network parameters fully exploit their information storage potential, it examines the development trajectory of INR video compression. Initially, the NeRV model uses simple CNNs to map coordinates to video frames, then fits the video and compresses the network for video compression. Subsequent efforts make minimal changes to video fitting and network compression, instead focusing on improving network architecture. Innovations like autoregressive embedding , complex ConvNext blocks , and auxiliary encoders significantly enhance the reconstruction quality of23 F1254283PCTINR models with the same number of parameters. These advancements demonstrate that architectural improvements effectively increase the information storage capacity of INR parameters.

[0137] Given this background, it is reasonable to assume that the information storage potential of INR network parameters has not been fully realized. To explore this further, it proposes a method that focuses on enhancing the utilization of existing network parameters through parameter reuse, rather than altering or adding new INR structures. Traditionally, a network’s expressive power increases with deeper and wider architectures, which often results in a higher parameter count. However, from a compression perspective, it hypothesizes that it is possible to increase the network’s depth and width without increasing the total number of parameters. By reusing parameters, it aims to enhance the network’s expressive capability while maintaining parameter number. This approach challenges the conventional trade-off between complexity and parameter count, suggesting that parameter reuse can achieve the desired expressive power without additional parameters. The method seeks to maximize the potential of current architectures, pushing boundaries in video compression and storage efficiency.

[0138] It implements parameter reuse in the ConvNeXt blocks of the HiNeRV network. This approach, depicted in Fig. 6C, augments the network’s depth while keeping the parameter count constant. The ConvNeXt block is only reused for once. It evaluates this technique using compression tests on the HEVC Class B dataset. The findings reveal that the upgraded HiNeRV network with parameter reuse exceeds the performance of the original network. This demonstrates that the parameter reuse strategy effectively unleashes the parameter potential of neural representation for efficient video compression.Parameter Reuse

[0139] The proposed parameter reuse mechanism aims to enhance the expressive capability of the HiNeRV network by reusing parameters within its ConvNeXt blocks. This approach allows for deepening and widening the network without increasing the overall parameter count. Below is a detailed description of the parameter reuse mechanism.Deepening the Network

[0140] Stacking ConvNeXt Blocks: It implements a scheme where ConvNeXt blocks are repeatedly stacked. Specifically, the layer structure of a ConvNeXt block is repeated for m times as shown in Fig. 6C. Each stacked ConvNeXt block uses the same set of parameters, which deepens the network without increasing the parameter count.

[0141] Implementation Steps: During the forward pass, for each input X, the output Y is obtained after passing through m identical ConvNeXt blocks. The process is represented as follows:Y = ConvNeXtm(X),

[0142] This method increases the network’s depth while keeping the total number of parameters unchanged since each layer reuses the same parameters. It has also conducted more exploration and experiments on the location, number of layers, and granularity of parameter reuse deepening. The specific results are shown in the experimental section.Widening the Network

[0143] Concatenation of Linear Layer Weights: Because when widening a neural network, it needs to consider aligning the number of input and output channels. Therefore, expanding the output channels of24 F1254283PCTthe first linear layer and the input channels of the second linear layer in ConvNeXt block is considered, ensuring that the input and output channels of each ConvNeXt block are aligned. To widen the linear layers, the weights of two linear layers in the ConvNeXt blocks are concatenate. This means that the weight matrices of these layers are repeated and concatenated to form a wider matrix without adding new parameters.

[0144] Implementation Steps: For an input , a new weight matrix is created by concatenating the original weight matrix. These concatenated weights process the input, thereby widening the network.H^inew=concat VF PF] , axis = 0), FF2new=concat(W2, W2,axis = 1),

[0145] Here, Wrand W2represent the original weight matrixes, and Wlnewand W2neware the new concatenated weight matrixes. The output Y is obtained by applying the two wider linear operations with H^lnewand Fl^newY = Conv WnewWnew,

[0146] This concatenation mechanism increases the network’s width while maintaining the same parameter count.[0147JA simple test of the network widening method on HEVC Class B is done. The test results show that widening the network does not improve the video compression performance. The scheme of widening the network through parameter reuse cannot stimulate the potential of the network to preserve information, so the network is not widen again in subsequent experiments.

[0148] Although widening the network cannot improve the video compression performance, deepening the network can help. Therefore, parameter reuse reduces redundancy, enabling a smaller set of parameters to perform more tasks, thus improving parameter utilization efficiency. The reuse mechanism allows the network to deepen its structure, enhancing its ability to model complex patterns in video data. This method improves reconstruction quality while keeping the parameter count stable.

[0149] Fig . 7 illustrates a flowchart of a method 700 for video processing in accordance with embodiments of the present disclosure. The method 700 is implemented during a conversion between a video unit of a video and a bitstream of the video.

[0150] At block 710, for a conversion between a video unit of a video and a bitstream of the video, a compression on the video unit is perfomred based on a neural representation. In this case, one or more parameters are reused in the neural representation. In this way, by reusing same network parameters in the video neural representation, the R-D performance of neural representation-based video compression can be improved.

[0151] At block 720, performing the conversion based on the compressed video unit. In some embodiments, the conversion includes encoding the video unit into the bitstream. In some other embodiments, the conversion includes decoding the video unit from the bitstream.

[0152] In some embodiments, the neural representation comprises one or more convolutional layers and stacks a same convolutional layer for multiple times. In some other embodiments, the neural representation comprises one or more linear layers and stacks a same convolutional linear for multiple times.25 F1254283PCT

[0153] In some embodiments, the neural representation comprises one or more complex blocks including one or more convolutional layers and one or more linear layers, and the neural representation stacks a same block for multiple times. In some other embodiments, the neural representation comprises one or more convolutional layers and expands a convolutional layers by concatenating a weight of the one or more convolutional layers for multiple times. In some further embodiments, the neural representation comprises one or more linear layers and expands a linear layer by concatenating a weight of the one or more linear layers for multiple times.

[0154] In some embodiments, a structure of a network for the compression comprises mapping an input patch coordinates to an initial input feature: 0 stem(.Ybase(i>j> y) where (i, j, t) represents the input patch coordinates, Xo represents the initial input feature, Ybase represents basic embeddings, and Fstemrepresents a stem convolutional layer that adjusts the basic embeddings to a number of channels. In some embodiments, the structure of the network for the compression further comprises one or more video compression with hierarchical encoding-based neural representation (HiNeRV) block where the initial feature is progressively processed and upsampled into a final video patch.

[0155] In some embodiments, each HiNeRV block performs a series of transformations includes: a bilinear interpolation to upsample an input feature X„-i by a scale factor S„; determining a hierarchical embedding y„(z, J, t) based on the coordinates (i, j, t) and mapping the hierarchical embedding y„(z, j, f) to the number of channels through a linear layer; adding the upsampled input feature X„- with a processed embedding; obtaining features by feeding the upsampled input feature X„-i added with a processed embedding into a plurality of ConvNeXt blocks, wherein a ConvNeXt block includes a depthwise separable convolution layer that extracts spatial information, followed by layer normalization; amplifying the features by a linear layer, activated by a GeLU function; obtaining a final output by reducing the features in channels by another linear layer; and processing the final output of a last HiNeRV block by a head layer that maps it to (R, G, B) video frames using a linear layer.

[0156] In some embodiments, a HiNeRV block is trained using mean-square-error (MSE) as a loss function to represent the video, and the network for the compression is enhanced with quantization- aware-training (QAT), during QAT, the network is finetuned, an arithmetic coding to is applied to quantized network parameters and the quantized network parameters are transformed into a bitstream suitable for transmission and storage. In some embodiments, a decoding process begins with arithmetic decoding, which retrieves the quantized network parameters, and the quantized network parameters are then reloaded into the HiNeRV block, one video unit corresponding to each input coordinate is reconstructed.

[0157] In some embodiments, convolutional next (ConvNext) blocks are repeatedly stacked in the neural representation, and each stacked ConvNeXt block uses the same parameters. For example, during a forward pass, for each input X, an output Y is obtained after passing through an identical ConvNeXt blocks, which is represented as follows:Y = ConvNeXtn( ),26 F1254283PCTwhere X represents the number of repetitions of ConvNext block.

[0158] In some embodiments, output channels of a first linear layer and input channels of the second linear layer in ConvNeXt are expanded. Further, weights of two linear layers in the ConvNext blocks may be concatenated.

[0159] In some embodiments, for an input, a weight matrix is created by concatenating an original weight matrix, which is represented as:VFlnew= concat(W , Wltaxis = 0),W2new=concat(W2, W2, axis = 1), where Wi and W2 represent original weight matrixes, and Winew and W2new represent concatenated weight matrixes, and where an output is obtained by applying the two wider linear operations with Winew and W2new:Y = Conv X)WinewW2new, where Y represents the output, X represents the input, and Winew and W2new represent concatenated weight matrixes.

[0160] In some embodiments, a structure of a network for the compression comprises a hierarchical encoding-based neural representation (HiNeRV) block, and the HiNeRV block begins by mapping an input patch coordinate to a basic input feature:where Xorepresents a basic feature, (i,j, t) represents the input patch coordinate, Ybase represents a basic grid which is an autoregressive embedding in HiNeRV, and Fiinearrepresents a linear layer that adjusts the basic grid to the required number of channels. In some embodiments, the basic input feature enters HiNeRV blocks where the basic input feature is progressively processed and upsampled into an output video patch.

[0161] In some embodiments, each HiNeRV block performs a series of transformations includes: a bilinear interpolation to upsample an input feature X„-i by a scale factor S„; determining a hierarchical grid y„(z, J, t) based on the coordinates (i, j, t) and mapping the hierarchical embedding y„(z, j, f) to the number of channels through a linear layer; adding the upsampled input feature Xn-\ with a processed grid; obtaining features by feeding the upsampled input feature X„-i added with a processed grid into a plurality of ConvNeXt blocks, wherein a ConvNeXt block includes a depthwise separable convolution layer that extracts spatial information, followed by layer normalization; amplifying the features by a linear layer, activated by a GeLU function; obtaining a final output by reducing the features in channels by another linear layer; and processing the final output of a last HiNeRV block by a head layer that maps it to (R, G, B) video frames using a linear layer. In some embodiments, a HiNeRV block is trained using meansquare-error (MSE) as a loss function to represent the video, and the network for the compression is enhanced with quantization-aware-training (QAT), during QAT, the network is finetuned, network parameters of the network is quantized, an arithmetic coding is applied to the quantized network parameters, and the quantized network parameters are transformed into a bitstream suitable for transmission and storage.

[0162] According to further embodiments of the present disclosure, a non-transitory computer-readable27 F1254283PCTrecording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video which is generated by a method performed by an apparatus for video processing. The method comprises: performing a compression on a video unit of the video based on a neural representation, wherein one or more parameters are reused in the neural representation; and generating the bitstream based on the compressed video unit.

[0163] According to still further embodiments of the present disclosure, a method for storing bitstream of a video is provided. The method comprises: performing a compression on a video unit of the video based on a neural representation, wherein one or more parameters are reused in the neural representation; generating the bitstream based on the compressed video unit; and storing the bitstream in a non-transitory computer-readable recording medium.

[0164] Implementations of the present disclosure can be described in view of the following clauses, the features of which can be combined in any reasonable manner.

[0165] Clause 1. A method of video processing, comprising: performing, for a conversion between a video unit of a video and a bitstream of the video, a compression on the video unit based on a neural representation, wherein one or more parameters are reused in the neural representation; and performing the conversion based on the compressed video unit.

[0166] Clause 2. The method of clause 1, wherein the neural representation comprises one or more convolutional layers and stacks a same convolutional layer for multiple times.

[0167] Clause 3. The method of clause 1, wherein the neural representation comprises one or more linear layers and stacks a same convolutional linear for multiple times.

[0168] Clause 4. The method of clause 1, wherein the neural representation comprises one or more complex blocks including one or more convolutional layers and one or more linear layers, and the neural representation stacks a same block for multiple times.

[0169] Clause 5. The method of clause 1, wherein the neural representation comprises one or more convolutional layers and expands a convolutional layers by concatenating a weight of the one or more convolutional layers for multiple times.

[0170] Clause 6. The method of clause 1, wherein the neural representation comprises one or more linear layers and expands a linear layer by concatenating a weight of the one or more linear layers for multiple times.

[0171] Clause 7. The method of clause 1, wherein a structure of a network for the compression comprises mapping an input patch coordinates to an initial input feature: 0 stem(.Ybase(i>j> y) wherein (i, j, t) represents the input patch coordinates, Xo represents the initial input feature, Ybase represents basic embeddings, and Fstemrepresents a stem convolutional layer that adjusts the basic embeddings to a number of channels.

[0172] Clause 8. The method of clause 7, wherein the structure of the network for the compression further comprises one or more video compression with hierarchical encoding-based neural representation (HiNeRV) block where the initial feature is progressively processed and upsampled into a final video patch.28 F1254283PCT

[0173] Clause 9. The method of clause 8, wherein each HiNeRV block performs a series of transformations comprises: a bilinear interpolation to upsample an input feature X„-i by a scale factor S„; determining a hierarchical embedding y„(z, j, f) based on the coordinates (i, j, t) and mapping the hierarchical embedding yni, j, f) to the number of channels through a linear layer; adding the upsampled input feature 1 with a processed embedding; obtaining features by feeding the upsampled input feature Xn- added with a processed embedding into a plurality of ConvNeXt blocks, wherein a ConvNeXt block includes a depthwise separable convolution layer that extracts spatial information, followed by layer normalization; amplifying the features by a linear layer, activated by a GeLU function; obtaining a final output by reducing the features in channels by another linear layer; and processing the final output of a last HiNeRV block by a head layer that maps it to (R, G, B) video frames using a linear layer.

[0174] Clause 10. The method of clause 8, wherein a HiNeRV block is trained using mean-square-error (MSE) as a loss function to represent the video, and the network for the compression is enhanced with quantization-aware-training (QAT), during QAT, the network is finetuned, an arithmetic coding to is applied to quantized network parameters and the quantized network parameters are transformed into a bitstream suitable for transmission and storage.

[0175] Clause 11. The method of clause 10, wherein a decoding process begins with arithmetic decoding, which retrieves the quantized network parameters, and the quantized network parameters are then reloaded into the HiNeRV block, one video unit corresponding to each input coordinate is reconstructed.

[0176] Clause 12. The method of clause 1, wherein convolutional next (ConvNext) blocks are repeatedly stacked in the neural representation, and each stacked ConvNeXt block uses the same parameters.

[0177] Clause 13. The method of clause 12, wherein during a forward pass, for each input X, an output Y is obtained after passing through an identical ConvNeXt blocks, which is represented as follows: , wherein X represents the number of repetitions of ConvNext block.

[0178] Clause 14. The method of clause 1, wherein during a forward pass, for each input X, an output Y is obtained after passing through an identical ConvNeXt blocks, which is represented as follows:Y = ConvNeXtn(X), wherein X represents the number of repetitions of ConvNext block.

[0179] Clause 15. The method of clause 1, wherein for an input, a weight matrix is created by concatenating an original weight matrix, which is represented as:wherein Wi and W2 represent original weight matrixes, and Winew and W2new represent concatenated weight matrixes, and wherein an output is obtained by applying the two wider linear operations with Winew and W2neW:Y = Conv(X)W^newW2new. wherein Y represents the output, X represents the input, and Winew and W2new represent concatenated weight matrixes.

[0180] Clause 16. The method of clause 1, wherein a structure of a network for the compression comprises a hierarchical encoding-based neural representation (HiNeRV) block, and the HiNeRV block begins by29 F1254283PCTmapping an input patch coordinate to a basic input feature:wherein Xorepresents a basic feature, (t,y, t) represents the input patch coordinate, Ybase represents a basic grid which is an autoregressive embedding in HiNeRV, and Fiinearrepresents a linear layer that adjusts the basic grid to the required number of channels.

[0181] Clause 17. The method of clause 16, wherein the basic input feature enters HiNeRV blocks where the basic input feature is progressively processed and upsampled into an output video patch.

[0182] Clause 18. The method of clause 17, wherein each HiNeRV block performs a series of transformations comprises: a bilinear interpolation to upsample an input feature X„-i by a scale factor S„; determining a hierarchical grid y„(z, J, t) based on the coordinates (i, j, t) and mapping the hierarchical embedding y„(z, j, f) to the number of channels through a linear layer; adding the upsampled input feature X„-i with a processed grid; obtaining features by feeding the upsampled input feature „-i added with a processed grid into a plurality of ConvNeXt blocks, wherein a ConvNeXt block includes a depthwise separable convolution layer that extracts spatial information, followed by layer normalization; amplifying the features by a linear layer, activated by a GeLU function; obtaining a final output by reducing the features in channels by another linear layer; and processing the final output of a last HiNeRV block by a head layer that maps it to (R, G, B) video frames using a linear layer.

[0183] Clause 19. The method of clause 16, wherein a HiNeRV block is trained using mean-square-error (MSE) as a loss function to represent the video, and the network for the compression is enhanced with quantization-aware-training (QAT), during QAT, the network is finetuned, network parameters of the network is quantized, an arithmetic coding is applied to the quantized network parameters, and the quantized network parameters are transformed into a bitstream suitable for transmission and storage.

[0184] Clause 20. The method of any of clauses 1-19, wherein the conversion includes encoding the video unit into the bitstream.

[0185] Clause 21. The method of any of clauses 1-19, wherein the conversion includes decoding the video unit from the bitstream.

[0186] Clause 22. An apparatus for video processing comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to perform a method in accordance with any of clauses 1-21.

[0187] Clause 23. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform a method in accordance with any of clauses 1-21.

[0188] Clause 24. A non-transitory computer-readable recording medium storing a bitstream of a video which is generated by a method performed by an apparatus for video processing, wherein the method comprises: performing a compression on a video unit of the video based on a neural representation, wherein one or more parameters are reused in the neural representation; and generating the bitstream based on the compressed video unit.

[0189] Clause 25. A method for storing a bitstream of a video, comprising: performing a compression on a video unit of the video based on a neural representation, wherein one or more parameters are reused in the neural representation; generating the bitstream based on the compressed video unit; and storing the30 F1254283PCTbitstream in a non-transitory computer-readable recording medium.Example Device

[0190] Fig. 8 illustrates a block diagram of a computing device 800 in which various embodiments of the present disclosure can be implemented. The computing device 800 may be implemented as or included in the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300).

[0191] It would be appreciated that the computing device 800 shown in Fig. 8 is merely for purpose of illustration, without suggesting any limitation to the functions and scopes of the embodiments of the present disclosure in any manner.

[0192] As shown in Fig. 8, the computing device 800 includes a general-purpose computing device 800. The computing device 800 may at least comprise one or more processors or processing units 810, a memory 820, a storage unit 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860.

[0193] In some embodiments, the computing device 800 may be implemented as any user terminal or server terminal having the computing capability. The server terminal may be a server, a large-scale computing device or the like that is provided by a service provider. The user terminal may for example be any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, station, unit, device, multimedia computer, multimedia tablet, Internet node, communicator, desktop computer, laptop computer, notebook computer, netbook computer, tablet computer, personal communication system (PCS) device, personal navigation device, personal digital assistant (PDA), audio / video player, digital camera / video camera, positioning device, television receiver, radio broadcast receiver, E-book device, gaming device, or any combination thereof, including the accessories and peripherals of these devices, or any combination thereof. It would be contemplated that the computing device 800 can support any type of interface to a user (such as “wearable” circuitry and the like).

[0194] The processing unit 810 may be a physical or virtual processor and can implement various processes based on programs stored in the memory 820. In a multi-processor system, multiple processing units execute computer executable instructions in parallel so as to improve the parallel processing capability of the computing device 800. The processing unit 810 may also be referred to as a central processing unit (CPU), a microprocessor, a controller or a microcontroller.

[0195] The computing device 800 typically includes various computer storage medium. Such medium can be any medium accessible by the computing device 800, including, but not limited to, volatile and nonvolatile medium, or detachable and non-detachable medium. The memory 820 can be a volatile memory (for example, a register, cache, Random Access Memory (RAM)), a non-volatile memory (such as a Read-Only Memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), or a flash memory), or any combination thereof. The storage unit 830 may be any detachable or non- detachable medium and may include a machine-readable medium such as a memory, flash memory drive, magnetic disk or another other media, which can be used for storing information and / or data and can be accessed in the computing device 800.31 F1254283PCT

[0196] The computing device 800 may further include additional detachable / non-detachable, volatile / non-volatile memory medium. Although not shown in Fig. 8, it is possible to provide a magnetic disk drive for reading from and / or writing into a detachable and non-volatile magnetic disk and an optical disk drive for reading from and / or writing into a detachable non-volatile optical disk. In such cases, each drive may be connected to a bus (not shown) via one or more data medium interfaces.

[0197] The communication unit 840 communicates with a further computing device via the communication medium. In addition, the functions of the components in the computing device 800 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, the computing device 800 can operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs) or further general network nodes.

[0198] The input device 850 may be one or more of a variety of input devices, such as a mouse, keyboard, tracking ball, voice-input device, and the like. The output device 860 may be one or more of a variety of output devices, such as a display, loudspeaker, printer, and the like. By means of the communication unit 840, the computing device 800 can further communicate with one or more external devices (not shown) such as the storage devices and display device, with one or more devices enabling the user to interact with the computing device 800, or any devices (such as a network card, a modem and the like) enabling the computing device 800 to communicate with one or more other computing devices, if required. Such communication can be performed via input / output (I / O) interfaces (not shown).

[0199] In some embodiments, instead of being integrated in a single device, some or all components of the computing device 800 may also be arranged in cloud computing architecture. In the cloud computing architecture, the components may be provided remotely and work together to implement the functionalities described in the present disclosure. In some embodiments, cloud computing provides computing, software, data access and storage service, which will not require end users to be aware of the physical locations or configurations of the systems or hardware providing these services. In various embodiments, the cloud computing provides the services via a wide area network (such as Internet) using suitable protocols. For example, a cloud computing provider provides applications over the wide area network, which can be accessed through a web browser or any other computing components. The software or components of the cloud computing architecture and corresponding data may be stored on a server at a remote position. The computing resources in the cloud computing environment may be merged or distributed at locations in a remote data center. Cloud computing infrastructures may provide the services through a shared data center, though they behave as a single access point for the users. Therefore, the cloud computing architectures may be used to provide the components and functionalities described herein from a service provider at a remote location. Alternatively, they may be provided from a conventional server or installed directly or otherwise on a client device.

[0200] The computing device 800 may be used to implement video encoding / decoding in embodiments of the present disclosure. The memory 820 may include one or more video coding modules 825 having one or more program instructions. These modules are accessible and executable by the processing unit 810 to perform the functionalities of the various embodiments described herein.32 F1254283PCT

[0201] In the example embodiments of performing video encoding, the input device 850 may receive video data as an input 870 to be encoded. The video data may be processed, for example, by the video coding module 825, to generate an encoded bitstream. The encoded bitstream may be provided via the output device 860 as an output 880.

[0202] In the example embodiments of performing video decoding, the input device 850 may receive an encoded bitstream as the input 870. The encoded bitstream may be processed, for example, by the video coding module 825, to generate decoded video data. The decoded video data may be provided via the output device 860 as the output 880.

[0203] While this disclosure has been particularly shown and described with references to example embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the present application as defined by the appended claims. Such variations are intended to be covered by the scope of this present application. As such, the foregoing description of embodiments of the present application is not intended to be limiting.33 F1254283PCT

Claims

I / We Claim:

1. A method of video processing, comprising: performing, for a conversion between a video unit of a video and a bitstream of the video, a compression on the video unit based on a neural representation, wherein one or more parameters are reused in the neural representation; and performing the conversion based on the compressed video unit.

2. The method of claim 1, wherein the neural representation comprises one or more convolutional layers and stacks a same convolutional layer for multiple times.

3. The method of claim 1, wherein the neural representation comprises one or more linear layers and stacks a same convolutional linear for multiple times.

4. The method of claim 1, wherein the neural representation comprises one or more complex blocks including one or more convolutional layers and one or more linear layers, and the neural representation stacks a same block for multiple times.

5. The method of claim 1, wherein the neural representation comprises one or more convolutional layers and expands a convolutional layers by concatenating a weight of the one or more convolutional layers for multiple times.

6. The method of claim 1, wherein the neural representation comprises one or more linear layers and expands a linear layer by concatenating a weight of the one or more linear layers for multiple times.

7. The method of claim 1, wherein a structure of a network for the compression comprises mapping an input patch coordinates to an initial input feature: 0 stem(.Ybase (i>j> y) wherein (i, j, t) represents the input patch coordinates, Xo represents the initial input feature, ybaserepresents basic embeddings, and Fstemrepresents a stem convolutional layer that adjusts the basic embeddings to a number of channels.

8. The method of claim 7, wherein the structure of the network for the compression further comprises one or more video compression with hierarchical encoding-based neural representation (HiNeRV) block where the initial feature is progressively processed and upsampled into a final video patch.

9. The method of claim 8, wherein each HiNeRV block performs a series of transformations comprises: a bilinear interpolation to upsample an input feature Xn-i by a scale factor S„;34 F1254283PCTdetermining a hierarchical embedding y„(i, j, t) based on the coordinates (i, j, t) and mapping the hierarchical embedding yn(i, j, f) to the number of channels through a linear layer; adding the upsampled input feature Xn-i with a processed embedding; obtaining features by feeding the upsampled input feature Xn-i added with a processed embedding into a plurality of ConvNeXt blocks, wherein a ConvNeXt block includes a depthwise separable convolution layer that extracts spatial information, followed by layer normalization; amplifying the features by a linear layer, activated by a GeLU function; obtaining a final output by reducing the features in channels by another linear layer; and processing the final output of a last HiNeRV block by a head layer that maps it to (R, G, B) video frames using a linear layer.

10. The method of claim 8, wherein a HiNeRV block is trained using mean-square-error (MSE) as a loss function to represent the video, and the network for the compression is enhanced with quantization-aware- training (QAT), during QAT, the network is finetuned, an arithmetic coding to is applied to quantized network parameters and the quantized network parameters are transformed into a bitstream suitable for transmission and storage.

11. The method of claim 10, wherein a decoding process begins with arithmetic decoding, which retrieves the quantized network parameters, and the quantized network parameters are then reloaded into the HiNeRV block, one video unit corresponding to each input coordinate is reconstructed.

12. The method of claim 1, wherein convolutional next (ConvNext) blocks are repeatedly stacked in the neural representation, and each stacked ConvNeXt block uses the same parameters.

13. The method of claim 12, wherein dining a forward pass, for each input X, an output Y is obtained after passing through an identical ConvNeXt blocks, which is represented as follows:Y = ConvNeXtn( ), wherein X represents the number of repetitions of ConvNext block.

14. The method of claim 1, wherein output channels of a first linear layer and input channels of the second linear layer in ConvNeXt are expanded, weights of two linear layers in the ConvNext blocks are concatenated.

15. The method of claim 1, wherein for an input, a weight matrix is created by concatenating an original weight matrix, which is represented as:YYnew=concat(W2, W2, axis = 1), wherein Wi and W2 represent original weight matrixes, and Winew and W2new represent concatenated weight matrixes, and wherein an output is obtained by applying the two wider linear operations with Winew and W2neW:35 F1254283PCTY = ConvWWnewWnew, wherein Y represents the output, X represents the input, and Winew and W2new represent concatenated weight matrixes.

16. The method of claim 1, wherein a structure of a network for the compression comprises a hierarchical encoding-based neural representation (HiNeRV) block, and the HiNeRV block begins by mapping an input patch coordinate to a basic input feature:wherein Xorepresents a basic feature, (i,j, t) represents the input patch coordinate, ybaserepresents a basic grid which is an autoregressive embedding in HiNeRV, and Flinearrepresents a linear layer that adjusts the basic grid to the required number of channels.

17. The method of claim 16, wherein the basic input feature enters HiNeRV blocks where the basic input feature is progressively processed and upsampled into an output video patch.

18. The method of claim 17, wherein each HiNeRV block performs a series of transformations comprises : a bilinear interpolation to upsample an input feature X„-i by a scale factor S„; determining a hierarchical grid yn(i, j, f) based on the coordinates (i, j, t) and mapping the hierarchical embedding yn(i, j, f) to the number of channels through a linear layer; adding the upsampled input feature X„-i with a processed grid; obtaining features by feeding the upsampled input feature X„-i added with a processed grid into a plurality of ConvNeXt blocks, wherein a ConvNeXt block includes a depthwise separable convolution layer that extracts spatial information, followed by layer normalization; amplifying the features by a linear layer, activated by a GeLU function; obtaining a final output by reducing the features in channels by another linear layer; and processing the final output of a last HiNeRV block by a head layer that maps it to (R, G, B) video frames using a linear layer.

19. The method of claim 16, wherein a HiNeRV block is trained using mean-square-error (MSE) as a loss function to represent the video, and the network for the compression is enhanced with quantization-aware- training (QAT), during QAT, the network is finetuned, network parameters of the network is quantized, an arithmetic coding is applied to the quantized network parameters, and the quantized network parameters are transformed into a bitstream suitable for transmission and storage.

20. The method of any of claims 1-19, wherein the conversion includes encoding the video unit into the bitstream.

21. The method of any of claims 1-19, wherein the conversion includes decoding the video unit from the bitstream.36 F1254283PCT22. An apparatus for video processing comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to perform a method in accordance with any of claims 1-21.

23. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform a method in accordance with any of claims 1-21.

24. A non-transitory computer-readable recording medium storing a bitstream of a video which is generated by a method performed by an apparatus for video processing, wherein the method comprises: performing a compression on a video unit of the video based on a neural representation, wherein one or more parameters are reused in the neural representation; and generating the bitstream based on the compressed video unit.

25. A method for storing a bitstream of a video, comprising: performing a compression on a video unit of the video based on a neural representation, wherein one or more parameters are reused in the neural representation; generating the bitstream based on the compressed video unit; and storing the bitstream in a non-transitory computer-readable recording medium.37 F1254283PCT

Citation Information

Patent Citations

  • Clustering-based quantization for neural network compression

    US20220261616A1

  • Use of embedded signalling for backward-compatible scaling improvements and super-resolution signalling

    US20220385911A1

  • Inter-frame prediction method, coder, decoder, and storage medium

    WO2023019407A1

  • Information processing device and method

    WO2023248486A1

  • Method, apparatus, and medium for video processing

    WO2024086568A1