Method, apparatus, and medium for video processing

A neural network-based video compression method with a synthesis transform module addresses inefficiencies in existing video coding standards by enhancing coding efficiency and reducing complexity through simplified attention and depth-wise convolution layers.

WO2024149308A9PCT designated stage expired Publication Date: 2025-08-21DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/071688
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-05-20
Filing Date
2024-01-10
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Existing video compression technologies, such as MPEG-2, MPEG-4, ITU-T H. 263, ITU-T H. 264/MPEG-4 AVC, and ITU-T H. 265 HEVC, require further improvements in coding efficiency.

Method used

A neural network-based image compression network with a synthesis transform module, including upscaling and attention modules, is applied to video units to enhance coding efficiency by reducing computational complexity through simplified attention and depth-wise convolution layers.

Benefits of technology

This approach improves video coding efficiency while maintaining quality, reducing computational complexity, and optimizing resource usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024071688_21082025_PF_FP_ABST
    Figure CN2024071688_21082025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is proposed. The method comprises: determining, for a conversion between a video unit of a video and a bitstream of the video, to apply a neural network based image compression network to the video unit, wherein the neural network based image compression network comprises a synthesis transform module that comprises at least one of: one or more upscaling layers and one or more attention modules; and performing the conversion based on the neural network based image compression network.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD, APPARATUS, AND MEDIUM FOR VIDEO PROCESSING

[0001] FIELDS

[0002] Embodiments of the present disclosure relates generally to video processing techniques, and more particularly, to neural network-based image and video compression method with Depth-wise and pixel-wise convolution for simplified synthesis transform.BACKGROUND

[0003] In nowadays, digital video capabilities are being applied in various aspects of peoples’ lives. Multiple types of video compression technologies, such as MPEG-2, MPEG-4, ITU-TH. 263, ITU-TH. 264 / MPEG-4 Part 10 Advanced Video Coding (AVC) , ITU-TH. 265 high efficiency video coding (HEVC) standard, versatile video coding (VVC) standard, have been proposed for video encoding / decoding. However, coding efficiency of video coding techniques is generally expected to be further improved.SUMMARY

[0004] Embodiments of the present disclosure provide a solution for video processing.

[0005] In a first aspect, a method for video processing is proposed. The method comprises: determining, for a conversion between a video unit of a video and a bitstream of the video, to apply a neural network based image compression network to the video unit, wherein the neural network based image compression network comprises a synthesis transform module that comprises at least one of: one or more upscaling layers and one or more attention modules; and performing the conversion based on the neural network based image compression network. According to embodiments of the present disclosure, the synthesis transform subnetwork is modified. In particular, it involves a simplified attention module by earlier downscaling the feature maps and then upscaling the feature maps, depth-wise convolution layers and point-wise convolution layers, such that the computational complexity can be decreased. Moreover, the position of the simplified attention module and depth-wise convolution layers can also be adjusted according to the budget of computational resources.

[0006] In a second aspect, an apparatus for video processing is proposed. The apparatus comprises a processor and a non-transitory memory with instructions thereon. The instructions upon execution by the processor, cause the processor to perform a method in  accordance with the first aspect of the present disclosure.

[0007] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to perform a method in accordance with the first aspect of the present disclosure.

[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream of a video which is generated by a method performed by an apparatus for video processing. The method comprises: determining to apply a neural network based image compression network to a video unit of the video, wherein the neural network based image compression network comprises a synthesis transform module that comprises at least one of: one or more upscaling layers and one or more attention modules; and generating the bitstream based on the neural network based image compression network.

[0009] In a fifth aspect, a method for storing a bitstream of a video is proposed. The method comprises: determining to apply a neural network based image compression network to a video unit of the video, wherein the neural network based image compression network comprises a synthesis transform module that comprises at least one of: one or more upscaling layers and one or more attention modules; generating the bitstream based on the neural network based image compression network; and storing the bitstream in a non-transitory computer-readable medium.

[0010] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Through the following detailed description with reference to the accompanying drawings, the above and other objectives, features, and advantages of example embodiments of the present disclosure will become more apparent. In the example embodiments of the present disclosure, the same reference numerals usually refer to the same components.

[0012] FIG. 1 illustrates a block diagram that illustrates an example video coding  system, in accordance with some embodiments of the present disclosure;

[0013] FIG. 2 illustrates a block diagram that illustrates a first example video encoder, in accordance with some embodiments of the present disclosure;

[0014] FIG. 3 illustrates a block diagram that illustrates an example video decoder, in accordance with some embodiments of the present disclosure;

[0015] FIG. 4 is a schematic diagram illustrating an example transform coding scheme;

[0016] FIG. 5 illustrates example latent representations of an image;

[0017] FIG. 6 is a schematic diagram illustrating an example autoencoder implementing a hyperprior model;

[0018] FIG. 7 is a schematic diagram illustrating an example combined model configured to jointly optimize a context model along with a hyperprior and the autoencoder;

[0019] FIG. 8 illustrates an example encoding process;

[0020] FIG. 9 illustrates an example decoding process;

[0021] FIG. 10 illustrates an example encoder and decoder with wavelet-based transform;

[0022] FIG. 11 illustrates an example output of a forward wavelet-based transform;

[0023] FIG. 12 illustrates an example partitioning of the output of a forward wavelet-based transform;

[0024] FIG. 13 shows an example of synthesis transform module;

[0025] FIG. 14 shows a residual calibrated module;

[0026] FIG. 15 shows a simplified attention module;

[0027] FIG. 16 shows a swin transformer layer;

[0028] FIG. 17 shows a simplified attention module;

[0029] FIG. 18 shows a simplified attention module with depth-wise separated convolution layers for down-scaling and up-scaling;

[0030] FIG. 19 shows a simplified attention module with depth-wise separated  convolution layer and point-wise convolution layer in the residual blocks;

[0031] FIG. 20 illustrates a flowchart of a method for video processing in accordance with embodiments of the present disclosure;

[0032] FIG. 21 shows an example of convolution-based attention block; and

[0033] FIG. 22 illustrates a block diagram of a computing device in which various embodiments of the present disclosure can be implemented.

[0034] Throughout the drawings, the same or similar reference numerals usually refer to the same or similar elements.DETAILED DESCRIPTION

[0035] Principle of the present disclosure will now be described with reference to some embodiments. It is to be understood that these embodiments are described only for the purpose of illustration and help those skilled in the art to understand and implement the present disclosure, without suggesting any limitation as to the scope of the disclosure. The disclosure described herein can be implemented in various manners other than the ones described below.

[0036] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.

[0037] References in the present disclosure to “one embodiment, ” “an embodiment, ” “an example embodiment, ” and the like indicate that the embodiment described may include a particular feature, structure, or characteristic, but it is not necessary that every embodiment includes the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.

[0038] It shall be understood that although the terms “first” and “second” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example,  a first element could be termed a second element, and similarly, a second element could be termed a first element, without departing from the scope of example embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0039] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments. As used herein, the singular forms “a” , “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” , “comprising” , “has” , “having” , “includes” and / or “including” , when used herein, specify the presence of stated features, elements, and / or components etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.

[0040] Example Environment

[0041] Fig. 1 is a block diagram that illustrates an example video coding system 100 that may utilize the techniques of this disclosure. As shown, the video coding system 100 may include a source device 110 and a destination device 120. The source device 110 can be also referred to as a video encoding device, and the destination device 120 can be also referred to as a video decoding device. In operation, the source device 110 can be configured to generate encoded video data and the destination device 120 can be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0042] The video source 112 may include a source such as a video capture device. Examples of the video capture device include, but are not limited to, an interface to receive video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.

[0043] The video data may comprise one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded picture is a coded representation of a picture. The associated data may include sequence parameter sets,  picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The encoded video data may be transmitted directly to destination device 120 via the I / O interface 116 through the network 130A. The encoded video data may also be stored onto a storage medium / server 130B for access by destination device 120.

[0044] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120 which is configured to interface with an external display device.

[0045] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard and other current and / or further standards.

[0046] Fig. 2 is a block diagram illustrating an example of a video encoder 200, which may be an example of the video encoder 114 in the system 100 illustrated in Fig. 1, in accordance with some embodiments of the present disclosure.

[0047] The video encoder 200 may be configured to implement any or all of the techniques of this disclosure. In the example of Fig. 2, the video encoder 200 includes a plurality of functional components. The techniques described in this disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.

[0048] In some embodiments, the video encoder 200 may include a partition unit 201, a predication unit 202 which may include a mode select unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.

[0049] In other examples, the video encoder 200 may include more, fewer, or different functional components. In an example, the predication unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform predication in an IBC mode in which at least one reference picture is a picture where the current video block is located.

[0050] Furthermore, although some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be integrated, but are represented in the example of Fig. 2 separately for purposes of explanation.

[0051] The partition unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0052] The mode select unit 203 may select one of the coding modes, intra or inter, e.g., based on error results, and provide the resulting intra-coded or inter-coded block to a residual generation unit 207 to generate residual block data and to a reconstruction unit 212 to reconstruct the encoded block for use as a reference picture. In some examples, the mode select unit 203 may select a combination of intra and inter predication (CIIP) mode in which the predication is based on an inter predication signal and an intra predication signal. The mode select unit 203 may also select a resolution for a motion vector (e.g., a sub-pixel or integer pixel precision) for the block in the case of inter-predication.

[0053] To perform inter prediction on a current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 to the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the buffer 213 other than the picture associated with the current video block.

[0054] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations for a current video block, for example, depending on whether the current video block is in an I-slice, a P-slice, or a B-slice. As used herein, an “I-slice” may refer to a portion of a picture composed of macroblocks, all of which are based upon macroblocks within the same picture. Further, as used herein, in some aspects, “P-slices” and “B-slices” may refer to portions of a picture composed of macroblocks that are not dependent on macroblocks in the same picture.

[0055] In some examples, the motion estimation unit 204 may perform uni-directional prediction for the current video block, and the motion estimation unit 204 may search reference pictures of list 0 or list 1 for a reference video block for the current video block. The motion estimation unit 204 may then generate a reference index that indicates the reference picture in list 0 or list 1 that contains the reference video block and a motion vector that indicates a spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, a prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate the predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.

[0056] Alternatively, in other examples, the motion estimation unit 204 may perform bi-directional prediction for the current video block. The motion estimation unit 204 may search the reference pictures in list 0 for a reference video block for the current video block and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 204 may then generate reference indexes that indicate the reference pictures in list 0 and list 1 containing the reference video blocks and motion vectors that indicate spatial displacements between the reference video blocks and the current video block. The motion estimation unit 204 may output the reference indexes and the motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate the predicted video block of the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0057] In some examples, the motion estimation unit 204 may output a full set of motion information for decoding processing of a decoder. Alternatively, in some embodiments, the motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.

[0058] In one example, the motion estimation unit 204 may indicate, in a syntax structure associated with the current video block, a value that indicates to the video decoder 300 that the current video block has the same motion information as the another video block.

[0059] In another example, the motion estimation unit 204 may identify, in a syntax structure associated with the current video block, another video block and a motion vector difference (MVD) . The motion vector difference indicates a difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0060] As discussed above, video encoder 200 may predictively signal the motion vector. Two examples of predictive signaling techniques that may be implemented by video encoder 200 include advanced motion vector predication (AMVP) and merge mode signaling.

[0061] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0062] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by the minus sign) the predicted video block (s) of the current video block from the current video block. The residual data of the current video block may include residual video blocks that correspond to different sample components of the samples in the current video block.

[0063] In other examples, there may be no residual data for the current video block for the current video block, for example in a skip mode, and the residual generation unit 207 may not perform the subtracting operation.

[0064] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to a residual video block associated with the current video block.

[0065] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0066] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transforms to the transform coefficient video block, respectively, to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by the predication unit 202 to produce a reconstructed video block associated with the current video block for storage in the buffer 213.

[0067] After the reconstruction unit 212 reconstructs the video block, loop filtering operation may be performed to reduce video blocking artifacts in the video block.

[0068] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.

[0069] Fig. 3 is a block diagram illustrating an example of a video decoder 300, which may be an example of the video decoder 124 in the system 100 illustrated in Fig. 1, in accordance with some embodiments of the present disclosure.

[0070] The video decoder 300 may be configured to perform any or all of the techniques of this disclosure. In the example of Fig. 3, the video decoder 300 includes a plurality of functional components. The techniques described in this disclosure may be shared among the various components of the video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.

[0071] In the example of Fig. 3, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transformation unit 305, and a reconstruction unit 306 and a buffer 307. The video decoder 300 may, in some examples, perform a decoding pass generally reciprocal to the encoding pass described with respect to video encoder 200.

[0072] The entropy decoding unit 301 may retrieve an encoded bitstream. The encoded bitstream may include entropy coded video data (e.g., encoded blocks of video data) . The entropy decoding unit 301 may decode the entropy coded video data, and from the entropy decoded video data, the motion compensation unit 302 may determine motion information  including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 may, for example, determine such information by performing the AMVP and merge mode. AMVP is used, including derivation of several most probable candidates based on data from adjacent PBs and the reference picture. Motion information typically includes the horizontal and vertical motion vector displacement values, one or two reference picture indices, and, in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, a “merge mode” may refer to deriving the motion information from spatially or temporally neighboring blocks.

[0073] The motion compensation unit 302 may produce motion compensated blocks, possibly performing interpolation based on interpolation filters. Identifiers for interpolation filters to be used with sub-pixel precision may be included in the syntax elements.

[0074] The motion compensation unit 302 may use the interpolation filters as used by the video encoder 200 during encoding of the video block to calculate interpolated values for sub-integer pixels of a reference block. The motion compensation unit 302 may determine the interpolation filters used by the video encoder 200 according to the received syntax information and use the interpolation filters to produce predictive blocks.

[0075] The motion compensation unit 302 may use at least part of the syntax information to determine sizes of blocks used to encode frame (s) and / or slice (s) of the encoded video sequence, partition information that describes how each macroblock of a picture of the encoded video sequence is partitioned, modes indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-encoded block, and other information to decode the encoded video sequence. As used herein, in some aspects, a “slice” may refer to a data structure that can be decoded independently from other slices of the same picture, in terms of entropy coding, signal prediction, and residual signal reconstruction. A slice can either be an entire picture or a region of a picture.

[0076] The intra prediction unit 303 may use intra prediction modes for example received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 304 inverse quantizes, i.e., de-quantizes, the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301.  The inverse transform unit 305 applies an inverse transform.

[0077] The reconstruction unit 306 may obtain the decoded blocks, e.g., by summing the residual blocks with the corresponding prediction blocks generated by the motion compensation unit 302 or intra-prediction unit 303. If desired, a deblocking filter may also be applied to filter the decoded blocks in order to remove blockiness artifacts. The decoded video blocks are then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra predication and also produces decoded video for presentation on a display device.

[0078] Some exemplary embodiments of the present disclosure will be described in detailed hereinafter. It should be understood that section headings are used in the present document to facilitate ease of understanding and do not limit the embodiments disclosed in a section to only that section. Furthermore, while certain embodiments are described with reference to Versatile Video Coding or other specific video codecs, the disclosed techniques are applicable to other video coding technologies also. Furthermore, while some embodiments describe video coding steps in detail, it will be understood that corresponding steps decoding that undo the coding will be implemented by a decoder. Furthermore, the term video processing encompasses video coding or compression, video decoding or decompression and video transcoding in which video pixels are represented from one compressed format into another compressed format or at a different compressed bitrate.

[0079] 1. Brief summary

[0080] This present disclosure is related to a neural network-based image and video compression approach where an autoregressive neural network is utilized. The examples target a high efficiency synthesis transform for the decoder, therefore enhancing the quality of reconstruction images with moderate computational complexity. The present disclosure is applicable to both luma and chroma components.

[0081] 2. Introduction

[0082] Deep learning is developing in a variety of areas, such as in computer vision and image processing. Inspired by the successful application of deep learning technology to computer vision areas, neural image / video compression technologies are being studied for application to image / video compression techniques. The neural network is designed based  on interdisciplinary research of neuroscience and mathematics. The neural network has shown strong capabilities in the context of non-linear transform and classification. An example neural network-based image compression algorithm achieves comparable R-D performance with Versatile Video Coding (VVC) , which is a video coding standard developed by the Joint Video Experts Team (JVET) with experts from motion picture experts group (MPEG) and Video coding experts group (VCEG) . Neural network-based video compression is an actively developing research area resulting in continuous improvement of the performance of neural image compression. However, neural network-based video coding is still a largely undeveloped discipline due to the inherent difficulty of the problems addressed by neural networks.

[0083] 2.1 Image / Video Compression

[0084] Image / video compression usually refers to a computing technology that compresses video images into binary code to facilitate storage and transmission. The binary codes may or may not support losslessly reconstructing the original image / video. Coding without data loss is known as lossless compression and coding while allowing for targeted loss of data in known as lossy compression, respectively. Most coding systems employ lossy compression since lossless reconstruction is not necessary in most scenarios. Usually the performance of image / video compression algorithms is evaluated based on a resulting compression ratio and reconstruction quality. Compression ratio is directly related to the number of binary codes resulting from compression, with fewer binary codes resulting in better compression. Reconstruction quality is measured by comparing the reconstructed image / video with the original image / video, with greater similarity resulting in better reconstruction quality.

[0085] Image / video compression techniques can be divided into video coding methods and neural-network-based video compression methods. Video coding schemes adopt transform-based solutions, in which statistical dependency in latent variables, such as discrete cosine transform (DCT) and wavelet coefficients, is employed to carefully hand-engineer entropy codes to model the dependencies in the quantized regime. Neural network-based video compression can be grouped into neural network-based coding tools and end-to-end neural network-based video compression. The former is embedded into existing video codecs as coding tools and only serves as part of the framework, while the latter is a separate framework developed based on neural networks without depending on video codecs.

[0086] A series of video coding standards have been developed to accommodate the increasing demands of visual content transmission. The international organization for standardization (ISO)  / International Electrotechnical Commission (IEC) has two expert groups, namely Joint Photographic Experts Group (JPEG) and Moving Picture Experts Group (MPEG) . International Telecommunication Union (ITU) telecommunication standardization sector (ITU-T) also has a Video Coding Experts Group (VCEG) , which is for standardization of image / video coding technology. The influential video coding standards published by these organizations include Joint Photographic Experts Group (JPEG) , JPEG 2000, H. 262, H. 264 / advanced video coding (AVC) and H. 265 / High Efficiency Video Coding (HEVC) . The Joint Video Experts Team (JVET) , formed by MPEG and VCEG, developed the Versatile Video Coding (VVC) standard. An average of 50%bitrate reduction is reported by VVC under the same visual quality compared with HEVC.

[0087] Neural network-based image / video compression / coding is also under development. Example neural network coding network architectures are relatively shallow, and the performance of such networks is not satisfactory. Neural network-based methods benefit from the abundance of data and the support of powerful computing resources, and are therefore better exploited in a variety of applications. Neural network-based image / video compression has shown promising improvements and is confirmed to be feasible. Nevertheless, this technology is far from mature and a lot of challenges should be addressed.

[0088] 2.2 Neural Networks

[0089] Neural networks, also known as artificial neural networks (ANN) , are computational models used in machine learning technology. Neural networks are usually composed of multiple processing layers, and each layer is composed of multiple simple but non-linear basic computational units. One benefit of such deep networks is a capacity for processing data with multiple levels of abstraction and converting data into different kinds of representations. Representations created by neural networks are not manually designed. Instead, the deep network including the processing layers is learned from massive data using a general machine learning procedure. Deep learning eliminates the necessity of handcrafted representations. Thus, deep learning is regarded useful especially for processing natively unstructured data, such as acoustic and visual signals. The processing of such data has been a longstanding difficulty in the artificial intelligence  field.

[0090] 2.3 Neural Networks For Image Compression

[0091] Neural networks for image compression can be classified in two categories, including pixel probability models and auto-encoder models. Pixel probability models employ a predictive coding strategy. Auto-encoder models employ a transform-based solution. Sometimes, these two methods are combined together.

[0092] 2.3.1 Pixel Probability Modeling

[0093] According to Shannon’s information theory, the optimal method for lossless coding can reach the minimal coding rate, which is denoted as -log2p (x) where p (x) is the probability of symbol x. Arithmetic coding is a lossless coding method that is believed to be among the optimal methods. Given a probability distribution p (x) , arithmetic coding causes the coding rate to be as close as possible to a theoretical limit -log2p (x) without considering the rounding error. Therefore, the remaining problem is to determine the probability, which is very challenging for natural image / video due to the curse of dimensionality. The curse of dimensionality refers to the problem that increasing dimensions causes data sets to become sparse, and hence rapidly increasing amounts of data is needed to effectively analyze and organize data as the number of dimensions increases.

[0094] Following the predictive coding strategy, one way to model p (x) is to predict pixel probabilities one by one in a raster scan order based on previous observations, where x is an image, can be expressed as follows: p (x) =p (x1) p (x2|x1) …p (xi|x1, …, xi-1) …p (xm×n|x1, …, xm×n-1)          (1)

[0095] where m and n are the height and width of the image, respectively. The previous observation is also known as the context of the current pixel. When the image is large, estimation of the conditional probability can be difficult. Thereby, a simplified method is to limit the range of the context of the current pixel as follows: p (x) =p (x1) p (x2|x1) …p (xi|xi-k, …, xi-1) …p (xm×n|xm×n-k, …, xm×n-1)        (2)

[0096] where k is a pre-defined constant controlling the range of the context.

[0097] It should be noted that the condition may also take the sample values of other color components into consideration. For example, when coding the red (R) , green (G) ,  and blue (B) (RGB) color component, the R sample is dependent on previously coded pixels (including R, G, and / or B samples) , the current G sample may be coded according to previously coded pixels and the current R sample. Further, when coding the current B sample, the previously coded pixels and the current R and G samples may also be taken into consideration.

[0098] Neural networks may be designed for computer vision tasks, and may also be effective in regression and classification problems. Therefore, neural networks may be used to estimate the probability of p (xi) given a context x1, x2, …, xi-1. In an example neural network design, the pixel probability is employed for binary images according to xi∈ {-1, +1} . The neural autoregressive distribution estimator (NADE) is designed for pixel probability modeling. NADE is a feed-forward network with a single hidden layer. In another example, the feed-forward network may include connections skipping the hidden layer. Further, the parameters may also be shared. Example designs perform experiments on the binarized MNIST dataset. In an example, NADE is extended to a real-valued NADE (RNADE) model, where the probability p (xi|x1, …, xi-1) is derived with a mixture of Gaussians. The RNADE model feed-forward network also has a single hidden layer, but the hidden layer employs rescaling to avoid saturation and uses a rectified linear unit (ReLU) instead of sigmoid. In another example, NADE and RNADE are improved by using reorganizing the order of the pixels and with deeper neural networks.

[0099] Designing advanced neural networks plays an important role in improving pixel probability modeling. In an example neural network, a multi-dimensional long short-term memory (LSTM) is used. The LSTM works together with mixtures of conditional Gaussian scale mixtures for probability modeling. LSTM is a special kind of recurrent neural networks (RNNs) and may be employed to model sequential data. The spatial variant of LSTM may also be used for images later. Several different neural networks may be employed, including recurrent neural networks (RNNs) and convolutional neural networks (CNNs) , such as Pixel RNN (PixelRNN) and Pixel CNN (PixelCNN) , respectively. In PixelRNN, two variants of LSTM, denoted as row LSTM and diagonal bidirectional LSTM (BiLSTM) are employed. Diagonal BiLSTM is specifically designed for images. PixelRNN incorporates residual connections to help train deep neural networks with up to twelve layers. In PixelCNN, masked convolutions are used to adjust for the shape of the context. PixelRNN and PixelCNN are more dedicated to natural images. For example, PixelRNN and PixelCNN consider pixels as discrete values (e.g.,  0, 1, …, 255) and predict a multinomial distribution over the discrete values. Further, PixelRNN and PixelCNN deal with color images in RGB color space. In addition, PixelRNN and PixelCNN work well on the large-scale image dataset image network (ImageNet) . In an example, a Gated PixelCNN is used to improve the PixelCNN. Gated PixelCNN achieves comparable performance with PixelRNN, but with much less complexity. In an example, a PixelCNN++ is employed with the following improvements upon PixelCNN: a discretized logistic mixture likelihood is used rather than a 256-way multinomial distribution; down-sampling is used to capture structures at multiple resolutions; additional short-cut connections are introduced to speed up training; dropout is adopted for regularization; and RGB is combined for one pixel. In another example, PixelSNAIL combines casual convolutions with self-attention.

[0100] Most of the above methods directly model the probability distribution in the pixel domain. Some designs also model the probability distribution as conditional based upon explicit or latent representations. Such a model can be expressed as:

[0101] where h is the additional condition and p (x) =p (h) p (x|h) indicates the modeling is split into an unconditional model and a conditional model. The additional condition can be image label information or high-level representations.

[0102] 2.3.2 Auto-encoder

[0103] An Auto-encoder is now described. The auto-encoder is trained for dimensionality reduction and include an encoding component and a decoding component. The encoding component converts the high-dimension input signal to low-dimension representations. The low-dimension representations may have reduced spatial size, but a greater number of channels. The decoding component recovers the high-dimension input from the low-dimension representation. The auto-encoder enables automated learning of representations and eliminates the need of hand-crafted features, which is also believed to be one of the most important advantages of neural networks.

[0104] FIG. 4 is a schematic diagram illustrating an example transform coding scheme 400. The original image x is transformed by the analysis network ga to achieve the latent representation y. The latent representation y is quantized (q) and compressed into bits. The number of bits R is used to measure the coding rate. The quantized latent  representation is then inversely transformed by a synthesis network gs to obtain the reconstructed image The distortion (D) is calculated in a perceptual space by transforming x and with the function gp, resulting in z and which are compared to obtain D.

[0105] An auto-encoder network can be applied to lossy image compression. The learned latent representation can be encoded from the well-trained neural networks. However, adapting the auto-encoder to image compression is not trivial since the original auto-encoder is not optimized for compression, and is thereby not efficient for direct use as a trained auto-encoder. In addition, other major challenges exist. First, the low-dimension representation should be quantized before being encoded. However, the quantization is not differentiable, which is required in backpropagation while training the neural networks. Second, the objective under a compression scenario is different since both the distortion and the rate need to be take into consideration. Estimating the rate is challenging. Third, a practical image coding scheme should support variable rate, scalability, encoding / decoding speed, and interoperability. In response to these challenges, various schemes are under development.

[0106] An example auto-encoder for image compression using the example transform coding scheme 400 can be regarded as a transform coding strategy. The original image x is transformed with the analysis network y=ga (x) , where y is the latent representation to be quantized and coded. The synthesis network inversely transforms the quantized latent representation back to obtain the reconstructed image The framework is trained with the rate-distortion loss function,  where D is the distortion between x and R is the rate calculated or estimated from the quantized representation  and λ is the Lagrange multiplier. D can be calculated in either pixel domain or perceptual domain. Most example systems follow this prototype and the differences between such systems might only be the network structure or loss function.

[0107] In terms of network structure, RNNs and CNNs are the most widely used architectures. In the RNNs relevant category, an example general framework for variable rate image compression uses RNN. The example uses binary quantization to generate codes and does not consider rate during training. The framework provides a scalable coding functionality, where RNN with convolutional and deconvolution layers performs well. Another example offers an improved version by upgrading the encoder with a neural  network similar to PixelRNN to compress the binary codes. The performance is better than JPEG on a Kodak image dataset using multi-scale structural similarity (MS-SSIM) evaluation metric. Another example further improves the RNN-based solution by introducing hidden-state priming. In addition, an SSIM-weighted loss function is also designed, and a spatially adaptive bitrates mechanism is included. This example achieves better results than better portable graphics (BPG) on the Kodak image dataset using MS-SSIM as evaluation metric. Another example system supports spatially adaptive bitrates by training stop-code tolerant RNNs.

[0108] Another example proposes a general framework for rate-distortion optimized image compression. The example system uses multiary quantization to generate integer codes and considers the rate during training. The loss is the joint rate-distortion cost, which can be mean square error (MSE) or other metrics. The example system adds random uniform noise to stimulate the quantization during training and uses the differential entropy of the noisy codes as a proxy for the rate. The example system uses generalized divisive normalization (GDN) as the network structure, which includes a linear mapping followed by a nonlinear parametric normalization. The effectiveness of GDN on image coding is verified. Another example system includes improved version that uses three convolutional layers each followed by a down-sampling layer and a GDN layer as the forward transform. Accordingly, this example version uses three layers of inverse GDN each followed by an up-sampling layer and convolution layer to stimulate the inverse transform. In addition, an arithmetic coding method is devised to compress the integer codes. The performance is reportedly better than JPEG and JPEG 2000 on Kodak dataset in terms of MSE. Another example improves the method by devising a scale hyper-prior into the auto-encoder. The system transforms the latent representation y with a subnet ha to z=ha (y) and z is quantized and transmitted as side information. Accordingly, the inverse transform is implemented with a subnet hs that decodes from the quantized side information to the standard deviation of the quantized which is further used during the arithmetic coding of On the Kodak image set, this method is slightly worse than BGP in terms of peak signal to noise ratio (PSNR) . Another example system further exploits the structures in the residue space by introducing an autoregressive model to estimate both the standard deviation and the mean. This example uses a Gaussian mixture model to further remove redundancy in the residue. The performance is on par with VVC on the Kodak image set using PSNR as evaluation metric.

[0109] 2.3.3 Hyper Prior Model

[0110] FIG. 5 illustrates example latent representations of an image. FIG. 5 includes an image 501 from the Kodak dataset, via isualization of the latent 502 representation y of the image 501, a standard deviations σ 503 of the latent 502, and latents y 504 after a hyper prior network is introduced. A hyper prior network includes a hyper encoder and decoder. In the transform coding approach to image compression, as shown in FIG. 4, the encoder subnetwork transforms the image vector x using a parametric analysis transform  into a latent representation y, which is then quantized to form Because is discrete-valued,  can be losslessly compressed using entropy coding techniques such as arithmetic coding and transmitted as a sequence of bits.

[0111] As evident from the latent 502 and the standard deviations σ 503 of FIG. 5, there are significant spatial dependencies among the elements of Notably, their scales (standard deviations σ 503) appear to be coupled spatially. An additional set of random variables may be introduced to capture the spatial dependencies and to further reduce the redundancies. In this case the image compression network is depicted in FIG. 6.

[0112] FIG. 6 is a schematic diagram 600 illustrating an example network architecture of an autoencoder implementing a hyperprior model. The upper side shows an image autoencoder network, and the lower side corresponds to the hyperprior subnetwork. The analysis and synthesis transforms are denoted as ga and gs, respectively. Q represents quantization, and AE, AD represent arithmetic encoder and arithmetic decoder, respectively. The hyperprior model includes two subnetworks, hyper encoder (denoted with ha) and hyper decoder (denoted with hs) . The hyper prior model generates a quantized hyper latent which comprises information related to the probability distribution of the samples of the quantized latent is included in the bitstream and transmitted to the receiver (decoder) along with

[0113] In schematic diagram 600, the upper side of the models is the encoder ga and decoder gs as discussed above. The lower side is the additional hyper encoder ha and hyper decoder hs networks that are used to obtain In this architecture the encoder subjects the input image x to ga, yielding the responses y with spatially varying standard deviations. The responses y are fed into ha, summarizing the distribution of standard deviations in z. z is then quantized compressed, and transmitted as side information. The encoder then uses the quantized vector to estimate σ, the spatial distribution of  standard deviations, and uses σ to compress and transmit the quantized image representation The decoder first recovers  from the compressed signal. The decoder then uses hs to obtain σ, which provides the decoder with the correct probability estimates to successfully recover as well. The decoder then feeds into gs to obtain the reconstructed image.

[0114] When the hyper encoder and hyper decoder are added to the image compression network, the spatial redundancies of the quantized latent are reduced. The latents y 504 in FIG. 5 correspond to the quantized latent when the hyper encoder / decoder are used. Compared to the standard deviations σ 503, the spatial redundancies are significantly reduced as the samples of the quantized latent are less correlated.

[0115] 2.3.4 Context Model

[0116] Although the hyperprior model improves the modelling of the probability distribution of the quantized latent additional improvement can be obtained by utilizing an autoregressive model that predicts quantized latents from their causal context, which may be known as a context model.

[0117] The term auto-regressive indicates that the output of a process is later used as an input to the process. For example, the context model subnetwork generates one sample of a latent, which is later used as input to obtain the next sample.

[0118] FIG. 7 is a schematic diagram 700 illustrating an example combined model configured to jointly optimize a context model along with a hyperprior and the autoencoder. The combined model jointly optimizes an autoregressive component that estimates the probability distributions of latents from their causal context (Context Model) along with a hyperprior and the underlying autoencoder. Real-valued latent representations are quantized (Q) to create quantized latents and quantized hyper-latents which are compressed into a bitstream using an arithmetic encoder (AE) and decompressed by an arithmetic decoder (AD) . The dashed region corresponds to the components that are executed by the receiver (e.g, a decoder) to recover an image from a compressed bitstream.

[0119] An example system utilizes a joint architecture where both a hyperprior model subnetwork (hyper encoder and hyper decoder) and a context model subnetwork are utilized. The hyperprior and the context model are combined to learn a probabilistic  model over quantized latents which is then used for entropy coding. As depicted in schematic diagram 700, the outputs of the context subnetwork and hyper decoder subnetwork are combined by the subnetwork called Entropy Parameters, which generates the mean μ and scale (or variance) σ parameters for a Gaussian probability model. The gaussian probability model is then used to encode the samples of the quantized latents into bitstream with the help of the arithmetic encoder (AE) module. In the decoder the gaussian probability model is utilized to obtain the quantized latents from the bitstream by arithmetic decoder (AD) module.

[0120] In an example, the latent samples are modeled as gaussian distribution or gaussian mixture models (not limited to) . In the example according to the schematic diagram 700, the context model and hyper prior are jointly used to estimate the probability distribution of the latent samples. Since a gaussian distribution can be defined by a mean and a variance (aka sigma or scale) , the joint model is used to estimate the mean and variance (denoted as μ and σ) .

[0121] 2.3.5 The encoding process using joint auto-regressive hyper prior model

[0122] The design in FIG 4. corresponds an example combined compression method. In this section and the next, the encoding and decoding processes are described separately.

[0123] FIG. 8 illustrates an example encoding process 800. The input image is first processed with an encoder subnetwork. The encoder transforms the input image into a transformed representation called latent, denoted by y. y is then input to a quantizer block, denoted by Q, to obtain the quantized latent is then converted to a bitstream (bits1) using an arithmetic encoding module (denoted AE) . The arithmetic encoding block converts each sample of the into a bitstream (bits1) one by one, in a sequential order.

[0124] The modules hyper encoder, context, hyper decoder, and entropy parameters subnetworks are used to estimate the probability distributions of the samples of the quantized latent the latent y is input to hyper encoder, which outputs the hyper latent (denoted by z) . The hyper latent is then quantized and a second bitstream (bits2) is generated using arithmetic encoding (AE) module. The factorized entropy module generates the probability distribution, that is used to encode the quantized hyper latent into bitstream. The quantized hyper latent includes information about the probability distribution of the quantized latent

[0125] The Entropy Parameters subnetwork generates the probability distribution estimations, that are used to encode the quantized latent The information that is generated by the Entropy Parameters typically include a mean μ and scale (or variance) σparameters, that are together used to obtain a gaussian probability distribution. A gaussian distribution of a random variable x is defined as where the parameter μ is the mean or expectation of the distribution (and also its median and mode) , while the parameter σ is its standard deviation (or variance, or scale) . In order to define a gaussian distribution, the mean and the variance need to be determined. The entropy parameters module are used to estimate the mean and the variance values.

[0126] The subnetwork hyper decoder generates part of the information that is used by the entropy parameters subnetwork, the other part of the information is generated by the autoregressive module called context module. The context module generates information about the probability distribution of a sample of the quantized latent, using the samples that are already encoded by the arithmetic encoding (AE) module. The quantized latent is typically a matrix composed of many samples. The samples can be indicated using indices, such as or depending on the dimensions of the matrix The samples are encoded by AE one by one, typically using a raster scan order. In a raster scan order the rows of a matrix are processed from top to bottom, where the samples in a row are processed from left to right. In such a scenario (where the raster scan order is used by the AE to encode the samples into bitstream) , the context module generates the information pertaining to a sample using the samples encoded before, in raster scan order. The information generated by the context module and the hyper decoder are combined by the entropy parameters module to generate the probability distributions that are used to encode the quantized latent into bitstream (bits1) .

[0127] Finally, the first and the second bitstream are transmitted to the decoder as result of the encoding process. It is noted that the other names can be used for the modules described above.

[0128] In the above description, all of the elements in FIG. 8 are collectively called an encoder. The analysis transform that converts the input image into latent representation is also called an encoder (or auto-encoder) .

[0129] 2.3.6 The decoding process using joint auto-regressive hyper prior model

[0130] FIG. 9 illustrates an example decoding process 900. FIG. 9 depicts a decoding process separately.

[0131] In the decoding process, the decoder first receives the first bitstream (bits1) and the second bitstream (bits2) that are generated by a corresponding encoder. The bits2 is first decoded by the arithmetic decoding (AD) module by utilizing the probability distributions generated by the factorized entropy subnetwork. The factorized entropy module typically generates the probability distributions using a predetermined template, for example using predetermined mean and variance values in the case of gaussian distribution. The output of the arithmetic decoding process of the bits2 is which is the quantized hyper latent. The AD process reverts to AE process that was applied in the encoder. The processes of AE and AD are lossless, meaning that the quantized hyper latent  that was generated by the encoder can be reconstructed at the decoder without any change.

[0132] After obtaining of it is processed by the hyper decoder, whose output is fed to entropy parameters module. The three subnetworks, context, hyper decoder and entropy parameters that are employed in the decoder are identical to the ones in the encoder. Therefore, the exact same probability distributions can be obtained in the decoder (as in encoder) , which is essential for reconstructing the quantized latent without any loss. As a result, the identical version of the quantized latent that was obtained in the encoder can be obtained in the decoder.

[0133] After the probability distributions (e.g. the mean and variance parameters) are obtained by the entropy parameters subnetwork, the arithmetic decoding module decodes the samples of the quantized latent one by one from the bitstream bits1. From a practical standpoint, autoregressive model (the context model) is inherently serial, and therefore cannot be sped up using techniques such as parallelization. Finally, the fully reconstructed quantized latent is input to the synthesis transform (denoted as decoder in FIG. 9) module to obtain the reconstructed image.

[0134] In the above description, the all of the elements in FIG. 9 are collectively called decoder. The synthesis transform that converts the quantized latent into reconstructed image is also called a decoder (or auto-decoder) .

[0135] 2.3.7 Wavelet based neural compression architecture

[0136] The analysis transform (denoted as encoder) in FIG. 8 and the synthesis transform (denoted as decoder) in FIG. 9 might be replaced by a wavelet based transform. FIG. 10 below shows an example such an implementation. In the figure first the input image is converted from an RGB color format to a YUV color format. This conversion process is optional, and can be missing in other implementations. If however such a conversion is applied at the input image, a back conversion (from YUV to RGB) is also applied before the output image is generated. Moreover there are 2 additional post processing modules (post-process 1 and 2) shown in the figure. These modules are also optional, hence might be missing in other implementations. The core of an encoder with wavelet-based transform is composed of a wavelet-based forward transform, a quantization module and an entropy coding module. After these 3 modules are applied to the input image, the bitstream is generated. The core of the decoding process is composed of entropy decoding, de-quantization process and an inverse wavelet-based transform operation. The decoding process convers the bitstream into output image. The encoding and decoding processes are depicted FIG. 10.

[0137] FIG. 10 illustrates an example encoder and decoder 1000 with wavelet-based transform.

[0138] After the wavelet-based forward transform is applied to the input image, in the output of the wavelet-based forward transform the image is split into its frequency components. The output of a 2-dimensional forward wavelet transform (depicted as iWave forward module in the figure above) might take the form depicted in FIG. 11. The input of the transform is an image of a castle. In the example, after the transform an output with 7 distinct regions are obtained. The number of distinct regions depend on the specific implementation of the transform and might different from 7. Potential number of regions are 4, 7, 10, 13, …

[0139] FIG. 11 illustrates an example output 1100 of a forward wavelet-based transform.

[0140] In FIG. 11, the input image is transformed into 7 regions with 3 small images and 4 even smaller images. The transformation is based on the frequency components, the small image at the bottom right quarter comprises the high frequency components in both horizontal and vertical directions. The smallest image at the top-left corner on the other hand comprises the lowest frequency components both in the vertical and horizontal directions. The small image on the top-right quarter comprises the high frequency  components in the horizontal direction and low frequency components in the vertical direction.

[0141] FIG. 12 illustrates an example partitioning 1200 of the output of a forward wavelet-based transform. FIG. 12 depicts a possible splitting of the latent representation after the 2D forward transform. The latent representation are the samples (latent samples, or quantized latent samples) that are obtained after the 2D forward transform. The latent samples are divided into 7 sections above, denoted as HH1, LH1, HL1, LL2, HL2, LH2 and HH2. The HH1 describes that the section comprises high frequency components in the vertical direction, high frequency components in the horizontal direction and that the splitting depth is 1. HL2 describes that the section comprises low frequency components in the vertical direction, high frequency components in the horizontal direction and that the splitting depth is 2.

[0142] After the latent samples are obtained at the encoder by the forward wavelet transform, they are transmitted to the decoder by using entropy coding. At the decoder, entropy decoding is applied to obtain the latent samples, which are then inverse transformed (by using iWave inverse module in FIG. 10) to obtain the reconstructed image.

[0143] 2.4 Neural Networks for Video Compression

[0144] Similar to video coding technologies, neural image compression serves as the foundation of intra compression in neural network-based video compression. Thus, development of neural network-based video compression technology is behind development of neural network-based image compression because neural network-based video compression technology is of greater complexity and hence needs far more effort to solve the corresponding challenges. Compared with image compression, video compression needs efficient methods to remove inter-picture redundancy. Inter-picture prediction is then a major step in these example systems. Motion estimation and compensation is widely adopted in video codecs, but is not generally implemented by trained neural networks.

[0145] Neural network-based video compression can be divided into two categories according to the targeted scenarios: random access and the low-latency. In random access case, the system allows decoding to be started from any point of the sequence, typically divides the entire sequence into multiple individual segments, and allows each segment to be decoded independently. In a low-latency case, the system aims to reduce decoding  time, and thereby temporally previous frames can be used as reference frames to decode subsequent frames.

[0146] 2.4.1 Low-latency

[0147] An example system employs a video compression scheme with trained neural networks. The system first splits the video sequence frames into blocks and each block is coded according to an intra coding mode or an inter coding mode. If intra coding is selected, there is an associated auto-encoder to compress the block. If inter coding is selected, motion estimation and compensation are performed and a trained neural network is used for residue compression. The outputs of auto-encoders are directly quantized and coded by the Huffman method.

[0148] Another neural network-based video coding scheme employs PixelMotionCNN. The frames are compressed in the temporal order, and each frame is split into blocks which are compressed in the raster scan order. Each frame is first extrapolated with the preceding two reconstructed frames. When a block is to be compressed, the extrapolated frame along with the context of the current block are fed into the PixelMotionCNN to derive a latent representation. Then the residues are compressed by a variable rate image scheme. This scheme performs on par with H. 264.

[0149] Another example system employs an end-to-end neural network-based video compression framework, in which all the modules are implemented with neural networks. The scheme accepts a current frame and a prior reconstructed frame as inputs. An optical flow is derived with a pre-trained neural network as the motion information. The motion information is warped with the reference frame followed by a neural network generating the motion compensated frame. The residues and the motion information are compressed with two separate neural auto-encoders. The whole framework is trained with a single rate-distortion loss function. The example system achieves better performance than H. 264.

[0150] Another example system employs an advanced neural network-based video compression scheme. The system inherits and extends video coding schemes with neural networks with the following major features. First the system uses only one auto-encoder to compress motion information and residues. Second, the system uses motion compensation with multiple frames and multiple optical flows. Third, the system uses an on-line state that is learned and propagated through the following frames over time. This scheme achieves better performance in MS-SSIM than HEVC reference software.

[0151] Another example system uses an extended end-to-end neural network-based video compression framework. In this example, multiple frames are used as references. The example system is thereby able to provide more accurate prediction of a current frame by using multiple reference frames and associated motion information. In addition, a motion field prediction is deployed to remove motion redundancy along temporal channel. Postprocessing networks are also used to remove reconstruction artifacts from previous processes. The performance of this system is better than H. 265 by a noticeable margin in terms of both PSNR and MS-SSIM.

[0152] Another example system uses scale-space flow to replace an optical flow by adding a scale parameter based on a framework. This example system may achieve better performance than H. 264. Another example system uses a multi-resolution representation for optical flows based. Concretely, the motion estimation network produces multiple optical flows with different resolutions and let the network learn which one to choose under the loss function. The performance is slightly better than H. 265.

[0153] 2.4.2 Random Access

[0154] Another example system uses a neural network-based video compression scheme with frame interpolation. The key frames are first compressed with a neural image compressor and the remaining frames are compressed in a hierarchical order. The system performs motion compensation in the perceptual domain by deriving the feature maps at multiple spatial scales of the original frame and using motion to warp the feature maps. The results are used for the image compressor. The method is on par with H. 264.

[0155] An example system uses a method for interpolation-based video compression. The interpolation model combines motion information compression and image synthesis. The same auto-encoder is used for image and residual. Another example system employs a neural network-based video compression method based on variational auto-encoders with a deterministic encoder. Concretely, the model includes an auto-encoder and an auto-regressive prior. Different from previous methods, this system accepts a group of pictures (GOP) as inputs and incorporates a three dimensional (3D) autoregressive prior by taking into account of the temporal correlation while coding the latent representations. This system provides comparative performance as H. 265.

[0156] 2.5 Preliminaries

[0157] Almost all the natural image and / or video is in digital format. A grayscale digital image can be represented by where is the set of values of a pixel, m is the image height, and n is the image width. For example,  is an example setting, and in this case Thus, the pixel can be represented by an 8-bit integer. An uncompressed grayscale digital image has 8 bits-per-pixel (bpp) , while compressed bits are definitely less.

[0158] A color image is typically represented in multiple channels to record the color information. For example, in the RGB color space an image can be denoted by with three separate channels storing Red, Green, and Blue information. Similar to the 8-bit grayscale image, an uncompressed 8-bit RGB image has 24 bpp. Digital images / videos can be represented in different color spaces. The neural network-based video compression schemes are mostly developed in RGB color space while the video codecs typically use a YUV color space to represent the video sequences. In YUV color space, an image is decomposed into three channels, namely luma (Y) , blue difference choma (Cb) and red difference chroma (Cr) . Y is the luminance component and Cb and Cr are the chroma components. The compression benefit to YUV occur because Cb and Cr are typically down sampled to achieve pre-compression since human vision system is less sensitive to chroma components.

[0159] A color video sequence is composed of multiple color images, also called frames, to record scenes at different timestamps. For example, in the RGB color space, a color video can be denoted by X= {x0, x1, …, xt, …, xT-1} where T is the number of frames in a video sequence and If m=1080, n=1920,  and the video has 50 frames-per-second (fps) , then the data rate of this uncompressed video is 1920×1080×8×3×50=2,488,320,000 bits-per-second (bps) . This results in about 2.32 gigabits per second (Gbps) , which uses a lot storage and should be compressed before transmission over the internet.

[0160] Usually the lossless methods can achieve a compression ratio of about 1.5 to 3 for natural images, which is clearly below streaming requirements. Therefore, lossy compression is employed to achieve a better compression ratio, but at the cost of incurred distortion. The distortion can be measured by calculating the average squared difference between the original image and the reconstructed image, for example based on MSE. For a grayscale image, MSE can be calculated with the following equation.

[0161] Accordingly, the quality of the reconstructed image compared with the original image can be measured by peak signal-to-noise ratio (PSNR) :

[0162] where is the maximal value in e.g., 255 for 8-bit grayscale images. There are other quality evaluation metrics such as structural similarity (SSIM) and multi-scale SSIM (MS-SSIM) .

[0163] To compare different lossless compression schemes, the compression ratio given the resulting rate, or vice versa, can be compared. However, to compare different lossy compression methods, the comparison has to take into account both the rate and reconstructed quality. For example, this can be accomplished by calculating the relative rates at several different quality levels and then averaging the rates. The average relative rate is known as Bjontegaard’s delta-rate (BD-rate) . There are other aspects to evaluate image and / or video coding schemes, including encoding / decoding complexity, scalability, robustness, and so on.

[0164] 3. Technical problems solved by disclosed technical solutions

[0165] 3.1. The core problem

[0166] The state-of-the-art image compression networks typically include a synthesis transform module, for example a convolutional neural network with attention modules, to reconstruct the image from the latent codes. A deep synthesis transform module can facilitate the recovery of visual signals. However, the optimal tradeoff between the complexity and its reconstruction capability is challenging. A heavy design is usually accompanied with high computational complexity, which impedes its practical usuage. In contrast, an oversimple synthesis transform module may have limited ability of reconstructing the image, which deteriorates the quality of the decompressed image. Existing methods rarely consider simplifying the design of synthesis module while maintaining its reconstruction capability.

[0167] 3.2 Background and details of the problem

[0168] FIG. 13 shows an example of synthesis transform module. In particular, in FIG.  13, an example of synthesis transform module employed in verification model of JPEG-AI is elaborated. Given the luma latent code  (i.e. input) , the luma reconstruction  (e.g. output) can be obtained through the synthesis transform module. The synthesis transform module may contain residual blocks, deconvolutional layers, non-linear activations, cropping layers, and attention modules, convolution layers, and post-processing modules. The computational complexity of attention module and post-processing are prominent, which takes over 80%of the overall synthesis complexity. In the example depicted in FIG. 13, totally 4 upscaling layers (e.g. deconvolution layers) are involved, corresponding to 4 steps of the reconstruction process, where each step may contain one deconvolutional layer, one cropping layer, and one non-linear activation layer. The upscaling (or upsampling or deconvolution) layer is depicted in the figure using Conv Cpx3x3 / 2↑. Furthermore, in the example in Figure 10, an attention module (RNAB) is depicted. The bock diagrams of the RNAB, ResAU and Residual Block (RB) are also depicted in addition to the block diagram of the synthesis transform.

[0169] The abovementioned upsampling (e.g. deconvolution or upscaling) layer increases the width and height of the input tensor. For example, the upsampling layer might double the width and height of the input tensor.

[0170] The attention module might comprise at least 2 branches. The output of a first branch and a second branch is multiplied by each other.

[0171] The attention module might comprise 3 branches. The output of a first branch and a second branch is multiplied by each other, whose output is added to the third branch. One of the branches might be identity branch, meaning that no operation is performed on the input by the identity branch.

[0172] 4. Detailed solutions

[0173] The detailed solutions below should be considered as examples to explain general concepts. These solutions should not be interpreted in a narrow way. Furthermore, these solutions can be combined in any manner.

[0174] 4.1 Core of solutions

[0175] The target of the solutions are to improve the reconstruction capability of the synthesis transform module with constraint on computational resources. The core of the present disclosure is to simplify the synthesis transform module while maintaining the  reconstruction capability. The structure of attention module and the position of attention module may be modified. The convolutional layers used by attention module and synthesis transform may be modified to further reduce the model complexity, which can be further applied / extended to other parts of the codec such as analysis transform, entropy coding module, context modeling module, prediction fusion module, hyper-prior coding module, etc.

[0176] 4.2 Details of the solutions

[0177] 1. Depth-wise separable convolution and / or pixel-wise convolution may be included in the synthesis transform.

[0178] a) In one example, all existing convolution layers (or deconvolution layers) are replaced by depth-wise separable convolution.

[0179] i. In one example, the group number of the depth-wise separable convolution layer is identical to the depth of features, i.e., channel number.

[0180] ii. In one example, the group number of the depth-wise separable convolution layer is set as F.

[0181] 1. In one example, F is equal to 1.

[0182] 2. In one example, F may be the power of 2.

[0183] iii. The kernel size of the depth-wise separable convolution may be KxP.

[0184] 1. In one example, K and P equal to 1.

[0185] 2. In one example, K and P equal to 2n+1, where n may be a positive integer number.

[0186] 3. In one example, the values K and P may not be identical.

[0187] b) In one example, the depth-wise separable convolution (or deconvolution) , pointwise convolution (or deconvolution) , and conventional convolution (or deconvolution) may be combined and used in the N-th step of synthesis transformation.

[0188] i. In one example, the depth-wise separable convolution (or deconvolution) , and pointwise convolution (or deconvolution) may be sequentially combined wherein depth-wise separable convolution (or deconvolution) is first used.

[0189] 1. Alternatively, pointwise convolution (or deconvolution) is first used.

[0190] ii. In one example, the depth-wise separable convolution (or deconvolution) , and pointwise convolution (or deconvolution) are used as separated branch.

[0191] iii. In one example, the order of the depth-wise separable convolution (or deconvolu-tion) , pointwise convolution (or deconvolution) , and conventional convolution (or deconvolution) can be arbitrary changed.

[0192] 2. The attention modules may involve the depth-wise separable convolution or deconvolution layers.

[0193] a) In one example, whether the depth-wise separable convolution is used in the attention modules or where the depth-wise separable convolution layers are placed may be de-termined by the available computational resources.

[0194] b) In one example, only upscaling layers are involved in the synthesis transform. The attention modules are excluded from the synthesis transform, catering to the extremely low complexity decoding scenario. Partial or all of upscaling layers are implemented as depth-wise separable convolution layers.

[0195] c) In one example, the attention module is placed at the N-th steps of synthesis transform.

[0196] i. In one example, N equals to 0, corresponding to the location before the first de-convolutional layer (e.g., upsampling layer of upscaling layer) . In this example the input is processed first by the attention module, and then by one or more deconvolution (upsampling or upscaling) layer. Herein, the deconvolution layer may be one or multiple depth-wise separable deconvolution layer and / or point-wise deconvolution layer.

[0197] ii. In one example, N equals to 1, corresponding to the location after the first decon-volutional layer. In this example the input is processed first by a deconvolution (upsampling or upscaling) layer, and then an attention module. Herein, the decon-volution layer may be one or multiple depth-wise separable deconvolution layer and / or point-wise deconvolution layer.

[0198] iii. In one example, N equals to 2, corresponding to the location after the second de-convolutional layer. Herein, the deconvolution layer may be one or multiple depth-wise separable deconvolution layer and / or point-wise deconvolution layer.

[0199] iv. In one example, N equals to 3, corresponding to the location after the third decon-volutional layer. Herein, the deconvolution layer may be one or multiple depth-wise separable deconvolution layer and / or point-wise deconvolution layer.

[0200] v. In one example, N equals to 4, corresponding to the location after the fourth de-convolutional layer. Herein, the deconvolution layer may be one or multiple depth-wise separable deconvolution layer and / or point-wise deconvolution layer.

[0201] vi. Alternatively, for example, the attention module may be placed before or after the cropping layers. Herein, the deconvolution layer may be one or multiple depth-wise separable deconvolution layer and / or point-wise deconvolution layer.

[0202] vii. Alternatively, for example, the attention module may be placed before or after the non-linear activation layers.

[0203] viii. Alternatively, for example, the attention module may be placed before or after the depth-wise separable convolution layers.

[0204] 3. The synthesis transform may contain upscaling layers and / or attention modules.

[0205] a) In on example, whether the attention modules are involved or where the attention mod-ules are placed may be determined by the available computational resources.

[0206] b) In one example, only upscaling layers are involved in the synthesis transform. The attention modules are excluded from the synthesis transform, catering to the extremely low complexity decoding scenario.

[0207] c) In one example, the attention module is placed at the N-th steps of synthesis transform.

[0208] i. In one example, N equals to 0, corresponding to the location before the first de-convolutional layer (e.g. upsampling layer of upscaling layer) . In this example the input is processed first by the attention module, and then by one or more de-convolution (upsampling or upscaling) layer.

[0209] ii. In one example, N equals to 1, corresponding to the location after the first decon-volutional layer. In this example the input is processed first by a deconvolution (upsampling or upscaling) layer, and then an attention module.

[0210] iii. In one example, N equals to 2, corresponding to the location after the second de-convolutional layer.

[0211] iv. In one example, N equals to 3, corresponding to the location after the third decon-volutional layer.

[0212] v. In one example, N equals to 4, corresponding to the location after the fourth de-convolutional layer.

[0213] vi. Alternatively, for example, the attention module may be placed before or after the cropping layers.

[0214] vii. Alternatively, for example, the attention module may be placed before or after the non-linear activation layers.

[0215] 4. The synthesis transform may contain upscaling layers and / or attention modules.

[0216] a) In one example, whether the attention modules are involved or where the attention modules are placed may be signaled.

[0217] b) In one example, one flag is used to signal whether the attention module is involved in the synthesis transform.

[0218] c) In one example, a syntax is used to indicate how many attention modules are involved in the synthesis transform.

[0219] d) In one example, a set of attention modules are provided. The usage of attention mod-ules may be determined according to the expected complexity or performance.

[0220] i. In one example, a flag is signaled indicating the priority of complexity. When low-complexity is much desired, a lightweight attention module may be enabled.

[0221] ii. In one example, a flag is signaled indicating the priority of reconstruction quality. When high-performance is much desired, a sophisticated attention module may be used.

[0222] iii. In another example a flag is included in the bitstream to indicate whether a first attention module is used or a second attention module is used.

[0223] e) In one example, multiple flags may be used to indicate whether the N-th synthesis step (i.e. a deconvolution layer, a cropping layer and a non-linear deconvolutional layer) includes an attention module.

[0224] f) In one example, multiple types of attention modules are provided with different or same complexity. How to combine theses attention modules with a synthesis step may be further signaled.

[0225] 5. A downscaling layer and upscaling layer are involved to save the computational complexity. The multiple depth-wise separable deconvolution layer and / or point-wise deconvolution layer are involved as part of downscaling layer and upscaling layer.

[0226] a) In one example, the downscaling and upscaling are operated in spatial-wise.

[0227] i. In one example, downscaling layer is applied at the beginning of the main trunk and mask trunk, decreasing the spatial ratio of feature maps. After the combination of mask trunk and main trunk, an upscaling layer is applied to recover the spatial resolution of feature maps.

[0228] ii. In one example, downscaling layer is applied at the beginning of the mask trunk, decreasing the spatial ratio of feature maps. Before the ending of mask trunk, an upscaling layer is applied to recover the spatial resolution of feature maps.

[0229] iii. In one example, downscaling layer is applied at the beginning of the main trunk, decreasing the spatial ratio of feature maps. Before the ending of main trunk, an upscaling layer is applied to recover the spatial resolution of feature maps.

[0230] iv. In one example, a combination of depth-wise separable deconvolution and / or point-wise deconvolution layer composes as upscaling and downscaling layers, changing the spatial resolution of feature maps.

[0231] b) Alternatively, for example, the downscaling layer and upscaling layer may operate at channel-wise that adjusts the feature channel numbers.

[0232] 6. The attention module may include a main trunk, a skip trunk and a mask trunk. The three trucks receive the identical input.

[0233] a) Alternatively, for example, the input of the mask trunk may be the intermediate output of the main trunk.

[0234] 7. Residual swin transformer blocks may be used in the synthesis transformation. Depth-wise separable deconvolution layer and / or point-wise deconvolution layer are involved as partial of residual swin transformer block.

[0235] a) In one example, the residual swin transformer block and attention modules may be both used in the synthesis transformation.

[0236] b) In one example, residual swin transformer block is used in the synthesis transformation and attention modules are removed from synthesis transformation.

[0237] c) In one example, the residual swin transformer block may be partial of the attention modules.

[0238] i. In one example, only the swin transformer layer is included in the attention mod-ules.

[0239] ii. In one example, residual swin transformer may act as a new branch and the asso-ciated output may combine with the output of attention module through adding  / subtraction  / concatenation  / fusion  / multiplication  / nonlinear activation.

[0240] d) In one example, the residual swin transformer block is a residual block with swin trans-former layers and depth-wise separable convolutional layers.

[0241] e) In one example, the residual swin transformer block is a residual block with swin trans-former layers with depth-wise separable deconvolutional layer (e.g. upsampling layer of upscaling layer) .

[0242] f) In one example, the residual swin transformer block is a modified residual block where one branch contains swin transformer layers and one branch contains depth-wise sep-arable convolution and / or deconvolutional layers.

[0243] g) In one example, the swin transformer layers may be configured with HEAD H, DEPTH D, WINDOW size W, and PATCH size P.

[0244] i. In one example, W=1, P=1, D=2, H=2.

[0245] ii. In one example, W=1, P=1, D=4, H=4.

[0246] iii. In one example, W=4, P=2.

[0247] iv. In one example, H is the same as D.

[0248] v. In one example, H equals to 2 and D equals to 2.

[0249] vi. In one example, the settings of swin transformer layers may be dependent on the position of residual swin transformer block.

[0250] 1. In one example, two residual swin transformer blocks are placed at the N-th step and M-th step of synthesis transform, respectively. The settings of swin transformer layers may be same or different.

[0251] 2. In one example, multiple residual swin transformer blocks are placed within each step of synthesis transform. The settings of swin transformer layers may be same or different.

[0252] h) In one example, the swin transformer layer is composed with multi-head self-attention layer, multi-layer perceptron, and layer normalization.

[0253] 8. In one example an attention module is used in synthesis transform, wherein the attention module comprises at least two branches.

[0254] a) The first branch comprises first a downscaling layer, followed by at least one convo-lution layer, or one depth-wise separable convolution layer or a residual block, finally an upscaling layer.

[0255] i. Alternatively or additionally the first branch might comprise an activation layer, e.g. a sigmoid layer, a hyperbolic tangent function, a relu layer or such, that is applied after the upscaling layer.

[0256] ii. Alternatively or additionally the first branch might comprise a depth-wise separa-ble convolution layer that is applied after the upscaling layer.

[0257] iii. The upscaling layer might increase the size of the input in height or width or chan-nel dimension) .

[0258] 1. In one example the upscaling layer might increase the width and height of the input by a factor of x2.

[0259] 2. In one example the upscaling layer might increase the number of channels (e.g. number of channel maps) of the input by a factor of x2.

[0260] iv. The downscaling layer might decrease the size of the input in height or width or channel dimension) .

[0261] 1. In one example the downscaling layer might decrease the width and height of the input by a factor of x2.

[0262] 2. In one example the downscaling layer might decrease the number of channels (e.g. number of channel maps) of the input by a factor of x2.

[0263] v. The number of samples of an input might be reduced by an downscaling layer and increased by an upscaling layer.

[0264] vi. Since the first branch begins with a downscaling layer, all operations performed till the upscaling layer are performed using reduced number of input samples, therefore reducing the complexity.

[0265] b) The second branch might comprise a depth-wise separable convolution layer or a re-sidual block.

[0266] c) The second branch might be an identity branch, meaning that the input of the second branch is output without any modification.

[0267] d) The output of the first branch and the second branch might be multiplied by each other.

[0268] i. Alternatively or additionally, the multiplication result (of first and second branches) might be added to the output of a third branch.

[0269] e) In one example the attention module comprises 3 branches.

[0270] i. A first branch comprises a first downscaling layer, 3 residual blocks, an upscaling layer, a convolution layer and an activation layer in that order. The depth-wise separable convolution layer may be used in the downscaling layer, residual blocks, and upscaling layers.

[0271] ii. A second branch comprises 2 residual blocks (RB) . The depth-wise separable con-volution layer may be involved in the residual blocks.

[0272] iii. The output of the first and the second branches are multiplied with each other.

[0273] iv. The multiplication result is added to the input (i.e. the third branch is an identity branch) .

[0274] f) In one example the attention module comprises 3 branches.

[0275] i. A first branch comprises a first downscaling layer, one or more residual blocks, an upscaling layer, a convolution layer and an activation layer in that order. The depth-wise separable convolution layer may be used in the downscaling layer, re-sidual blocks and upscaling layers.

[0276] ii. A second branch comprises one or more residual blocks (RB) .

[0277] iii. The output of the first and the second branches are multiplied with each other.

[0278] iv. The multiplication result is added to the input (i.e. the third branch is an identity branch) .

[0279] g) In one example the attention module comprises 3 branches.

[0280] i. A first branch comprises a first downscaling layer, one or more residual blocks, an upscaling layer, a convolution layer and an activation layer in that order. The depth-wise separable convolution layer may be used in the downscaling layer, re-sidual blocks and upscaling layers.

[0281] ii. A second branch is an identity branch.

[0282] iii. The output of the first and the second branches are multiplied with each other.

[0283] iv. The multiplication result is added to the input (i.e. the third branch is an identity branch) .

[0284] h) In one example the attention module comprises 2 branches.

[0285] i. A first branch comprises a first downscaling layer, one or more residual blocks, an upscaling layer, a convolution layer and an activation layer in that order. The depth-wise separable convolution layer may be used in the downscaling layer, re-sidual blocks and upscaling layers.

[0286] ii. A second branch is an identity branch.

[0287] iii. The output of the first and the second branches are multiplied with each other, which is the output of the attention module.

[0288] i) In one example the attention module comprises at least 2 branches.

[0289] i. A first branch comprises a first downscaling layer, one or more residual blocks, an upscaling layer, a convolution layer and an activation layer in that order. No residual block is applied before the downscaling layer and after the upscaling layer (therefore the computationally complicated residual block operations are per-formed all on reduced number of samples) . The depth-wise separable convolution layer may be used in the downscaling layer, and upscaling layers.

[0290] j) In one example (e.g. Figure 11) the attention module comprises 2 branches.

[0291] i. The first branch is an identity branch.

[0292] ii. The second branch comprises firstly one or more residual blocks (in other words the input of the second branch is first processed with one or more residual blocks) , after which the second branch splits into 2 branches (branches 3 and 4) :

[0293] 1. The third branch is an identity branch.

[0294] 2. The fourth branch comprises one or more residual blocks and it might com-prise an activation layer (e.g. a sigmoid layer or a hyperbolic tangent function or such) .

[0295] 3. The output of the third and fourth branches are multiplied with each other to form the output of the second branch.

[0296] General aspects

[0297] 9. Whether to and / or how to apply the disclosed methods above may be signalled at block level / sequence level / group of pictures level / picture level / slice level / tile group level, such as in coding structures of CTU / CU / TU / PU / CTB / CB / TB / PB, or sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / tile group header.

[0298] 10. Whether to and / or how to apply the disclosed methods above may be dependent on coded information, such as block size, colour format, single / dual tree partitioning, colour compo-nent, slice / picture type.

[0299] 11. The proposed methods disclosed in this document may be used in other coding tools which require chroma fusion.

[0300] 12. A syntax element disclosed above may be binarized as a flag, a fixed length code, an EG (x) code, a unary code, a truncated unary code, a truncated binary code, etc. It can be signed or unsigned.

[0301] 13. A syntax element disclosed above may be coded with at least one context model. Or it may be bypass coded.

[0302] 14. A syntax element disclosed above may be signaled in a conditional way.

[0303] a. The SE is signaled only if the corresponding function is applicable.

[0304] b. The SE is signaled only if the dimensions (width and / or height) of the block sat-isfy a condition.

[0305] 15. A syntax element disclosed above may be signaled at block level / sequence level / group of pictures level / picture level / slice level / tile group level, such as in coding structures of CTU / CU / TU / PU / CTB / CB / TB / PB, or sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / tile group header.

[0306] 4.3 Benefit of the present disclosure

[0307] According to the present disclosure, the synthesis transform subnetwork is modified. The present disclosure involves a simplified attention module by earlier downscaling the feature maps and then upscaling the feature maps, depth-wise convolution layers and point-wise convolution layers, such that the computational complexity can be decreased. Moreover, the position of the simplified attention module and depth-wise convolution layers can also be adjusted according to the budget of computational resources.

[0308] 5. Embodiments

[0309] A residual calibrated attention module is used in synthesis transform, where the structure of the residual calibrated attention module is illustrated in FIG. 14. FIG. 14 shows a residual calibrated module.

[0310] A simplified attention module is used in the synthesis transform, where the structure is shown in FIG. 15. FIG. 15 shows a simplified attention module.

[0311] Residual swin transformer blocks are used in the synthesis transform, where the structure of the swin transformer layer is shown in FIG. 16. FIG. 16 shows a swin transformer layer.

[0312] A simplified attention module is used in synthesis transform, where the structure of the residual attention module is illustrated in FIG. 17. FIG. 17 shows a simplified attention module.

[0313] A simplified attention module with depth-wise convolution layers for up-scaling and down-scaling is used in the synthesis transform, where the structure is shown in FIG. 18.FIG. 18 shows a simplified attention module with depth-wise separated convolution layers for down-scaling and up-scaling.

[0314] A simplified attention module with depth-wise convolution layer and point-wise convolution layer is used in the residual blocks in the synthesis transform, where the structure is shown in FIG. 19. FIG. 19 shows a simplified attention module with depth-wise separated convolution layer and point-wise convolution layer in the residual blocks.

[0315] As used herein, the term “video unit” or “video block” may be a sequence, a picture, a slice, a tile, a brick, a subpicture, a coding tree unit (CTU)  / coding tree block (CTB) , a CTU / CTB row, one or multiple coding units (CUs)  / coding blocks (CBs) , one ore multiple CTUs / CTBs, one or multiple Virtual Pipeline Data Unit (VPDU) , a sub-region within a picture / slice / tile / brick.

[0316] Fig. 20 illustrates a flowchart of a method 2000 for video processing in accordance with embodiments of the present disclosure. The method 2000 is implemented during a conversion between a video unit of a video and a bitstream of the video.

[0317] At block 2010, for a conversion between a video unit of a video and a bitstream of the video, it is determined to apply a neural network based image compression network to the video unit. The neural network based image compression network includes a synthesis transform module that includes at least one of: one or more upscaling layers and one or more attention modules.

[0318] At block 2020, the conversion is performed based on the neural network based image compression network. In some embodiments, the conversion may include encoding the video unit into the bitstream. Alternatively, or in addition, the conversion may include decoding the video unit from the bitstream. According to embodiments of the present disclosure, the synthesis transform subnetwork is modified. In particular, it involves a simplified attention module by earlier downscaling the feature maps and then upscaling the feature maps, depth-wise convolution layers and point-wise convolution layers, such  that the computational complexity can be decreased. Moreover, the position of the simplified attention module and depth-wise convolution layers can also be adjusted according to the budget of computational resources.

[0319] In some embodiments, the synthesis transform module only comprises the one or more upscaling layers, and the one or more attention modules are excluded from the synthesis transform module. For example, the one or more attention modules are excluded from the synthesis transform, catering to the extremely low complexity decoding scenario.

[0320] In some embodiments, the one or more attention modules are placed at N-th steps of the synthesis transform module, where N is an integer number. For example, N is equal to 2, and the one or more attention modules are placed a location after a second deconvolutional layer. As another example, N is equal to 3, and the one or more attention modules are placed a location after a third deconvolutional layer. As a further example, N is equal to 4, and the one or more attention modules are placed a location after a fourth deconvolutional layer.

[0321] In some embodiments, N is equal to 0, and the one or more attention modules are placed a location before a first deconvolutional layer (for example, upsampling layer of upscaling layer) . In this case, an input may be processed by the one or more attention modules, and then by one or more deconvolution (for example, upsampling or upscaling) layers.

[0322] In some embodiments, N is equal to 1, and the one or more attention modules are placed a location after the first deconvolutional layer. In this case, an input is processed by a deconvolution layer, and then by the one or more attention modules. In other words, the input is processed first by a deconvolution (upsampling or upscaling) layer, and then an attention module.

[0323] In some embodiments, the one or more attention modules are placed before cropping layers. Alternatively, the one or more attention modules are placed after the cropping layers.

[0324] In some embodiments, the one or more attention modules are placed before non-linear activation layers. Alternatively, the one or more attention modules are placed after the non-linear activation layers.

[0325] In some embodiments, whether the one or more modules are excluded from the  synthesis transform module based on currently available computational resources. In some other embodiments, where the one or more modules are placed in the synthesis transform module based on currently available computational resources.

[0326] In some embodiments, a plurality of types of attention modules is provided with different complexity. Alternatively, the plurality of types of attention modules is provided with same complexity. In some embodiments, a way to combine the plurality of types of attention modules with a synthesis step may be indicated.

[0327] In some embodiments, whether the one or more modules are involved in the synthesis transform module is indicated. Alternatively, or in addition, where the one or more modules are placed the synthesis transform module is indicated.

[0328] In some embodiments, a flag is used to indicate whether the one or more modules are involved in the synthesis transform module. In some other embodiments, a syntax element is used to indicate the number of attention modules involved in the synthesis transform module.

[0329] In some embodiments, a set of attention modules are provided. In this case, a usage of the set of attention modules may be determined according to an expected complexity or performance.

[0330] In some embodiments, a flag is signaled to indicate a priority of complexity. In this case, if a low-complexity is desired, a lightweight attention module may be enabled.

[0331] Alternatively, a flag is signaled to indicate a priority of reconstruction quality. In this case, if a high-performance is desired, a sophisticated attention module may be used. In some other embodiments, a flag is included in the bitstream to indicate whether a first attention module is used or a second attention module is used.

[0332] In some embodiments, a plurality of flags is used to indicate whether the N-th synthesis step includes an attention module, where N is an integer number. For example, the N-th synthesis step may include a deconvolution layer, a cropping layer and a non-linear deconvolutional layer.

[0333] In some embodiments, a downscaling layer and an upscaling layer are involved in the synthesis transform module. In some embodiments, at least one of: a depth-wise separable deconvolution layer or a point-wise deconvolution layer is involved as part of the downscaling layer and upscaling layer.

[0334] In some embodiments, the downscaling layer and the upscaling layer are operated in channel-wise that adjusts feature channel numbers. Alternatively, the downscaling layer and the upscaling layer are operated in spatial-wise.

[0335] By way of example, the downscaling layer is applied at a beginning of a mask trunk, and the upscaling layer is applied before an ending of mask trunk. For example, the downscaling layer is applied at a beginning of a main trunk and a mask trunk, and the upscaling layer is applied after a combination of mask trunk and main trunk. As another example, the downscaling layer is applied at a beginning of a main trunk, and the upscaling layer is applied before an ending of the main trunk. As a further example, a combination of the depth-wise separable deconvolution layer the point-wise deconvolution layer acts as the upscaling and downscaling layers and changes a spatial resolution of feature maps. In some embodiments, the downscaling layer decreases a spatial ratio of feature maps, and the upscaling layer recovers a spatial resolution of feature maps.

[0336] In some embodiments, an attention module of the one or more attention modules comprises a main trunk, a skip trunk and a mask trunk. In some embodiments, the main trunk, the skip trunk and the mask trunk receive an identical input. Alternatively, an input of the mask trunk is an intermediate output of the main trunk.

[0337] In some embodiments, an attention module used in the synthesis transform module comprises at least two branches. In some embodiments, a first branch in the attention module comprises: a downscaling layer, at least one convolution layer or a residual block, and an upscaling layer. In some embodiments, the first branch comprises a depth-wise separable convolution layer that is applied after the upscaling layer. Alternatively, or in addition, the first branch comprises an activation layer that is applied after the upscaling layer. The activation layer may include one or more of: a sigmoid layer, a hyperbolic tangent function, or a relu layer.

[0338] In some embodiments, the upscaling layers increase a size or channel dimension of an input. In some embodiments, the upscaling layer increases a width and height of the input by a factor of 2. In some other embodiments, the upscaling layer increases the number of channels (for example, the number of channel maps) of the input by a factor of 2.

[0339] In some embodiments, the downscaling layer decreases a size or channel  dimension of an input. In some embodiments, the downscaling layer decreases a width and height of the input by a factor of 2. In some other embodiments, the downscaling layer decreases the number of channels (for example, the number of channel maps) of the input by a factor of 2.

[0340] In some embodiments, the number of samples of an input is reduced by the downscaling layer and increased by the upscaling layer. In some other embodiments, all operations performed till the upscaling layer are performed using reduced number of input samples. In this case, since the first branch begins with a downscaling layer, it can reduce the complexity.

[0341] In some embodiments, a second branch in the attention module comprises: a depth-wise separable convolution layer or a residual block. In some other embodiments, a second branch in the attention module comprises an identity branch, and an input of the second branch is output without modification. In this case, it means that an output of the first branch and an output of the second branch are multiplied by each other.

[0342] In some embodiments, outputs of the first and second branches may be multiplied by each other. In some other embodiments, a multiplication result of the first and second branches is added to an output of a third branch in the attention module.

[0343] In some embodiments, the attention module comprises three branches. For example, a first branch in the three branches comprises: a first downscaling layer, one or more residual blocks, an upscaling layer, a convolution layer and an activation layer in that order. In some embodiments, a depth-wise separable convolution layer is used in at least one of: the first downscaling layer, the one or more residual blocks and the upscaling layer. Alternately, or in addition, a second branch in the three branches comprises one or more residual blocks (RB) .

[0344] In some embodiments, a first branch in the three branches comprises: a first downscaling layer, three residual blocks, an upscaling layer, a convolution layer and an activation layer in that order. In some embodiments, a depth-wise separable convolution layer is used in at least one of: the first downscaling layer, the three residual blocks and the upscaling layer. In some embodiments, a second branch in the three branches comprises two residual blocks (RBs) , and a depth-wise separable convolution layer is involved in the two RBs. In some other embodiments, a second branch in the three branches comprises an identity branch. In some embodiments, outputs of first and second  branches in the three branches are multiplied with each other. In some other embodiments, a multiplication result of the first and second branches is added to an input of a third branch which is an identity branch.

[0345] In some embodiments, the attention module comprises two branches. In some embodiments, a first branch in the two branches comprises: a first downscaling layer, one or more residual blocks, an upscaling layer, a convolution layer and an activation layer in that order. In some embodiments, a depth-wise separable convolution layer is used in at least one of: the first downscaling layer, the one or more residual blocks and the upscaling layer. In some other embodiments, a depth-wise separable convolution layer is used in at least one of: the first downscaling layer and the upscaling layer. In some embodiments, no residual block is applied before the first downscaling layer and after the upscaling layer.

[0346] In some embodiments, a second branch in the two branches comprises an identity branch. In some embodiments, outputs of first and second branches in the two branches are multiplied with each other, which is an output of the attention module.

[0347] In some embodiments, as shown in FIG. 14, a first branch in the two branches is an identity branch. In some embodiments, a second branch in the two branches comprises one or more residual blocks, and after one or more residual blocks, the second branch is split into two branches comprising a third branch and a fourth branch. In some embodiments, the third branch is an identity branch. In some embodiments, the fourth branch comprises one or more other residual blocks. The fourth branch may further include an activation layer. For example, the fourth branch may include a sigmoid layer or a hyperbolic tangent function. In some embodiments, outputs of the third and fourth branches are multiplied with each other to form an output of the second branch.

[0348] In some embodiments, a residual swin transformer block is used in the synthesis transform module. For example, at least one of: a depth-wise separable deconvolution layer or a point-wise deconvolution layer are involved as partial of the residual swin transformer block.

[0349] In some embodiments, the residual swin transformer block and the one or more attention modules are used in the synthesis transform module. In some other embodiments, the residual swin transformer block is used in the synthesis transform module and the one or more attention modules are removed from the synthesis transform module.

[0350] In some embodiments, the residual swin transformer block is partial of the one or more attention modules. For example, only a swin transformer layer is included in the one or more attention modules. As another example, the residual swin transformer block acts as a new branch and an associated output is combine with an output of the one or more attention modules through one of: adding, subtraction, concatenation, fusion, multiplication, or nonlinear activation.

[0351] In some embodiments, the residual swin transformer block is a residual block with a swin transformer layer and a depth-wise separable convolutional layer. In some other embodiments, the residual swin transformer block is a residual block with a swin transformer layer with a depth-wise separable deconvolutional layer (for example, upsampling layer of upscaling layer) . In some other embodiments, the residual swin transformer block is a modified residual block where one branch comprises a swin transformer layer and one branch comprises at least one of: a depth-wise separable convolution layer or a deconvolutional layer.

[0352] In some embodiments, a swin transformer layer is configured with a head size, a depth size, a window size and a patch size. In some embodiments, the window size is equal to 1, the path size is equal to 1, the depth size is equal to 2, and the head size is equal to 2. Alternatively, the window size is equal to 1, the path size is equal to 1, the depth size is equal to 4, and the head size is equal to 4. In some other embodiments, the window size is equal to 4 and the patch size is equal to 2..

[0353] In some embodiments, the head size is the same as the depth size. Alternatively, the head size is equal to 2 and the depth size is equal to 2.

[0354] In some embodiments, a setting of the swin transformer layer is dependent on a position of residual swin transformer block. For example, one residual swin transformer block is placed at the N-th step of the synthesis transform module, another residual swin transformer block is placed at the M-th step of the synthesis transform module, where N and M are integer numbers.

[0355] In some embodiments, a setting of the residual swin transformer block and a setting of the other residual swin transformer block are same. Alternatively, the setting of the residual swin transformer block and the setting of the other residual swin transformer block are different.

[0356] In some embodiments, a plurality of residual swin transformer blocks is placed within each step of the synthesis transform module. In some embodiments, settings of swin transformer layers are same. Alternatively, the settings of swin transformer layers are different. In some embodiments, a swin transformer layer comprises a multi-head self-attention layer, a multi-layer perceptron, and a layer normalization.

[0357] In some embodiments, at least one of: a depth-wise separable convolution layer or a pixel-wise convolution layer is included in the synthesis transform module. For example, all existing convolution layers or deconvolution layers are replaced by the depth-wise separable convolution layer.

[0358] In some embodiments, the group number of the depth-wise separable convolution layer is identical to a depth of features. For example, the depth pf features may be channel number.

[0359] In some embodiments, the group number of the depth-wise separable convolution layer is set as a predetermined number. For example, the predetermined number is equal to 1. As another example, the predetermined number is the power of 2.

[0360] In some embodiments, a kernel size of the depth-wise separable convolution layer be KxP, where K and P are integer numbers. For example, both K and P are equal to 1. As an example, both K and P are equal to 2n+1, where n is a positive integer number. As another example, values of K and P are not identical.

[0361] In some embodiments, the depth-wise separable convolution layer, a pointwise convolution layer and a conventional convolution layer are combined and used in the N-th step of the synthesis transform module. Alternatively, or in addition, a depth-wise separable deconvolution layer, a pointwise deconvolution layer and a conventional deconvolution layer are combined and used in the N-th step of the synthesis transform module. In this case, N may be an integer number.

[0362] In some embodiments, a depth-wise separable convolution layer and a pointwise convolution layer are sequentially combined, and the depth-wise separable convolution layer is first used. Alternatively, or in addition, the depth-wise separable deconvolution layer and the pointwise deconvolution layer are sequentially combined, and the depth-wise separable convolution layer is first used.

[0363] In some embodiments, a depth-wise separable convolution layer and a pointwise  convolution layer are sequentially combined, and the pointwise convolution layer is first used. Alternatively, or in addition, the depth-wise separable deconvolution layer and the pointwise deconvolution layer are sequentially combined, and the pointwise convolution layer is first used.

[0364] In some embodiments, a depth-wise separable convolution layer and a pointwise convolution layer are used as separated branches. Alternatively, or in addition, the depth-wise separable deconvolution layer and the pointwise deconvolution layer are used as separated branches.

[0365] In some embodiments, an order of the depth-wise separable convolution layer, a pointwise convolution layer and a conventional convolution layer are arbitrary changed. Alternately, or in addition, an order of a depth-wise separable deconvolution layer, a pointwise deconvolution layer and a conventional deconvolution layer are arbitrary changed.

[0366] In some embodiments, the one or more attention modules comprise at least one of a depth-wise separable convolution layer or a depth-wise separable deconvolution layer. In some embodiments, whether the depth-wise separable convolution layer is used in the one or more attention modules is determined based on available computational resources. Alternatively, or in addition, where the depth-wise separable convolution layer is placed is determined based on the available computational resources.

[0367] In some embodiments, only upscaling layers are involved in the synthesis transform module and the one or more attention modules are excluded from the synthesis transform module. In this case, partial or all of the upscaling layers may be implemented as depth-wise separable convolution layers.

[0368] In some embodiments, the one or more attention modules are placed at the N-th steps of the synthesis transform module, wherein N is an integer number. For example, N is equal to 0, the one or more attention modules are placed at a location before a first deconvolutional layer. In some embodiments, an input is processed by the one or more attention modules, and then by one or more deconvolution layers.

[0369] In some embodiments, N is equal to 1, and the one or more attention modules are placed a location after the first deconvolutional layer. In some embodiments, an input is processed by a deconvolution layer, and then by the one or more attention modules. In  some embodiments, N is equal to 2, and the one or more attention modules are placed a location after a second deconvolutional layer. In some other embodiments, N is equal to 3, and the one or more attention modules are placed a location after a third deconvolutional layer. In some further embodiments, N is equal to 4, and the one or more attention modules are placed a location after a fourth deconvolutional layer. In some embodiments, a deconvolution layer comprises one or more depth-wise separable deconvolution layers or one or more point-wise deconvolution layers.

[0370] In some embodiments, the one or more attention modules are placed before cropping layers. Alternatively, the one or more attention modules are placed after the cropping layers.

[0371] In some embodiments, the one or more attention modules are placed before non-linear activation layers. Alternatively, the one or more attention modules are placed after the non-linear activation layers.

[0372] In some embodiments, the one or more attention modules are placed before depth-wise separable convolution layers. Alternatively, the one or more attention modules are placed after the depth-wise separable convolution layers.

[0373] In some embodiments, an indication of whether to and / or how to determine to apply the neural network based image compression network to the video unit is indicated at one of the followings: sequence level, group of pictures level, picture level, slice level, or tile group level. In some embodiments, an indication of whether to and / or how to determine to apply the neural network based image compression network to the video unit is indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS) , a video parameter set (VPS) , a dependency parameter set (DPS) , a decoding capability information (DCI) , a picture parameter set (PPS) , an adaptation parameter sets (APS) , a slice header, or a tile group header. In some embodiments, an indication of whether to and / or how to determine to apply the neural network based image compression network to the video unit is included in one of the following: a prediction block (PB) , a transform block (TB) , a coding block (CB) , a prediction unit (PU) , a transform unit (TU) , a coding unit (CU) , a coding tree block (CTB) , or a coding tree unit (CTU) .

[0374] In some embodiments, the method 2000 further comprises: determining, based on coded information of the video unit, whether and / or how to determine to apply the  neural network based image compression network to the video unit. The coded information may include at least one of: a block size, a colour format, a single and / or dual tree partitioning, a colour component, a slice type, or a picture type. In some embodiments, the video unit is applied with a coding tool that requires chroma fusion.

[0375] In some embodiments, the SE is binarized as one of a flag, a fixed length code, an EG (x) code, a unary code, a truncated unary code, or a truncated binary code. In some embodiments, the SE is signed or unsigned. In some embodiments, the SE is coded with at least one context model. Alternatievly, the SE is bypass coded. In some embodiments, the SE is signaled in a conditional way. In some embodiments, the SE is signaled only if a corresponding function is applicable. Alternatively, the SE is signaled only if dimensions of the video unit satisfy a condition.

[0376] In some embodiments, the SE is indicated at one of the followings: sequence level, group of pictures level, picture level, slice level, or tile group level. In some embodiments, the SE is indicated at one of the followings: a prediction block (PB) , a transform block (TB) , a coding block (CB) , a prediction unit (PU) , a transform unit (TU) , a coding unit (CU) , a coding tree block (CTB) , or a coding tree unit (CTU) .

[0377] According to further embodiments of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video which is generated by a method performed by an apparatus for video processing. The method comprises: determining to apply a neural network based image compression network to a video unit of the video, wherein the neural network based image compression network comprises a synthesis transform module that comprises at least one of: one or more upscaling layers and one or more attention modules; and generating the bitstream based on the neural network based image compression network.

[0378] According to still further embodiments of the present disclosure, a method for storing bitstream of a video is provided. The method comprises: determining to apply a neural network based image compression network to a video unit of the video, wherein the neural network based image compression network comprises a synthesis transform module that comprises at least one of: one or more upscaling layers and one or more attention modules; generating the bitstream based on the neural network based image compression network; and storing the bitstream in a non-transitory computer-readable  medium.

[0379] FIG. 21 shows example of convolution-based attention block. A convolution-based attention block receives a tensor input of size [C, h, w] and outputs tensor output of same size after performing the sequence of steps depicted in FIG. 21. In particular, the process may be split into two branches, operating in different spatial resolution. The first branch may include two residual blocks. The second branch may start with down-sampling convolution with stride = 2, followed by two residual blocks and transposed convolution with stride = 2. The second branch may be concluded by sigmoid. Two branches may be joined in per-element multiplication ⊙. Then resulting tensor may be multiplied by parameter α, and added to the input tensor. With α=0 all operations in CAB are essentially by-passed.

[0380] Implementations of the present disclosure can be described in view of the following clauses, the features of which can be combined in any reasonable manner.

[0381] Clause 1. A method for video processing, comprising: determining, for a conversion between a video unit of a video and a bitstream of the video, to apply a neural network based image compression network to the video unit, wherein the neural network based image compression network comprises a synthesis transform module that comprises at least one of: one or more upscaling layers and one or more attention modules; and performing the conversion based on the neural network based image compression network.

[0382] Clause 2. The method of clause 1, wherein the synthesis transform module only comprises the one or more upscaling layers, and the one or more attention modules are excluded from the synthesis transform module.

[0383] Clause 3. The method of clause 1, wherein the one or more attention modules are placed at N-th steps of the synthesis transform module, wherein N is an integer number.

[0384] Clause 4. The method of clause 3, wherein N is equal to 2, and the one or more attention modules are placed a location after a second deconvolutional layer.

[0385] Clause 5. The method of clause 3, wherein N is equal to 3, and the one or more attention modules are placed a location after a third deconvolutional layer.

[0386] Clause 6. The method of clause 3, wherein N is equal to 4, and the one or more attention modules are placed a location after a fourth deconvolutional layer.

[0387] Clause 7. The method of clause 3, wherein N is equal to 0, and the one or more attention modules are placed a location before a first deconvolutional layer.

[0388] Clause 8. The method of clause 7, wherein an input is processed by the one or more attention modules, and then by one or more deconvolution layers.

[0389] Clause 9. The method of clause 3, wherein N is equal to 1, and the one or more attention modules are placed a location after the first deconvolutional layer.

[0390] Clause 10. The method of clause 9, wherein an input is processed by a deconvolution layer, and then by the one or more attention modules.

[0391] Clause 11. The method of clause 3, wherein the one or more attention modules are placed before cropping layers, or wherein the one or more attention modules are placed after the cropping layers.

[0392] Clause 12. The method of clause 3, wherein the one or more attention modules are placed before non-linear activation layers, or wherein the one or more attention modules are placed after the non-linear activation layers.

[0393] Clause 13. The method of clause 1, wherein whether the one or more modules are excluded from the synthesis transform module based on currently available computational resources.

[0394] Clause 14. The method of clause 1, wherein where the one or more modules are placed in the synthesis transform module based on currently available computational resources.

[0395] Clause 15. The method of clause 1, wherein a plurality of types of attention modules is provided with different complexity, or wherein the plurality of types of attention modules is provided with same complexity.

[0396] Clause 16. The method of clause 15, wherein a way to combine the plurality of types of attention modules with a synthesis step is indicated.

[0397] Clause 17. The method of clause 1, wherein whether the one or more modules are involved in the synthesis transform module is indicated, and / or wherein where the one or more modules are placed the synthesis transform module is indicated.

[0398] Clause 18. The method of clause 1, wherein a flag is used to indicate whether the one or more modules are involved in the synthesis transform module.

[0399] Clause 19. The method of clause 1, wherein a syntax element is used to indicate the number of attention modules involved in the synthesis transform module.

[0400] Clause 20. The method of clause 1, wherein a set of attention modules are provided, and a usage of the set of attention modules is determined according to an expected complexity or performance.

[0401] Clause 21. The method of clause 20, wherein a flag is signaled to indicate a priority of complexity, and wherein if a low-complexity is desired, a lightweight attention module is enabled.

[0402] Clause 22. The method of clause 20, wherein a flag is signaled to indicate a priority of reconstruction quality, and wherein if a high-performance is desired, a sophisticated attention module is used.

[0403] Clause 23. The method of clause 20, wherein a flag is included in the bitstream to indicate whether a first attention module is used or a second attention module is used.

[0404] Clause 24. The method of clause 1, wherein a plurality of flags is used to indicate whether the N-th synthesis step includes an attention module, wherein N is an integer number.

[0405] Clause 25. The method of clause 1, wherein a downscaling layer and an upscaling layer are involved in the synthesis transform module.

[0406] Clause 26. The method of clause 25, wherein at least one of: a depth-wise separable deconvolution layer or a point-wise deconvolution layer is involved as part of the downscaling layer and upscaling layer.

[0407] Clause 27. The method of clause 25 or 26, wherein the downscaling layer and the upscaling layer are operated in channel-wise that adjusts feature channel numbers.

[0408] Clause 28. The method of clause 25 or 26, wherein the downscaling layer and the upscaling layer are operated in spatial-wise.

[0409] Clause 29. The method of clause 28, wherein the downscaling layer is applied at a beginning of a mask trunk, and the upscaling layer is applied before an ending of mask trunk.

[0410] Clause 30. The method of clause 28, wherein the downscaling layer is applied at a beginning of a main trunk and a mask trunk, and the upscaling layer is applied after  a combination of mask trunk and main trunk.

[0411] Clause 31. The method of clause 28, wherein the downscaling layer is applied at a beginning of a main trunk, and the upscaling layer is applied before an ending of the main trunk.

[0412] Clause 32. The method of any of clauses 29 to 31, wherein the downscaling layer decreases a spatial ratio of feature maps, and the upscaling layer recovers a spatial resolution of feature maps.

[0413] Clause 33. The method of clause 28, wherein a combination of the depth-wise separable deconvolution layer the point-wise deconvolution layer acts as the upscaling and downscaling layers and changes a spatial resolution of feature maps.

[0414] Clause 34. The method of clause 1, wherein an attention module of the one or more attention modules comprises a main trunk, a skip trunk and a mask trunk.

[0415] Clause 35. The method of clause 34, wherein the main trunk, the skip trunk and the mask trunk receive an identical input.

[0416] Clause 36. The method of clause 34, wherein an input of the mask trunk is an intermediate output of the main trunk.

[0417] Clause 37. The method of clause 1, wherein an attention module used in the synthesis transform module comprises at least two branches.

[0418] Clause 38. The method of clause 37, wherein a first branch in the attention module comprises: a downscaling layer, at least one convolution layer or a residual block, and an upscaling layer.

[0419] Clause 39. The method of clause 38, wherein the first branch comprises an activation layer that is applied after the upscaling layer.

[0420] Clause 40. The method of clause 39, wherein the activation layer comprises at least one of: a sigmoid layer, a hyperbolic tangent function, or a relu layer.

[0421] Clause 41. The method of clause 38, wherein the upscaling layers increase a size or channel dimension of an input.

[0422] Clause 42. The method of clause 41, wherein the upscaling layer increases a width and height of the input by a factor of 2.

[0423] Clause 43. The method of clause 41, wherein the upscaling layer increases the number of channels of the input by a factor of 2.

[0424] Clause 44. The method of clause 38, wherein the downscaling layer decreases a size or channel dimension of an input.

[0425] Clause 45. The method of clause 44, wherein the downscaling layer decreases a width and height of the input by a factor of 2.

[0426] Clause 46. The method of clause 44, wherein the downscaling layer decreases the number of channels of the input by a factor of 2.

[0427] Clause 47. The method of clause 38, wherein the number of samples of an input is reduced by the downscaling layer and increased by the upscaling layer.

[0428] Clause 48. The method of clause 38, wherein all operations performed till the upscaling layer are performed using reduced number of input samples.

[0429] Clause 49. The method of clause 38, wherein the first branch comprises a depth-wise separable convolution layer that is applied after the upscaling layer.

[0430] Clause 50. The method of clause 37, wherein a second branch in the attention module comprises: a depth-wise separable convolution layer or a residual block.

[0431] Clause 51. The method of clause 37, wherein a second branch in the attention module comprises an identity branch, and an input of the second branch is output without modification.

[0432] Clause 52. The method of clause 37, wherein an output of the first branch and an output of the second branch are multiplied by each other.

[0433] Clause 53. The method of clause 52, wherein a multiplication result of the first and second branches is added to an output of a third branch in the attention module.

[0434] Clause 54. The method of clause 37, wherein the attention module comprises three branches.

[0435] Clause 55. The method of clause 54, wherein a first branch in the three branches comprises: a first downscaling layer, one or more residual blocks, an upscaling layer, a convolution layer and an activation layer in that order.

[0436] Clause 56. The method of clause 55, wherein a depth-wise separable convolution  layer is used in at least one of: the first downscaling layer, the one or more residual blocks and the upscaling layer.

[0437] Clause 57. The method of clause 54, wherein a second branch in the three branches comprises one or more residual blocks (RB) .

[0438] Clause 58. The method of clause 54, wherein a first branch in the three branches comprises: a first downscaling layer, three residual blocks, an upscaling layer, a convolution layer and an activation layer in that order.

[0439] Clause 59. The method of clause 58, wherein a depth-wise separable convolution layer is used in at least one of: the first downscaling layer, the three residual blocks and the upscaling layer.

[0440] Clause 60. The method of clause 54, wherein a second branch in the three branches comprises two residual blocks (RBs) , and a depth-wise separable convolution layer is involved in the two RBs.

[0441] Clause 61. The method of clause 54, wherein a second branch in the three branches comprises an identity branch.

[0442] Clause 62. The method of any of clauses 54-61, wherein outputs of first and second branches in the three branches are multiplied with each other.

[0443] Clause 63. The method of clause 62, wherein a multiplication result of the first and second branches is added to an input of a third branch which is an identity branch.

[0444] Clause 64. The method of clause 37, wherein the attention module comprises two branches.

[0445] Clause 65. The method of clause 37 or 64, wherein a first branch in the two branches comprises: a first downscaling layer, one or more residual blocks, an upscaling layer, a convolution layer and an activation layer in that order.

[0446] Clause 66. The method of clause 65, wherein a depth-wise separable convolution layer is used in at least one of: the first downscaling layer, the one or more residual blocks and the upscaling layer.

[0447] Clause 67. The method of clause 65, wherein a depth-wise separable convolution layer is used in at least one of: the first downscaling layer and the upscaling layer.

[0448] Clause 68. The method of clause 65, wherein no residual block is applied before the first downscaling layer and after the upscaling layer.

[0449] Clause 69. The method of clause 65, wherein a second branch in the two branches comprises an identity branch.

[0450] Clause 70. The method of clause 65, wherein outputs of first and second branches in the two branches are multiplied with each other, which is an output of the attention module.

[0451] Clause 71. The method of clause 64, wherein a first branch in the two branches is an identity branch.

[0452] Clause 72. The method of clause 64, wherein a second branch in the two branches comprises one or more residual blocks, and after one or more residual blocks, the second branch is split into two branches comprising a third branch and a fourth branch.

[0453] Clause 73. The method of clause 72, wherein the third branch is an identity branch.

[0454] Clause 74. The method of clause 72, wherein the fourth branch comprises one or more other residual blocks.

[0455] Clause 75. The method of clause 74, wherein the fourth branch further comprises an activation layer.

[0456] Clause 76. The method of clause 72, wherein outputs of the third and fourth branches are multiplied with each other to form an output of the second branch.

[0457] Clause 77. The method of clause 1, wherein a residual swin transformer block is used in the synthesis transform module.

[0458] Clause 78. The method of clause 77, wherein at least one of: a depth-wise separable deconvolution layer or a point-wise deconvolution layer are involved as partial of the residual swin transformer block.

[0459] Clause 79. The method of clause 77, wherein the residual swin transformer block and the one or more attention modules are used in the synthesis transform module.

[0460] Clause 80. The method of clause 77, wherein the residual swin transformer block is used in the synthesis transform module and the one or more attention modules are  removed from the synthesis transform module.

[0461] Clause 81. The method of clause 77, wherein the residual swin transformer block is partial of the one or more attention modules.

[0462] Clause 82. The method of clause 81, wherein only a swin transformer layer is included in the one or more attention modules.

[0463] Clause 83. The method of clause 81, wherein the residual swin transformer block acts as a new branch and an associated output is combine with an output of the one or more attention modules through one of: adding, subtraction, concatenation, fusion, multiplication, or nonlinear activation.

[0464] Clause 84. The method of clause 77, wherein the residual swin transformer block is a residual block with a swin transformer layer and a depth-wise separable convolutional layer.

[0465] Clause 85. The method of clause 77, wherein the residual swin transformer block is a residual block with a swin transformer layer with a depth-wise separable deconvolutional layer.

[0466] Clause 86. The method of clause 77, wherein the residual swin transformer block is a modified residual block where one branch comprises a swin transformer layer and one branch comprises at least one of: a depth-wise separable convolution layer or a deconvolutional layer.

[0467] Clause 87. The method of clause 77, wherein a swin transformer layer is configured with a head size, a depth size, a window size and a patch size.

[0468] Clause 88. The method of clause 87, wherein the window size is equal to 1, the path size is equal to 1, the depth size is equal to 2, and the head size is equal to 2, or wherein the window size is equal to 1, the path size is equal to 1, the depth size is equal to 4, and the head size is equal to 4, or wherein the window size is equal to 4 and the patch size is equal to 2.

[0469] Clause 89. The method of clause 87, wherein the head size is the same as the depth size, or wherein the head size is equal to 2 and the depth size is equal to 2.

[0470] Clause 90. The method of clause 87, wherein a setting of the swin transformer layer is dependent on a position of residual swin transformer block.

[0471] Clause 91. The method of clause 90, wherein one residual swin transformer block is placed at the N-th step of the synthesis transform module, another residual swin transformer block is placed at the M-th step of the synthesis transform module, wherein N and M are integer numbers.

[0472] Clause 92. The method of clause 91, wherein a setting of the residual swin transformer block and a setting of the other residual swin transformer block are same, or wherein the setting of the residual swin transformer block and the setting of the other residual swin transformer block are different.

[0473] Clause 93. The method of clause 90, wherein a plurality of residual swin transformer blocks are placed within each step of the synthesis transform module.

[0474] Clause 94. The method of clause 93, wherein settings of swin transformer layers are same, or wherein the settings of swin transformer layers are different.

[0475] Clause 95. The method of clause 77, wherein a swin transformer layer comprises a multi-head self-attention layer, a multi-layer perceptron, and a layer normalization.

[0476] Clause 96. The method of clause 1, wherein at least one of: a depth-wise separable convolution layer or a pixel-wise convolution layer is included in the synthesis transform module.

[0477] Clause 97. The method of clause 96, wherein all existing convolution layers or deconvolution layers are replaced by the depth-wise separable convolution layer.

[0478] Clause 98. The method of clause 97, wherein the group number of the depth-wise separable convolution layer is identical to a depth of features.

[0479] Clause 99. The method of clause 97, wherein the group number of the depth-wise separable convolution layer is set as a predetermined number.

[0480] Clause 100. The method of clause 99, wherein the predetermined number is equal to 1, or wherein the predetermined number is the power of 2.

[0481] Clause 101. The method of clause 97, wherein a kernel size of the depth-wise separable convolution layer be KxP, wherein K and P are integer numbers.

[0482] Clause 102. The method of clause 101, wherein both K and P are equal to 1, or wherein both K and P are equal to 2n+1, wherein n is a positive integer number, or wherein values of K and P are not identical.

[0483] Clause 103. The method of clause 96, wherein the depth-wise separable convolution layer, a pointwise convolution layer and a conventional convolution layer are combined and used in the N-th step of the synthesis transform module , and / or wherein a depth-wise separable deconvolution layer, a pointwise deconvolution layer and a conventional deconvolution layer are combined and used in the N-th step of the synthesis transform module, and wherein N is an integer number.

[0484] Clause 104. The method of clause 103, wherein a depth-wise separable convolution layer and a pointwise convolution layer are sequentially combined, and the depth-wise separable convolution layer is first used, and / or wherein the depth-wise separable deconvolution layer and the pointwise deconvolution layer are sequentially combined, and the depth-wise separable convolution layer is first used.

[0485] Clause 105. The method of clause 103, wherein a depth-wise separable convolution layer and a pointwise convolution layer are sequentially combined, and the pointwise convolution layer is first used, and / or wherein the depth-wise separable deconvolution layer and the pointwise deconvolution layer are sequentially combined, and the pointwise convolution layer is first used.

[0486] Clause 106. The method of clause 103, wherein a depth-wise separable convolution layer and a pointwise convolution layer are used as separated branches, and / or wherein the depth-wise separable deconvolution layer and the pointwise deconvolution layer are used as separated branches.

[0487] Clause 107. The method of clause 103, wherein an order of the depth-wise separable convolution layer, a pointwise convolution layer and a conventional convolution layer are arbitrary changed, and / or wherein an order of a depth-wise separable deconvolution layer, a pointwise deconvolution layer and a conventional deconvolution layer are arbitrary changed.

[0488] Clause 108. The method of clause 1, wherein the one or more attention modules comprise at least one of a depth-wise separable convolution layer or a depth-wise separable deconvolution layer.

[0489] Clause 109. The method of clause 108, wherein whether the depth-wise separable convolution layer is used in the one or more attention modules is determined based on available computational resources, and / or wherein where the depth-wise  separable convolution layer is placed is determined based on the available computational resources.

[0490] Clause 110. The method of clause 108, wherein only upscaling layers are involved in the synthesis transform module and the one or more attention modules are excluded from the synthesis transform module.

[0491] Clause 111. The method of clause 110, wherein partial or all of the upscaling layers are implemented as depth-wise separable convolution layers.

[0492] Clause 112. The method of clause 108, wherein the one or more attention modules are placed at the N-th steps of the synthesis transform module, wherein N is an integer number.

[0493] Clause 113. The method of clause 112, wherein N is equal to 0, the one or more attention modules are placed at a location before a first deconvolutional layer.

[0494] Clause 114. The method of clause 113, wherein an input is processed by the one or more attention modules, and then by one or more deconvolution layers.

[0495] Clause 115. The method of clause 112, wherein N is equal to 1, and the one or more attention modules are placed a location after the first deconvolutional layer.

[0496] Clause 116. The method of clause 115, wherein an input is processed by a deconvolution layer, and then by the one or more attention modules.

[0497] Clause 117. The method of clause 112, wherein N is equal to 2, and the one or more attention modules are placed a location after a second deconvolutional layer.

[0498] Clause 118. The method of clause 112, wherein N is equal to 3, and the one or more attention modules are placed a location after a third deconvolutional layer.

[0499] Clause 119. The method of clause 112, wherein N is equal to 4, and the one or more attention modules are placed a location after a fourth deconvolutional layer.

[0500] Clause 120. The method of any of clauses 113-116, wherein a deconvolution layer comprises one or more depth-wise separable deconvolution layers or one or more point-wise deconvolution layers.

[0501] Clause 121. The method of clause 112, wherein the one or more attention modules are placed before cropping layers, or wherein the one or more attention modules  are placed after the cropping layers.

[0502] Clause 122. The method of clause 112, wherein the one or more attention modules are placed before non-linear activation layers, or wherein the one or more attention modules are placed after the non-linear activation layers.

[0503] Clause 123. The method of clause 112, wherein the one or more attention modules are placed before depth-wise separable convolution layers, or wherein the one or more attention modules are placed after the depth-wise separable convolution layers.

[0504] Clause 124. The method of any of clauses 1-123, wherein an indication of whether to and / or how to determine to apply the neural network based image compression network to the video unit is indicated at one of the followings: sequence level, group of pictures level, picture level, slice level, or tile group level.

[0505] Clause 125. The method of any of clauses 1-123, wherein an indication of whether to and / or how to determine to apply the neural network based image compression network to the video unit is indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS) , a video parameter set (VPS) , a dependency parameter set (DPS) , a decoding capability information (DCI) , a picture parameter set (PPS) , an adaptation parameter sets (APS) , a slice header, or a tile group header.

[0506] Clause 126. The method of any of clauses 1-123, wherein an indication of whether to and / or how to determine to apply the neural network based image compression network to the video unit is included in one of the following: a prediction block (PB) , a transform block (TB) , a coding block (CB) , a prediction unit (PU) , a transform unit (TU) , a coding unit (CU) , a coding tree block (CTB) , or a coding tree unit (CTU) .

[0507] Clause 127. The method of any of clauses 1-123, further comprising: determining, based on coded information of the video unit, whether and / or how to determine to apply the neural network based image compression network to the video unit, the coded information including at least one of: a block size, a colour format, a single and / or dual tree partitioning, a colour component, a slice type, or a picture type.

[0508] Clause 128. The method of any of clauses 1-127, wherein the video unit is applied with a coding tool that requires chroma fusion.

[0509] Clause 129. The method of any of clauses 1-128, wherein the SE is binarized as one of a flag, a fixed length code, an EG (x) code, a unary code, a truncated unary code,  or a truncated binary code.

[0510] Clause 130. The method of clause 129, wherein the SE is signed or unsigned.

[0511] Clause 131. The method of any of clauses 1-130, wherein the SE is coded with at least one context model, or wherein the SE is bypass coded.

[0512] Clause 132. The method of any of clauses 1-131, wherein the SE is signaled in a conditional way.

[0513] Clause 133. The method of clause 132, wherein the SE is signaled only if a corresponding function is applicable, or wherein the SE is signaled only if dimensions of the video unit satisfy a condition.

[0514] Clause 134. The method of any of clauses 1-133, wherein the SE is indicated at one of the followings: sequence level, group of pictures level, picture level, slice level, or tile group level.

[0515] Clause 135. The method of any of clauses 1-133, wherein the SE is indicated at one of the followings: a prediction block (PB) , a transform block (TB) , a coding block (CB) , a prediction unit (PU) , a transform unit (TU) , a coding unit (CU) , a coding tree block (CTB) , or a coding tree unit (CTU) .

[0516] Clause 136. The method of any of clauses 1-135, wherein the conversion includes encoding the video unit into the bitstream.

[0517] Clause 137. The method of any of clauses 1-135, wherein the conversion includes decoding the video unit from the bitstream.

[0518] Clause 138. An apparatus for video processing comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to perform a method in accordance with any of clauses 1-137.

[0519] Clause 139. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform a method in accordance with any of clauses 1-137.

[0520] Clause 140. A non-transitory computer-readable recording medium storing a bitstream of a video which is generated by a method performed by an apparatus for video processing, wherein the method comprises: determining to apply a neural network based  image compression network to a video unit of the video, wherein the neural network based image compression network comprises a synthesis transform module that comprises at least one of: one or more upscaling layers and one or more attention modules; and generating the bitstream based on the neural network based image compression network.

[0521] Clause 141. A method for storing a bitstream of a video, comprising: determining to apply a neural network based image compression network to a video unit of the video, wherein the neural network based image compression network comprises a synthesis transform module that comprises at least one of: one or more upscaling layers and one or more attention modules; generating the bitstream based on the neural network based image compression network; and storing the bitstream in a non-transitory computer-readable medium.

[0522] Example Device

[0523] FIG. 22 illustrates a block diagram of a computing device 2100 in which various embodiments of the present disclosure can be implemented. The computing device 2100 may be implemented as or included in the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300) .

[0524] It would be appreciated that the computing device 2100 shown in FIG. 22 is merely for purpose of illustration, without suggesting any limitation to the functions and scopes of the embodiments of the present disclosure in any manner.

[0525] As shown in FIG. 22, the computing device 2100 includes a general-purpose computing device 2100. The computing device 2100 may at least comprise one or more processors or processing units 2110, a memory 2120, a storage unit 2130, one or more communication units 2140, one or more input devices 2150, and one or more output devices 2160.

[0526] In some embodiments, the computing device 2100 may be implemented as any user terminal or server terminal having the computing capability. The server terminal may be a server, a large-scale computing device or the like that is provided by a service provider. The user terminal may for example be any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, station, unit, device, multimedia computer, multimedia tablet, Internet node, communicator, desktop computer, laptop computer, notebook computer, netbook computer, tablet computer, personal  communication system (PCS) device, personal navigation device, personal digital assistant (PDA) , audio / video player, digital camera / video camera, positioning device, television receiver, radio broadcast receiver, E-book device, gaming device, or any combination thereof, including the accessories and peripherals of these devices, or any combination thereof. It would be contemplated that the computing device 2100 can support any type of interface to a user (such as “wearable” circuitry and the like) .

[0527] The processing unit 2110 may be a physical or virtual processor and can implement various processes based on programs stored in the memory 2120. In a multi-processor system, multiple processing units execute computer executable instructions in parallel so as to improve the parallel processing capability of the computing device 2100. The processing unit 2110 may also be referred to as a central processing unit (CPU) , a microprocessor, a controller or a microcontroller.

[0528] The computing device 2100 typically includes various computer storage medium. Such medium can be any medium accessible by the computing device 2100, including, but not limited to, volatile and non-volatile medium, or detachable and non-detachable medium. The memory 2120 can be a volatile memory (for example, a register, cache, Random Access Memory (RAM) ) , a non-volatile memory (such as a Read-Only Memory (ROM) , Electrically Erasable Programmable Read-Only Memory (EEPROM) , or a flash memory) , or any combination thereof. The storage unit 2130 may be any detachable or non-detachable medium and may include a machine-readable medium such as a memory, flash memory drive, magnetic disk or another other media, which can be used for storing information and / or data and can be accessed in the computing device 2100.

[0529] The computing device 2100 may further include additional detachable / non-detachable, volatile / non-volatile memory medium. Although not shown in FIG. 22, it is possible to provide a magnetic disk drive for reading from and / or writing into a detachable and non-volatile magnetic disk and an optical disk drive for reading from and / or writing into a detachable non-volatile optical disk. In such cases, each drive may be connected to a bus (not shown) via one or more data medium interfaces.

[0530] The communication unit 2140 communicates with a further computing device via the communication medium. In addition, the functions of the components in the computing device 2100 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore,  the computing device 2100 can operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs) or further general network nodes.

[0531] The input device 2150 may be one or more of a variety of input devices, such as a mouse, keyboard, tracking ball, voice-input device, and the like. The output device 2160 may be one or more of a variety of output devices, such as a display, loudspeaker, printer, and the like. By means of the communication unit 2140, the computing device 2100 can further communicate with one or more external devices (not shown) such as the storage devices and display device, with one or more devices enabling the user to interact with the computing device 2100, or any devices (such as a network card, a modem and the like) enabling the computing device 2100 to communicate with one or more other computing devices, if required. Such communication can be performed via input / output (I / O) interfaces (not shown) .

[0532] In some embodiments, instead of being integrated in a single device, some or all components of the computing device 2100 may also be arranged in cloud computing architecture. In the cloud computing architecture, the components may be provided remotely and work together to implement the functionalities described in the present disclosure. In some embodiments, cloud computing provides computing, software, data access and storage service, which will not require end users to be aware of the physical locations or configurations of the systems or hardware providing these services. In various embodiments, the cloud computing provides the services via a wide area network (such as Internet) using suitable protocols. For example, a cloud computing provider provides applications over the wide area network, which can be accessed through a web browser or any other computing components. The software or components of the cloud computing architecture and corresponding data may be stored on a server at a remote position. The computing resources in the cloud computing environment may be merged or distributed at locations in a remote data center. Cloud computing infrastructures may provide the services through a shared data center, though they behave as a single access point for the users. Therefore, the cloud computing architectures may be used to provide the components and functionalities described herein from a service provider at a remote location. Alternatively, they may be provided from a conventional server or installed directly or otherwise on a client device.

[0533] The computing device 2100 may be used to implement video encoding / decoding  in embodiments of the present disclosure. The memory 2120 may include one or more video coding modules 2125 having one or more program instructions. These modules are accessible and executable by the processing unit 2110 to perform the functionalities of the various embodiments described herein.

[0534] In the example embodiments of performing video encoding, the input device 2150 may receive video data as an input 2170 to be encoded. The video data may be processed, for example, by the video coding module 2125, to generate an encoded bitstream. The encoded bitstream may be provided via the output device 2160 as an output 2180.

[0535] In the example embodiments of performing video decoding, the input device 2150 may receive an encoded bitstream as the input 2170. The encoded bitstream may be processed, for example, by the video coding module 2125, to generate decoded video data. The decoded video data may be provided via the output device 2160 as the output 2180.

[0536] While this disclosure has been particularly shown and described with references to preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the present application as defined by the appended claims. Such variations are intended to be covered by the scope of this present application. As such, the foregoing description of embodiments of the present application is not intended to be limiting.

Claims

1.A method for video processing, comprising:determining, for a conversion between a video unit of a video and a bitstream of the video, to apply a neural network based image compression network to the video unit, wherein the neural network based image compression network comprises a synthesis transform module that comprises at least one of: one or more upscaling layers and one or more attention modules; andperforming the conversion based on the neural network based image compression network.2.The method of claim 1, wherein the synthesis transform module only comprises the one or more upscaling layers, and the one or more attention modules are excluded from the synthesis transform module.3.The method of claim 1, wherein the one or more attention modules are placed at N-th steps of the synthesis transform module, wherein N is an integer number.4.The method of claim 3, wherein N is equal to 2, and the one or more attention modules are placed a location after a second deconvolutional layer.5.The method of claim 3, wherein N is equal to 3, and the one or more attention modules are placed a location after a third deconvolutional layer.6.The method of claim 3, wherein N is equal to 4, and the one or more attention modules are placed a location after a fourth deconvolutional layer.7.The method of claim 3, wherein N is equal to 0, and the one or more attention modules are placed a location before a first deconvolutional layer.8.The method of claim 7, wherein an input is processed by the one or more attention modules, and then by one or more deconvolution layers.9.The method of claim 3, wherein N is equal to 1, and the one or more attention modules are placed a location after the first deconvolutional layer.10.The method of claim 9, wherein an input is processed by a deconvolution layer, and then by the one or more attention modules.11.The method of claim 3, wherein the one or more attention modules are placed before cropping layers, orwherein the one or more attention modules are placed after the cropping layers.12.The method of claim 3, wherein the one or more attention modules are placed before non-linear activation layers, orwherein the one or more attention modules are placed after the non-linear activation layers.13.The method of claim 1, wherein whether the one or more modules are excluded from the synthesis transform module based on currently available computational resources.14.The method of claim 1, wherein where the one or more modules are placed in the synthesis transform module based on currently available computational resources.15.The method of claim 1, wherein a plurality of types of attention modules is provided with different complexity, orwherein the plurality of types of attention modules is provided with same complexity.16.The method of claim 15, wherein a way to combine the plurality of types of attention modules with a synthesis step is indicated.17.The method of claim 1, wherein whether the one or more modules are involved in the synthesis transform module is indicated, and / orwherein where the one or more modules are placed the synthesis transform module is indicated.18.The method of claim 1, wherein a flag is used to indicate whether the one or more modules are involved in the synthesis transform module.19.The method of claim 1, wherein a syntax element is used to indicate the number of attention modules involved in the synthesis transform module.20.The method of claim 1, wherein a set of attention modules are provided, and a usage of the set of attention modules is determined according to an expected complexity or performance.21.The method of claim 20, wherein a flag is signaled to indicate a priority of complexity, andwherein if a low-complexity is desired, a lightweight attention module is enabled.22.The method of claim 20, wherein a flag is signaled to indicate a priority of reconstruction quality, andwherein if a high-performance is desired, a sophisticated attention module is used.23.The method of claim 20, wherein a flag is included in the bitstream to indicate whether a first attention module is used or a second attention module is used.24.The method of claim 1, wherein a plurality of flags is used to indicate whether the N-th synthesis step includes an attention module, wherein N is an integer number.25.The method of claim 1, wherein a downscaling layer and an upscaling layer are involved in the synthesis transform module.26.The method of claim 25, wherein at least one of: a depth-wise separable deconvolution layer or a point-wise deconvolution layer is involved as part of the downscaling layer and upscaling layer.27.The method of claim 25 or 26, wherein the downscaling layer and the upscaling layer are operated in channel-wise that adjusts feature channel numbers.28.The method of claim 25 or 26, wherein the downscaling layer and the upscaling layer are operated in spatial-wise.29.The method of claim 28, wherein the downscaling layer is applied at a beginning of a mask trunk, and the upscaling layer is applied before an ending of mask trunk.30.The method of claim 28, wherein the downscaling layer is applied at a beginning of a main trunk and a mask trunk, and the upscaling layer is applied after a combination of mask trunk and main trunk.31.The method of claim 28, wherein the downscaling layer is applied at a beginning of a main trunk, and the upscaling layer is applied before an ending of the main trunk.32.The method of any of claims 29 to 31, wherein the downscaling layer decreases a spatial ratio of feature maps, and the upscaling layer recovers a spatial resolution of feature maps.33.The method of claim 28, wherein a combination of the depth-wise separable deconvolution layer the point-wise deconvolution layer acts as the upscaling and downscaling layers and changes a spatial resolution of feature maps.34.The method of claim 1, wherein an attention module of the one or more attention modules comprises a main trunk, a skip trunk and a mask trunk.35.The method of claim 34, wherein the main trunk, the skip trunk and the mask trunk receive an identical input.36.The method of claim 34, wherein an input of the mask trunk is an intermediate output of the main trunk.37.The method of claim 1, wherein an attention module used in the synthesis transform module comprises at least two branches.38.The method of claim 37, wherein a first branch in the attention module comprises: a downscaling layer, at least one convolution layer or a residual block, and an upscaling layer.39.The method of claim 38, wherein the first branch comprises an activation layer that is applied after the upscaling layer.40.The method of claim 39, wherein the activation layer comprises at least one of: a sigmoid layer, a hyperbolic tangent function, or a relu layer.41.The method of claim 38, wherein the upscaling layers increase a size or channel dimension of an input.42.The method of claim 41, wherein the upscaling layer increases a width and height of the input by a factor of 2.43.The method of claim 41, wherein the upscaling layer increases the number of channels of the input by a factor of 2.44.The method of claim 38, wherein the downscaling layer decreases a size or channel dimension of an input.45.The method of claim 44, wherein the downscaling layer decreases a width and height of the input by a factor of 2.46.The method of claim 44, wherein the downscaling layer decreases the number of channels of the input by a factor of 2.47.The method of claim 38, wherein the number of samples of an input is reduced by the downscaling layer and increased by the upscaling layer.48.The method of claim 38, wherein all operations performed till the upscaling layer are performed using reduced number of input samples.49.The method of claim 38, wherein the first branch comprises a depth-wise separable convolution layer that is applied after the upscaling layer.50.The method of claim 37, wherein a second branch in the attention module comprises: a depth-wise separable convolution layer or a residual block.51.The method of claim 37, wherein a second branch in the attention module comprises an identity branch, and an input of the second branch is output without modification.52.The method of claim 37, wherein an output of the first branch and an output of the second branch are multiplied by each other.53.The method of claim 52, wherein a multiplication result of the first and second branches is added to an output of a third branch in the attention module.54.The method of claim 37, wherein the attention module comprises three branches.55.The method of claim 54, wherein a first branch in the three branches comprises: a first downscaling layer, one or more residual blocks, an upscaling layer, a convolution layer and an activation layer in that order.56.The method of claim 55, wherein a depth-wise separable convolution layer is used in at least one of: the first downscaling layer, the one or more residual blocks and the upscaling layer.57.The method of claim 54, wherein a second branch in the three branches comprises one or more residual blocks (RB) .58.The method of claim 54, wherein a first branch in the three branches comprises: a first downscaling layer, three residual blocks, an upscaling layer, a convolution layer and an activation layer in that order.59.The method of claim 58, wherein a depth-wise separable convolution layer is used in at least one of: the first downscaling layer, the three residual blocks and the upscaling layer.60.The method of claim 54, wherein a second branch in the three branches comprises two residual blocks (RBs) , and a depth-wise separable convolution layer is involved in the two RBs.61.The method of claim 54, wherein a second branch in the three branches comprises an identity branch.62.The method of any of claims 54-61, wherein outputs of first and second branches in the three branches are multiplied with each other.63.The method of claim 62, wherein a multiplication result of the first and second branches is added to an input of a third branch which is an identity branch.64.The method of claim 37, wherein the attention module comprises two branches.65.The method of claim 37 or 64, wherein a first branch in the two branches comprises: a first downscaling layer, one or more residual blocks, an upscaling layer, a convolution layer and an activation layer in that order.66.The method of claim 65, wherein a depth-wise separable convolution layer is used in at least one of: the first downscaling layer, the one or more residual blocks and the upscaling layer.67.The method of claim 65, wherein a depth-wise separable convolution layer is used in at least one of: the first downscaling layer and the upscaling layer.68.The method of claim 65, wherein no residual block is applied before the first downscaling layer and after the upscaling layer.69.The method of claim 65, wherein a second branch in the two branches comprises an identity branch.70.The method of claim 65, wherein outputs of first and second branches in the two branches are multiplied with each other, which is an output of the attention module.71.The method of claim 64, wherein a first branch in the two branches is an identity branch.72.The method of claim 64, wherein a second branch in the two branches comprises one or more residual blocks, and after one or more residual blocks, the second branch is split into two branches comprising a third branch and a fourth branch.73.The method of claim 72, wherein the third branch is an identity branch.74.The method of claim 72, wherein the fourth branch comprises one or more other residual blocks.75.The method of claim 74, wherein the fourth branch further comprises an activation layer.76.The method of claim 72, wherein outputs of the third and fourth branches are multiplied with each other to form an output of the second branch.77.The method of claim 1, wherein a residual swin transformer block is used in the synthesis transform module.78.The method of claim 77, wherein at least one of: a depth-wise separable deconvolution layer or a point-wise deconvolution layer are involved as partial of the residual swin transformer block.79.The method of claim 77, wherein the residual swin transformer block and the one or more attention modules are used in the synthesis transform module.80.The method of claim 77, wherein the residual swin transformer block is used in the synthesis transform module and the one or more attention modules are removed from the synthesis transform module.81.The method of claim 77, wherein the residual swin transformer block is partial of the one or more attention modules.82.The method of claim 81, wherein only a swin transformer layer is included in the one or more attention modules.83.The method of claim 81, wherein the residual swin transformer block acts as a new branch and an associated output is combine with an output of the one or more attention modules through one of: adding, subtraction, concatenation, fusion, multiplication, or nonlinear activation.84.The method of claim 77, wherein the residual swin transformer block is a residual block with a swin transformer layer and a depth-wise separable convolutional layer.85.The method of claim 77, wherein the residual swin transformer block is a residual block with a swin transformer layer with a depth-wise separable deconvolutional layer.86.The method of claim 77, wherein the residual swin transformer block is a modified residual block where one branch comprises a swin transformer layer and one branch comprises at least one of: a depth-wise separable convolution layer or a deconvolutional layer.87.The method of claim 77, wherein a swin transformer layer is configured with a head size, a depth size, a window size and a patch size.88.The method of claim 87, wherein the window size is equal to 1, the path size is equal to 1, the depth size is equal to 2, and the head size is equal to 2, orwherein the window size is equal to 1, the path size is equal to 1, the depth size is equal to 4, and the head size is equal to 4, orwherein the window size is equal to 4 and the patch size is equal to 2.89.The method of claim 87, wherein the head size is the same as the depth size, orwherein the head size is equal to 2 and the depth size is equal to 2.90.The method of claim 87, wherein a setting of the swin transformer layer is dependent on a position of residual swin transformer block.91.The method of claim 90, wherein one residual swin transformer block is placed at the N-th step of the synthesis transform module, another residual swin transformer block is placed at the M-th step of the synthesis transform module, wherein N and M are integer numbers.92.The method of claim 91, wherein a setting of the residual swin transformer block and a setting of the other residual swin transformer block are same, orwherein the setting of the residual swin transformer block and the setting of the other residual swin transformer block are different.93.The method of claim 90, wherein a plurality of residual swin transformer blocks are placed within each step of the synthesis transform module.94.The method of claim 93, wherein settings of swin transformer layers are same, orwherein the settings of swin transformer layers are different.95.The method of claim 77, wherein a swin transformer layer comprises a multi-head self-attention layer, a multi-layer perceptron, and a layer normalization.96.The method of claim 1, wherein at least one of: a depth-wise separable convolution layer or a pixel-wise convolution layer is included in the synthesis transform module.97.The method of claim 96, wherein all existing convolution layers or deconvolution layers are replaced by the depth-wise separable convolution layer.98.The method of claim 97, wherein the group number of the depth-wise separable convolution layer is identical to a depth of features.99.The method of claim 97, wherein the group number of the depth-wise separable convolution layer is set as a predetermined number.100.The method of claim 99, wherein the predetermined number is equal to 1, orwherein the predetermined number is the power of 2.101.The method of claim 97, wherein a kernel size of the depth-wise separable convolution layer be KxP, wherein K and P are integer numbers.102.The method of claim 101, wherein both K and P are equal to 1, orwherein both K and P are equal to 2n+1, wherein n is a positive integer number, orwherein values of K and P are not identical.103.The method of claim 96, wherein the depth-wise separable convolution layer, a pointwise convolution layer and a conventional convolution layer are combined and used in the N-th step of the synthesis transform module, and / orwherein a depth-wise separable deconvolution layer, a pointwise deconvolution layer and a conventional deconvolution layer are combined and used in the N-th step of the synthesis transform module, andwherein N is an integer number.104.The method of claim 103, wherein a depth-wise separable convolution layer and a pointwise convolution layer are sequentially combined, and the depth-wise separable convolution layer is first used, and / orwherein the depth-wise separable deconvolution layer and the pointwise deconvolution layer are sequentially combined, and the depth-wise separable convolution layer is first used.105.The method of claim 103, wherein a depth-wise separable convolution layer and a pointwise convolution layer are sequentially combined, and the pointwise convolution layer is first used, and / orwherein the depth-wise separable deconvolution layer and the pointwise deconvolution layer are sequentially combined, and the pointwise convolution layer is first used.106.The method of claim 103, wherein a depth-wise separable convolution layer and a pointwise convolution layer are used as separated branches, and / orwherein the depth-wise separable deconvolution layer and the pointwise deconvolution layer are used as separated branches.107.The method of claim 103, wherein an order of the depth-wise separable convolution layer, a pointwise convolution layer and a conventional convolution layer are arbitrary changed, and / orwherein an order of a depth-wise separable deconvolution layer, a pointwise deconvolution layer and a conventional deconvolution layer are arbitrary changed.108.The method of claim 1, wherein the one or more attention modules comprise at least one of a depth-wise separable convolution layer or a depth-wise separable deconvolution layer.109.The method of claim 108, wherein whether the depth-wise separable convolution layer is used in the one or more attention modules is determined based on available computational resources, and / orwherein where the depth-wise separable convolution layer is placed is determined based on the available computational resources.110.The method of claim 108, wherein only upscaling layers are involved in the synthesis transform module and the one or more attention modules are excluded from the synthesis transform module.111.The method of claim 110, wherein partial or all of the upscaling layers are implemented as depth-wise separable convolution layers.112.The method of claim 108, wherein the one or more attention modules are placed at the N-th steps of the synthesis transform module, wherein N is an integer number.113.The method of claim 112, wherein N is equal to 0, the one or more attention modules are placed at a location before a first deconvolutional layer.114.The method of claim 113, wherein an input is processed by the one or more attention modules, and then by one or more deconvolution layers.115.The method of claim 112, wherein N is equal to 1, and the one or more attention modules are placed a location after the first deconvolutional layer.116.The method of claim 115, wherein an input is processed by a deconvolution layer, and then by the one or more attention modules.117.The method of claim 112, wherein N is equal to 2, and the one or more attention modules are placed a location after a second deconvolutional layer.118.The method of claim 112, wherein N is equal to 3, and the one or more attention modules are placed a location after a third deconvolutional layer.119.The method of claim 112, wherein N is equal to 4, and the one or more attention modules are placed a location after a fourth deconvolutional layer.120.The method of any of claims 113-116, wherein a deconvolution layer comprises one or more depth-wise separable deconvolution layers or one or more point-wise deconvolution layers.121.The method of claim 112, wherein the one or more attention modules are placed before cropping layers, orwherein the one or more attention modules are placed after the cropping layers.122.The method of claim 112, wherein the one or more attention modules are placed before non-linear activation layers, orwherein the one or more attention modules are placed after the non-linear activation layers.123.The method of claim 112, wherein the one or more attention modules are placed before depth-wise separable convolution layers, orwherein the one or more attention modules are placed after the depth-wise separable convolution layers.124.The method of any of claims 1-123, wherein an indication of whether to and / or how to determine to apply the neural network based image compression network to the video unit is indicated at one of the followings:sequence level,group of pictures level,picture level,slice level, ortile group level.125.The method of any of claims 1-123, wherein an indication of whether to and / or how to determine to apply the neural network based image compression network to the video unit is indicated in one of the following:a sequence header,a picture header,a sequence parameter set (SPS) ,a video parameter set (VPS) ,a dependency parameter set (DPS) ,a decoding capability information (DCI) ,a picture parameter set (PPS) ,an adaptation parameter sets (APS) ,a slice header, ora tile group header.126.The method of any of claims 1-123, wherein an indication of whether to and / or how to determine to apply the neural network based image compression network to the video unit is included in one of the following:a prediction block (PB) ,a transform block (TB) ,a coding block (CB) ,a prediction unit (PU) ,a transform unit (TU) ,a coding unit (CU) ,a coding tree block (CTB) , ora coding tree unit (CTU) .127.The method of any of claims 1-123, further comprising:determining, based on coded information of the video unit, whether and / or how to determine to apply the neural network based image compression network to the video unit, the coded information including at least one of:a block size,a colour format,a single and / or dual tree partitioning,a colour component,a slice type, ora picture type.128.The method of any of claims 1-127, wherein the video unit is applied with a coding tool that requires chroma fusion.129.The method of any of claims 1-128, wherein the SE is binarized as one of a flag, a fixed length code, an EG (x) code, a unary code, a truncated unary code, or a truncated binary code.130.The method of claim 129, wherein the SE is signed or unsigned.131.The method of any of claims 1-130, wherein the SE is coded with at least one context model, orwherein the SE is bypass coded.132.The method of any of claims 1-131, wherein the SE is signaled in a conditional way.133.The method of claim 132, wherein the SE is signaled only if a corresponding function is applicable, orwherein the SE is signaled only if dimensions of the video unit satisfy a condition.134.The method of any of claims 1-133, wherein the SE is indicated at one of the followings:sequence level,group of pictures level,picture level,slice level, ortile group level.135.The method of any of claims 1-133, wherein the SE is indicated at one of the followings:a prediction block (PB) ,a transform block (TB) ,a coding block (CB) ,a prediction unit (PU) ,a transform unit (TU) ,a coding unit (CU) ,a coding tree block (CTB) , ora coding tree unit (CTU) .136.The method of any of claims 1-135, wherein the conversion includes encoding the video unit into the bitstream.137.The method of any of claims 1-135, wherein the conversion includes decoding the video unit from the bitstream.138.An apparatus for video processing comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to perform a method in accordance with any of claims 1-137.139.A non-transitory computer-readable storage medium storing instructions that cause a processor to perform a method in accordance with any of claims 1-137.140.A non-transitory computer-readable recording medium storing a bitstream of a video which is generated by a method performed by an apparatus for video processing, wherein the method comprises:determining to apply a neural network based image compression network to a video unit of the video, wherein the neural network based image compression network comprises a synthesis transform module that comprises at least one of: one or more upscaling layers and one or more attention modules; andgenerating the bitstream based on the neural network based image compression network.141.A method for storing a bitstream of a video, comprising:determining to apply a neural network based image compression network to a video unit of the video, wherein the neural network based image compression network comprises a synthesis transform module that comprises at least one of: one or more upscaling layers and one or more attention modules;generating the bitstream based on the neural network based image compression network; andstoring the bitstream in a non-transitory computer-readable medium.