Method and device for video processing and medium
By introducing a simplified attention module and a neural network compression method with deep convolutional layers into video coding and decoding technology, the problems of high computational complexity and low reconstruction quality are solved, and more efficient video coding and decoding effects are achieved.
Patent Information
- Application Number
- CN202480007447.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-20
- Filing Date
- 2024-01-10
- Publication Date
- 2025-10-03
AI Technical Summary
Existing video coding and decoding technologies still have room for improvement in coding and decoding efficiency, especially neural network-based video compression methods, which face challenges in computational complexity and reconstruction quality.
A neural network-based image compression network is adopted, through a simplified attention module and deep convolutional layers, combined with upscaling and downscaling operations, to adjust the allocation of computing resources to reduce computational complexity and improve reconstruction quality.
It improves the computational efficiency of video encoding and decoding, enhances the quality of reconstructed images and videos, and is suitable for simplified synthesis transformation of luminance and chrominance components.
Smart Images

Figure CN120752644A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate generally to video processing techniques, and more particularly to a neural network-based image and video compression method utilizing depthwise convolution and pixel-level convolution for simplified synthetic transformations. Background Art
[0002] Digital video capabilities are now being used in every aspect of our lives. For video encoding and decoding, various video compression technologies have been proposed, including MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-T H.265 High Efficiency Video Codec (HEVC), and Versatile Video Codec (VVC). However, there is a general desire to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method includes: for conversion between a video unit of a video and a bit stream of the video, determining to apply a neural network-based image compression network to the video unit, wherein the neural network-based image compression network includes a synthetic transformation module, and the synthetic transformation module includes at least one of the following: one or more upscaling layers and one or more attention modules; and performing conversion according to the neural network-based image compression network. According to an embodiment of the present disclosure, the synthetic transformation subnetwork is modified. Specifically, it involves a simplified attention module, a deep convolution layer, and a point-by-point convolution layer that downscales the feature map earlier and then upscales the feature map, so that the computational complexity can be reduced. In addition, the positions of the simplified attention module and the deep convolution layer can also be adjusted according to the budget of computing resources.
[0005] In a second aspect, a device for video processing is provided. The device includes a processor and a non-volatile memory having instructions. When the instructions are executed by the processor, the processor performs the method according to the first aspect of the present disclosure.
[0006] In a third aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions, and the instructions cause a processor to execute the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video, the bitstream of the video being generated by a method performed by an apparatus for video processing. The method includes: determining to apply a neural network-based image compression network to a video unit of the video, wherein the neural network-based image compression network includes a synthetic transformation module, the synthetic transformation module including at least one of the following: one or more upscaling layers and one or more attention modules; and generating a bitstream based on the neural network-based image compression network.
[0008] In a fifth aspect, a method for storing a bitstream of a video is provided. The method includes: determining to apply a neural network-based image compression network to a video unit of the video, wherein the neural network-based image compression network includes a synthetic transform module, the synthetic transform module including at least one of the following: one or more upscaling layers and one or more attention modules; generating a bitstream based on the neural network-based image compression network; and storing the bitstream in a non-transitory computer-readable medium.
[0009] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings.In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0011] Figure 1 A block diagram illustrating an example video encoding and decoding system is shown according to some embodiments of the present disclosure;
[0012] Figure 2 shows a block diagram illustrating a first example video encoder according to some embodiments of the present disclosure;
[0013] Figure 3 shows a block diagram illustrating an example video decoder according to some embodiments of the present disclosure;
[0014] Figure 4 is a schematic diagram illustrating an example transform coding scheme;
[0015] Figure 5 An example latent representation of an image is shown;
[0016] Figure 6 is a schematic diagram illustrating an example autoencoder implementing a hyper-prior model;
[0017] Figure 7 is a schematic diagram illustrating an example combined model configured to jointly optimize a context model and a hyper-prior and an autoencoder;
[0018] Figure 8 An example encoding process is shown;
[0019] Figure 9 An example decoding process is shown;
[0020] Figure 10 An example encoder and decoder with a wavelet-based transform is shown;
[0021] Figure 11 shows an example output of a forward wavelet-based transform;
[0022] Figure 12 An example segmentation of the output of a forward wavelet-based transform is shown;
[0023] Figure 13 An example of a synthetic transform module is shown;
[0024] Figure 14 The residual calibration module is shown;
[0025] Figure 15 A simplified attention module is shown;
[0026] Figure 16 The Swin transformer layer is shown;
[0027] Figure 17 A simplified attention module is shown;
[0028] Figure 18 shows a simplified attention module with depthwise separated convolutional layers for downscaling and upscaling;
[0029] Figure 19 Shows a simplified attention module with depthwise separated convolutional layers and pointwise convolutional layers in the residual block;
[0030] Figure 20 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown;
[0031] Figure 21 shows an example of a convolution-based attention block; and
[0032] Figure 22 A block diagram is shown of a computing device in which various embodiments of the present disclosure may be implemented.
[0033] Throughout the drawings, the same or similar reference numbers generally refer to the same or similar elements. DETAILED DESCRIPTION
[0034] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, and do not imply any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.
[0035] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0036] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment will include that particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is intended that such feature, structure, or characteristic, whether or not explicitly described, be applicable to other embodiments and that it is within the knowledge of those skilled in the art to apply that feature, structure, or characteristic.
[0037] It should be understood that although the terms "first" and "second" and the like may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0038] The terms used herein are used only for the purpose of describing specific embodiments and are not intended to limit the example embodiments. As used herein, the singular forms "a," "an," and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "comprise," "including," "having," "including," and / or "comprising" when used herein indicate the presence of the features, elements, and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Sample Environment
[0039] Figure 1is a block diagram illustrating an example video codec system 100 that can utilize the techniques of the present disclosure. As shown, the video codec system 100 can include a source device 110 and a destination device 120. The source device 110 can also be referred to as a video encoding device, and the destination device 120 can also be referred to as a video decoding device. In operation, the source device 110 can be configured to generate encoded video data, and the destination device 120 can be configured to decode the encoded video data generated by the source device 110. The source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0040] The video source 112 may include a source such as a video capture device. Examples of a video capture device include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.
[0041] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include a coded picture and associated data. The coded picture is a coded representation of the picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded video data may be directly sent to the destination device 120 via the network 130A via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130B for access by the destination device 120.
[0042] Destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, the destination device 120 being configured to interface with an external display device.
[0043] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or future standards.
[0044] Figure 2is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of the video encoder 114 in the system 100 is shown.
[0045] Video encoder 200 may be configured to implement any or all of the techniques of this disclosure. Figure 2 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0046] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a cache 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.
[0047] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0048] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are described in detail in the following sections. Figure 2 are shown separately in the example.
[0049] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0050] The mode selection unit 203 can, for example, select one of a plurality of codec modes (intra-frame codec or inter-frame codec) based on the error result, and provide the resulting intra-frame coded block or inter-frame coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combined intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).
[0051] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the buffer 213 other than the picture associated with the current video block.
[0052] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture consisting of macroblocks, all of which are based on macroblocks within the same picture. Furthermore, as used herein, in some aspects, "P slices" and "B slices" may refer to portions of a picture consisting of macroblocks that are independent of macroblocks in the same picture.
[0053] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0054] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 for a reference video block for the current video block, and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 204 may then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference pictures in list 0 and list 1 containing multiple reference video blocks, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. The motion estimation unit 204 may output the multiple reference indices and multiple motion vectors for the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.
[0055] In some examples, motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder. Alternatively, in some embodiments, motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.
[0056] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0057] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0058] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.
[0059] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0060] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0061] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.
[0062] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to the residual video block associated with the current video block.
[0063] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0064] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0065] After reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.
[0066] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0067] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 An example of the video decoder 124 in the system 100 is shown.
[0068] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 3 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0069] exist Figure 3 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally opposite to the encoding process described with respect to the video encoder 200.
[0070] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information, which includes motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, including deriving several most likely candidates based on data from adjacent PBs and reference pictures. The motion information typically includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indexes, and in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially neighboring blocks or temporally neighboring blocks.
[0071] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. Identifiers for the interpolation filters used with sub-pixel precision may be included in the syntax elements.
[0072] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information, and motion compensation unit 302 may use the interpolation filters to produce a prediction block.
[0073] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of the encoded video sequence, partition information describing how each macroblock of the picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame coded block, and other information used to decode the encoded video sequence. As used herein, in some aspects, "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding and decoding, signal prediction, and residual signal reconstruction. A slice can be an entire picture or a region of a picture.
[0074] The intra prediction unit 303 can use, for example, an intra prediction mode received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0075] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on a display device.
[0076] Some exemplary embodiments of the present disclosure are described in detail below. It should be noted that the section headings used in this document are for ease of understanding and do not limit the embodiments disclosed in a section to that section. In addition, although some embodiments are described with reference to a multifunctional video codec or other specific video codec, the disclosed technology is also applicable to other video codec technologies. In addition, although some embodiments describe the video coding and decoding steps in detail, it should be understood that the corresponding decoding steps of the decoding will be implemented by the decoder. In addition, the term video processing includes video coding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at a different compression bit rate. 1. Brief Overview
[0077] This disclosure relates to neural network-based image and video compression methods that utilize autoregressive neural networks. Examples aim to synthesize transforms efficiently for the decoder, thereby improving the quality of the reconstructed image with moderate computational complexity. This disclosure applies to both luminance and chrominance components. 2. Introduction
[0078] Deep learning is developing in various fields, such as computer vision and image processing. Inspired by the successful application of deep learning techniques in the field of computer vision, neural image / video compression techniques are being studied for application in image / video compression techniques. Neural networks are designed based on interdisciplinary research in neuroscience and mathematics. Neural networks have shown powerful capabilities in the context of nonlinear transformations and classification. Examples of neural network-based image compression algorithms achieve RD performance comparable to Versatile Video Codec (VVC), a video codec standard developed by the Joint Video Experts Group (JVET) composed of experts from the Moving Picture Experts Group (MPEG) and the Video Coding Experts Group (VCEG). Neural network-based video compression is an actively developing research field, resulting in continuous improvements in the performance of neural image compression. However, due to the inherent difficulty of the problems solved by neural networks, neural network-based video coding and decoding remains a largely undeveloped discipline. 2.1 Image / Video Compression
[0079] Image / video compression generally refers to computing techniques that compress video images into binary codes to facilitate storage and transmission. Binary codecs may or may not support lossless reconstruction of the original image / video. Codecs with no data loss are called lossless compression, while codecs that allow targeted data loss are called lossy compression. Most codec systems use lossy compression because lossless reconstruction is not required in most scenarios. Typically, the performance of image / video compression algorithms is evaluated based on the resulting compression ratio and reconstruction quality. The compression ratio is directly related to the number of binary codes produced by the compression, where fewer binary codes result in better compression. Reconstruction quality is measured by comparing the reconstructed image / video to the original image / video, where greater similarity results in better reconstruction quality.
[0080] Image / video compression techniques can be divided into video codec methods and neural network-based video compression methods. Video codec schemes adopt transform-based solutions, where statistical dependencies in latent variables (such as discrete cosine transform (DCT) and wavelet coefficients) are exploited to carefully hand-design entropy codecs to model dependencies in the quantization domain. Neural network-based video compression can be grouped into neural network-based codec tools and end-to-end neural network-based video compression. The former is embedded into existing video codecs as a codec tool and acts only as part of the framework, while the latter is a separate framework developed based on neural networks without relying on video codecs.
[0081] A range of video codec standards have been developed to meet the growing demand for visual content delivery. The International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) has two expert groups: the Joint Photographic Experts Group (JPEG) and the Moving Picture Experts Group (MPEG). The International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) also has the Video Codec Experts Group (VCEG), which focuses on standardizing image / video codec technologies. Influential video codec standards released by these organizations include the Joint Photographic Experts Group (JPEG), JPEG 2000, H.262, H.264 / Advanced Video Codec (AVC), and H.265 / High Efficiency Video Codec (HEVC). The Joint Video Experts Group (JVET), formed by MPEG and VCEG, developed the Versatile Video Codec (VVC) standard. VVC is reported to reduce bitrate by an average of 50% compared to HEVC while maintaining the same visual quality.
[0082] Neural network-based image / video compression / encoding is also under development. Exemplary neural network codecs have relatively shallow architectures, and their performance has been unsatisfactory. Neural network-based approaches benefit from abundant data and powerful computing resources, and are therefore better utilized in various applications. Neural network-based image / video compression has shown promising improvements and has been proven to be feasible. However, the technology is far from mature, and many challenges must be addressed. 2.2 Neural Networks
[0083] Neural networks (also known as artificial neural networks (ANNs)) are computational models used in machine learning techniques. Neural networks typically consist of multiple processing layers, and each layer consists of multiple simple but nonlinear basic computational units. One advantage of such deep networks is their ability to process data with multiple levels of abstraction and convert the data into different kinds of representations. The representations created by neural networks are not manually designed. Instead, using general machine learning programs, deep networks comprising processing layers are learned from large amounts of data. Deep learning eliminates the need for handcrafted representations. Therefore, deep learning is considered particularly useful for processing native unstructured data such as acoustic and visual signals. Processing such data has been a long-standing challenge in the field of artificial intelligence. 2.3 Neural Networks for Image Compression
[0084] Neural networks for image compression can be divided into two categories: pixel-wise probabilistic models and autoencoder models. Pixel-wise probabilistic models employ a predictive encoding / decoding strategy, while autoencoder models employ a transform-based solution. Sometimes, these two approaches are combined. 2.3.1 Pixel Probabilistic Modeling
[0085] According to Shannon's information theory, the optimal method for lossless coding achieves the minimum codec rate, expressed as -log2p(x), where p(x) is the probability of symbol x. Arithmetic coding is a lossless coding method considered one of the optimal methods. Given a probability distribution p(x), arithmetic coding achieves a codec rate as close to the theoretical limit -log2p(x) as possible without taking into account rounding errors. Therefore, the remaining problem is to determine the probability, which is particularly challenging for natural images and videos due to the curse of dimensionality. The curse of dimensionality refers to the problem that increasing the number of dimensions makes the dataset sparse, so as the number of dimensions increases, the amount of data required to effectively analyze and organize the data increases rapidly.
[0086] Following the predictive encoding and decoding strategy, one way to model p(x) is to predict pixel probabilities one by one in raster scan order based on previous observations, where x is an image, which can be expressed as follows: p(x)=p(x1)p(x2|x1)…p(x i |x1,…,x i-1 )…p(x m×n |x1,…,x m×n-1 )(1) Where m and n are the height and width of the image, respectively. The previous observation is also called the context of the current pixel. When the image is large, the estimation of the conditional probability may be difficult. Therefore, a simplified approach is to limit the context range of the current pixel as follows: p(x)=p(x1)p(x2|x1)…p(x i |x i-k ,…,x i-1 )…p(x m×n |x m×n-k ,…,x m×n-1 )(2) Where k is a predefined constant that controls the scope of the context.
[0087] It should be noted that this condition can also take into account the sample values of other color components. For example, when encoding and decoding red (R), green (G), and blue (B) (RGB) color components, the R sample depends on the previously encoded and decoded pixels (including R, G, and / or B samples), and the current G sample can be encoded and decoded based on the previously encoded and decoded pixels and the current R sample. In addition, when encoding and decoding the current B sample, the previously encoded and decoded pixels as well as the current R and G samples can also be considered.
[0088] Neural networks can be designed for computer vision tasks and are also effective in regression and classification problems. Therefore, neural networks can be used to predict the probability of a given context x1, x2, ..., x i-1 In the case of estimating p(x i ). In the example neural network design, according to x i ∈{-1,+1}, pixel probabilities are adopted for binary images. A neural autoregressive distribution estimator (NADE) is designed for pixel probability modeling. NADE is a feedforward network with a single hidden layer. In another example, the feedforward network can include connections that skip the hidden layer. In addition, parameters can also be shared. The example design performs experiments on the binarized MNIST dataset. In one example, NADE is extended to a real-valued NADE (RNADE) model, where the probability p(x i |x1,…,x i-1) is derived using a Gaussian mixture. The RNADE model feed-forward network also has a single hidden layer, but the hidden layer is rescaled to avoid saturation and uses rectified linear units (ReLU) instead of sigmoids. In another example, NADE and RNADE are improved by reorganizing the order of pixels and using deeper neural networks.
[0089] Designing advanced neural networks plays a key role in improving pixel-wise probabilistic modeling. In the example neural network, a multidimensional long short-term memory (LSTM) is used. LSTM is used in conjunction with a mixture of conditional Gaussian scaled mixtures for probabilistic modeling. LSTM is a special type of recurrent neural network (RNN) and can be used to model sequential data. Spatial variants of LSTM can also be used for images. Several different neural networks can be used, including recurrent neural networks (RNNs) and convolutional neural networks (CNNs), such as the pixel RNN (PixelRNN) and the pixel CNN (PixelCNN). In the PixelRNN, two variants of LSTM are used, called row LSTM and diagonal bidirectional LSTM (BiLSTM). The diagonal BiLSTM is specifically designed for images. The PixelRNN incorporates residual connections to facilitate the training of deep neural networks with up to twelve layers. In the PixelCNN, masked convolutions are used to adjust for the shape of the context. The PixelRNN and PixelCNN are more specialized for natural images. For example, the PixelRNN and PixelCNN treat pixels as discrete values (e.g., 0, 1, ..., 255) and predict a multinomial distribution over these discrete values. In addition, PixelRNN and PixelCNN process color images in the RGB color space. In addition, PixelRNN and PixelCNN work well on the large-scale image dataset ImageNet. In one example, a gated PixelCNN is used to improve PixelCNN. The gated PixelCNN achieves comparable performance to PixelRNN, but with lower complexity. In one example, PixelCNN++ is used, and the following improvements are made to PixelCNN: using discretized logistic mixture likelihood instead of 256-way multinomial distribution; using downsampling to capture structure at multiple resolutions; introducing additional shortcut connections to speed up training; using dropout for regularization; and combining RGB into one pixel. In another example, PixelSNAIL combines accidental convolution with self-attention.
[0090] Most of the above methods directly model the probability distribution in the pixel domain. Some designs also model the probability distribution as a conditional probability distribution based on explicit or latent representations. Such a model can be expressed as: Where h is an additional condition, and p(x)=p(h)p(x|h) indicates that the modeling is divided into an unconditional model and a conditional model. The additional condition can be image label information or a high-level representation. 2.3.2 Autoencoder
[0091] Now let's describe autoencoders. Autoencoders are trained for dimensionality reduction and include an encoding component and a decoding component. The encoding component converts a high-dimensional input signal into a low-dimensional representation. The low-dimensional representation may have a reduced spatial size but a greater number of channels. The decoding component recovers the high-dimensional input from the low-dimensional representation. Autoencoders enable automatic learning of representations and eliminate the need for manually engineered features, which is considered one of the most important advantages of neural networks.
[0092] Figure 4 is a schematic diagram illustrating an example transform coding scheme 400. The original image x is analyzed by the analysis network g a The latent representation y is quantized (q) and compressed into bits. The number of bits R is used to measure the codec rate. Then the synthetic network g s Inverse transform to obtain the reconstructed image Distortion (D) is calculated in the perceptual space by using the function g p Transform x and To calculate, we get z and Compare z and To get D.
[0093] Autoencoder networks can be applied to lossy image compression. The learned latent representation can be encoded from a well-trained neural network. However, adapting autoencoders to image compression is not easy because the original autoencoder is not optimized for compression and is therefore not efficient to use directly as a trained autoencoder. In addition, there are other major challenges: first, the low-dimensional representation should be quantized before being encoded, however quantization is not differentiable, which is required in backpropagation while training the neural network. Second, the goals in the compression scenario are different because both distortion and ratio need to be considered. Estimating the ratio is challenging. Third, practical image coding and decoding schemes should support variable ratios, scalability, encoding / decoding speed, and interoperability. In response to these challenges, various schemes are under development.
[0094] The example autoencoder for image compression can be viewed as a transform coding strategy using the example transform coding scheme 400. The original image x is obtained using the analysis network y = g a (x) is transformed, where y is the latent representation, which is quantized and encoded. The synthesis network inverse transforms the quantized latent representation To obtain the reconstructed image The framework is trained using the rate-distortion loss function, i.e. where D is x and The distortion between them, R is the quantization representation is the ratio of the calculated or estimated values, and λ is the Lagrange multiplier. D can be calculated in the pixel domain or the receptive domain. Most example systems follow this prototype, and the only difference between them is the network structure or loss function.
[0095] In terms of network structure, RNN and CNN are the most widely used architectures. In the RNN-related category, an example general framework for variable-rate image compression uses RNN. The example uses binary quantization to generate codecs and does not consider the ratio during training. The framework provides scalable codec capabilities, in which the RNN has convolutional and deconvolution layers. Another example provides an improved version by upgrading the encoder with a neural network similar to PixelRNN to compress binary codes. Using the multi-scale structural similarity (MS-SSIM) evaluation metric, the performance on the Kodak image dataset is better than JPEG. By introducing hidden state activation, another example further improves the RNN-based solution. In addition, an SSIM weighted loss function is designed, and a spatial domain adaptive bitrate mechanism is included. Using MS-SSIM as the evaluation metric, the example achieves better results than Better Portable Graphics (BPG) on the Kodak image dataset. Another example system supports spatial domain adaptive bitrate by training a stop-codec tolerant RNN.
[0096] Another example proposes a general framework for rate-distortion optimized image compression. The example system uses multi-base quantization to generate integer codecs and considers the ratio during training. The loss is a joint rate-distortion cost, which can be the mean squared error (MSE) or other metric. The example system adds random uniform noise to stimulate quantization during training and uses the differential entropy of the noise codec as a proxy for the ratio. The example system uses generalized division normalization (GDN) as the network structure, which includes a linear mapping followed by nonlinear parameter normalization. The effectiveness of GDN in image codecs is verified. Another example system includes an improved version that uses three convolutional layers, each followed by a downsampling layer and a GDN layer as the forward transform. Therefore, this example version uses a three-layer inverse GDN (IGDN), each followed by an upscaling layer and a convolution layer to stimulate the inverse transform. In addition, an arithmetic coding method is designed to compress integer codecs. It is reported that the performance on the Kodak dataset is better than JPEG and JPEG2000 in terms of MSE. Another example improves this approach by designing a scale hyper-prior into the autoencoder. The system utilizes a subnetwork h a Transform the latent representation y into z = h a(y), and z is quantized and transmitted as side information. Therefore, the inverse transform is performed using the sub-network h s The sub-network h s From the quantitative side information Decoded to quantized The standard deviation of This method is further used during arithmetic coding and decoding of the image. On the Kodak image set, this method slightly underperforms BPG in terms of peak signal-to-noise ratio (PSNR). Another example system further exploits the structure in the residual space by introducing an autoregressive model to estimate the standard deviation and mean. This example uses a Gaussian mixture model to further eliminate redundancy in the residuals. Using PSNR as the evaluation metric, performance is comparable to VVC on the Kodak image set. 2.3.3 Super-prior model
[0097] Figure 5 An example latent representation of an image is shown. Figure 5 The image 501 from the Kodak dataset, the visualization of the potential 502 representation y of the image 501, the standard deviation σ503 of the potential 502, and the potential y 504 after the introduction of the super-prior network. The super-prior network includes an encoder and a decoder that utilize the super-prior information. In the transform coding method of image compression, such as Figure 4 As shown, the encoder subnetwork uses parameter analysis transformation g a (x,φ g ) transforms the image vector x into a latent representation y, which is then quantized to form because is a discrete value, so It can be losslessly compressed using entropy coding techniques such as arithmetic coding and transmitted as a bit sequence.
[0098] from Figure 5 The potential 502 and standard deviation σ503 can be clearly seen, There is a significant spatial dependence between the elements of . It is noteworthy that their scales (standard deviation σ503) appear to be coupled in the spatial domain. An additional set of random variables can be introduced to capture spatial dependencies and further reduce redundancy. In this case, the image compression network such as Figure 6 shown.
[0099] Figure 6 600 is a diagram illustrating an example network architecture of an autoencoder implementing a super prior model. The upper side shows the image autoencoder network, and the lower side corresponds to the super prior subnetwork. The analysis transform and the synthesis transform are denoted as g a and g sQ represents quantization, AE and AD represent arithmetic encoder and arithmetic decoder respectively. The super prior model consists of two sub-networks, the encoder (denoted as h a ) and the decoder using super prior information (denoted as h s ). The super-prior model generates a quantitative super-prior information potential value This includes quantifying potential value Information related to the probability distribution of the sample points. is included in the bitstream and is are transmitted together to the receiver (decoder).
[0100] In the diagram 600, the upper side of the model is the encoder g as discussed above. a and decoder g s The lower side is used to obtain The additional encoder h using super prior information a and decoder h using super prior information s In this architecture, the encoder subjects the input image x to g a , producing a response y with a spatially varying standard deviation. The response y is fed to h a , the distribution of standard deviations in z is summarized. z is then quantized is compressed and transmitted as side information. The encoder then uses the quantized vector To estimate the spatial distribution of the standard deviation σ, and use σ to compress and transmit the quantized image representation The decoder first recovers the compressed signal The decoder then uses h s To obtain σ, this provides the decoder with the correct probability estimate to successfully recover The decoder will then Feed to g s to obtain the reconstructed image.
[0101] When an encoder utilizing super-prior information and a decoder utilizing super-prior information are added to an image compression network, the quantized latent value The spatial redundancy is reduced. Figure 5 The potential value y 504 in corresponds to the quantized potential value when using an encoder / decoder that utilizes super-prior information. Compared with the standard deviation σ 503, the spatial redundancy is significantly reduced because the correlation of the samples of the quantized potential value is low. 2.3.4 Context Model
[0102] Although the hyper-prior model improves the quantitative potential value However, further improvement can be achieved by exploiting autoregressive models that predict the quantized latent value from its causal context (referred to as contextual models).
[0103] The term autoregressive indicates that the output of a process is then used as input to that process. For example, the context model subnetwork generates a sample of potential values, which is then used as input to get the next sample.
[0104] Figure 7 700 is a diagram illustrating an example combined model configured to jointly optimize a context model with a hyperprior and an autoencoder. The combined model jointly optimizes an autoregressive component that estimates the probability distribution of a latent value from its causal context (context model) with the hyperprior and the underlying autoencoder. The real-valued latent representation is quantized (Q) to create a quantized latent value. and quantify the potential value of super-prior information They are compressed into a bitstream using an arithmetic encoder (AE) and decompressed by an arithmetic decoder (AD). The dotted area corresponds to the components executed by the receiver (e.g., decoder) to recover the image from the compressed bitstream.
[0105] The example system utilizes a joint architecture where both the super-prior model sub-network (encoder utilizing super-prior information and decoder utilizing super-prior information) and the context model sub-network are utilized. The super-prior and context models are combined to learn to quantize the potential value The probability model on the quantized potential value is then used for entropy coding and decoding. As shown in schematic 700, the outputs of the context sub-network and the decoder sub-network that utilizes the super-prior information are combined by a sub-network called the entropy parameter, which generates the mean μ and scale (or variance) σ parameters for the Gaussian probability model. The Gaussian probability model is then used to encode the samples of the quantized potential value into a bitstream with the help of the arithmetic encoder (AE) module. In the decoder, the Gaussian probability model is used to obtain the quantized potential value from the bitstream through the arithmetic decoder (AD) module.
[0106] In an example, the potential sample is modeled as a Gaussian distribution or a Gaussian mixture model (without limitation). In the example according to schematic diagram 700, the context model and the hyper-prior are jointly used to estimate the probability distribution of the potential sample. Since the Gaussian distribution can be defined by a mean and a variance (also known as sigma or scale), the joint model is used to estimate the mean and variance (denoted as μ and σ). 2.3.5 Encoding Process Using Joint Autoregressive Hyperprior Model
[0107] Figure 4The design in corresponds to the example combined compression method. In this section and the next, the encoding and decoding processes are described separately.
[0108] Figure 8 An example encoding process 800 is shown. The input image is first processed by the encoder sub-network. The encoder transforms the input image into a transformed representation called a potential value, denoted by y. y is then input to the quantizer block, denoted by Q, to obtain a quantized potential value Then it is converted into a bit stream (bit stream 1, bits1) using the arithmetic coding module (denoted as AE). The arithmetic coding block sequentially converts Each sample point is converted into a bit stream (bits1) one by one.
[0109] The module uses an encoder with super prior information, context, a decoder with super prior information, and an entropy parameter subnetwork to estimate the quantized potential value. The potential value y is input to the encoder using super prior information, and the encoder using super prior information outputs the super prior information potential value (denoted as z). The super prior information potential value is then quantized The arithmetic coding (AE) module is used to generate a second bit stream (bit stream 2, bits2). The decomposition entropy module generates a probability distribution, which is used to encode the quantized super prior information potential value into a bit stream. The quantized super prior information potential value includes the quantized potential value information about the probability distribution of .
[0110] The entropy parameter subnetwork is generated to encode the quantized potential value The information generated by the entropy parameters usually includes the mean μ and scale (or variance) σ parameters, which are used together to obtain a Gaussian probability distribution. The Gaussian distribution of a random variable x is defined as The parameter μ is the mean or expectation of the distribution (also its median and mode), and the parameter σ is its standard deviation (or variance or scale). To define a Gaussian distribution, you need to determine the mean and variance. The Entropy Parameter module is used to estimate the mean and variance values.
[0111] The decoder of the sub-network using the super prior information generates part of the information used by the entropy parameter sub-network, and the other part of the information is generated by an autoregressive module called the context module. The context module uses the samples that have been encoded by the arithmetic coding (AE) module to generate information about the probability distribution of the samples of the quantized potential value. It is usually a matrix composed of many sample points. Sample points can be indicated by using indices, such as or It depends on the matrix Dimensions of . Sample points Encoded one by one by the AE, usually using raster scan order. In raster scan order, the rows of the matrix are processed from top to bottom, where the samples in the row are processed from left to right. In such a scenario (where the AE encodes the samples into the bitstream using raster scan order), the context module uses the samples that were previously encoded in raster scan order to generate the same samples as the ones in the previous step. The information generated by the context module and the decoder using the super prior information is combined by the entropy parameter module to generate the information used to quantize the potential value Encoded as a probability distribution of the bitstream (bits1).
[0112] Finally, as a result of the encoding process, the first bit stream and the second bit stream are transmitted to a decoder. It should be noted that other names may also be used for the above modules.
[0113] In the above description, Figure 8 All elements in are collectively called encoders. The analysis transformation that transforms the input image into the latent representation is also called an encoder (or autoencoder). 2.3.6 Decoding Process Using Joint Autoregressive Hyper-Prior Model
[0114] Figure 9 An example decoding process 900 is shown. Figure 9 Describe the decoding process separately.
[0115] During the decoding process, the decoder first receives the first bit stream (bits1) and the second bit stream (bits2) generated by the corresponding encoder. Bits2 is first decoded by the arithmetic decoding (AD) module using the probability distribution generated by the factorized entropy sub-network. The factorized entropy module typically uses a predetermined template to generate the probability distribution, such as a predetermined mean and variance value in the case of a Gaussian distribution. The output of the arithmetic decoding process for bits2 is It is the quantized super-prior information potential value. The AD process is restored to the AE process applied in the encoder. The AE and AD processes are lossless, which means that the quantized super-prior information potential value generated by the encoder is Can be reconstructed at the decoder without any changes.
[0116] In obtaining After that, it is processed by the decoder using super prior information, and the output of the decoder using super prior information is fed into the entropy parameter module. The three sub-networks used in the decoder, context, decoder using super prior information, and entropy parameters are the same as those in the encoder. Therefore, the exact same probability distribution can be obtained in the decoder (as in the encoder), which is very important for lossless reconstruction of the quantized potential value. As a result, the quantized potential value obtained in the encoder is The same version of can be obtained in the decoder.
[0117] After obtaining the probability distribution (e.g., mean and variance parameters) through the entropy parameter subnetwork, the arithmetic decoding module decodes the quantized potential value samples one by one from the bitstream bits1. From a practical point of view, the autoregressive model (context model) is serial in nature and therefore cannot be accelerated using techniques such as parallelization. Finally, the fully reconstructed quantized potential value is input to the composite transform (in Figure 9 ) module to obtain the reconstructed image.
[0118] In the above description, Figure 9 All elements in are collectively referred to as the decoder. The synthetic transform that transforms the quantized latent values into the reconstructed image is also called the decoder (or autodecoder). 2.3.7 Wavelet-based Neural Compression Architecture
[0119] Figure 8 The analysis transform in (represented as an encoder) and Figure 9 The synthetic transform in (denoted as decoder) can be replaced by a wavelet-based transform. Figure 8 An example diagram 800 of an encoder and decoder with a wavelet-based transform is shown. Figure 10 An example of such an implementation is shown. In this figure, the input image is first converted from RGB color format to YUV color format. This conversion process is optional and may be missing in other implementations. However, if such a conversion is applied at the input image, the inverse conversion (from YUV to RGB) is also applied before the output image is generated. In addition, 2 additional post-processing modules (Post-processing 1 and 2) are shown in this figure. These modules are also optional and may be missing in other implementations. The core of the encoder with wavelet-based transform consists of a wavelet-based forward transform, a quantization module and an entropy encoding and decoding module. After these 3 modules are applied to the input image, a bitstream is generated. The core of the decoding process consists of entropy decoding, an inverse quantization process and an inverse wavelet-based transform operation. The decoding process converts the bitstream into an output image. The encoding and decoding process is as follows Figure 10 shown.
[0120] Figure 10 An example encoder and decoder 1000 with a wavelet-based transform is shown.
[0121] After applying the wavelet-based forward transform to the input image, the image is divided into its frequency components in the output of the wavelet-based forward transform. The output of the 2D forward wavelet transform (depicted as the iWave forward module in the above figure) can be obtained using Figure 11The input to the transformation is an image of a castle. In this example, after the transformation, the output has seven distinct regions. The number of distinct regions depends on the specific implementation of the transformation and can be different from 7. Possible numbers of regions are 4, 7, 10, 13, etc.
[0122] Figure 11 Example output 1100 of a forward wavelet-based transform.
[0123] exist Figure 11 In the image above, you can see that the input image is transformed into seven regions, with three small images and four even smaller images. The transformation is based on frequency components, with the small image in the lower right quarter containing high-frequency components in both the horizontal and vertical directions. On the other hand, the smallest image in the upper left corner contains the lowest frequency components in both the vertical and horizontal directions. The small image in the upper right quarter contains high-frequency components in the horizontal direction and low-frequency components in the vertical direction.
[0124] Figure 12 An example segmentation 1200 of the output of a forward wavelet-based transform is shown. Figure 12 Depicts possible partitioning of the latent representation after a 2D forward transform. The latent representation is the sample point (latent sample point or quantized latent sample point) obtained after a 2D forward transform. The latent sample point is divided into the seven parts above, denoted as HH1, LH1, HL1, LL2, HL2, LH2, and HH2. HH1 describes that the part includes high-frequency components in the vertical direction and high-frequency components in the horizontal direction, and the partition depth is 1. HL2 describes that the part includes low-frequency components in the vertical direction and high-frequency components in the horizontal direction, and the partition depth is 2.
[0125] After the latent samples are obtained by forward wavelet transform at the encoder, entropy coding and decoding are used to transmit the latent samples to the decoder. At the decoder, entropy decoding is applied to obtain the latent samples, which are then inversely transformed (by using Figure 10 The iWave inverse module in
[15] is used to obtain the reconstructed image. 2.4 Neural Networks for Video Compression
[0126] Similar to video codecs, neural image compression serves as the foundation for intra-frame compression in neural network-based video compression. The development of neural network-based video compression has lagged behind that of neural network-based image compression due to its greater complexity and the associated challenges that require more effort to address. Compared to image compression, video compression requires effective methods to eliminate inter-frame picture redundancy. Inter-frame picture prediction is a key step in these example systems. Motion estimation and compensation are widely used in video codecs, but are not typically implemented using trained neural networks.
[0127] Neural network-based video compression can be divided into two categories based on the target scenario: random access and low latency. In the random access case, the system allows decoding to begin at any point in the sequence, typically dividing the entire sequence into multiple separate segments and allowing each segment to be decoded independently. In the low latency case, the system aims to reduce decoding time, allowing subsequent frames to be decoded using temporally previous frames as reference frames. 2.4.1 Low Latency
[0128] The example system employs a video compression scheme that utilizes a trained neural network. The system first divides a video sequence frame into blocks, and each block is encoded or decoded according to either intra or inter codec mode. If intra codec is selected, an associated autoencoder compresses the block. If inter codec is selected, motion estimation and compensation are performed, and a trained neural network is used for residual compression. The output of the autoencoder is directly quantized and encoded using the Huffman method.
[0129] Another neural network-based video codec employs the PixelMotionCNN. Frames are compressed sequentially in time, and each frame is divided into blocks, which are compressed in raster scan order. Each frame is first inferred using the two previous reconstructed frames. When a block is to be compressed, the inferred frame, along with the current block's context, is fed into the PixelMotionCNN to derive a latent representation. The residual is then compressed using a variable-rate image scheme. This scheme achieves performance comparable to H.264.
[0130] Another example system employs an end-to-end neural network-based video compression framework, in which all modules are implemented using neural networks. This scheme accepts the current frame and a previously reconstructed frame as input. Optical flow is derived using a pretrained neural network as motion information. Motion information is warped using a reference frame, and a motion-compensated frame is then generated by the neural network. The residual and motion information are compressed using two separate neural autoencoders. The entire framework is trained using a single rate-distortion loss function. The example system achieves better performance than H.264.
[0131] Another example system utilizes an advanced neural network-based video compression scheme. This system inherits and extends the video codec scheme using a neural network with the following key features. First, it uses only a single autoencoder to compress motion information and residuals. Second, it employs motion compensation using multiple frames and multiple optical flows. Third, it uses online states that are learned and propagated over time through subsequent frames. This scheme achieves better performance than the HEVC reference software on MS-SSIM.
[0132] Another example system uses an extended end-to-end neural network-based video compression framework. In this example, multiple frames are used as references. Therefore, by utilizing multiple reference frames and their associated motion information, the example system is able to provide more accurate predictions of the current frame. Furthermore, motion field prediction is implemented to eliminate motion redundancy along temporal channels. A post-processing network is also used to remove reconstruction artifacts from previous passes. This system significantly outperforms H.265 in both PSNR and MS-SSIM.
[0133] Another example system uses a scaled spatial stream, replacing the optical flow with a frame-by-frame scaling parameter. This system can achieve better performance than H.264. Another example system uses a multi-resolution representation based on optical flow. Specifically, the motion estimation network generates multiple optical flows at different resolutions and learns which one to select under a loss function. This achieves slightly better performance than H.265. 2.4.2 Random Access
[0134] Another system uses a neural network-based video compression scheme with frame interpolation. Keyframes are compressed first using a neural image compressor, and the remaining frames are compressed in a hierarchical order. This system performs motion compensation in the perceptual domain by deriving feature maps at multiple spatial scales of the original frame and using the motion to warp the feature maps. The result is used in the image compressor. This approach is comparable to H.264.
[0135] One example system uses an approach for interpolation-based video compression. The interpolation model combines motion information compression and image synthesis. The same autoencoder is used for both the image and the residual. Another example system employs a neural network-based video compression approach based on a variational autoencoder with a deterministic encoder. Specifically, the model includes an autoencoder and an autoregressive prior. Unlike previous approaches, this system accepts a group of pictures (GOP) as input and incorporates a three-dimensional (3D) autoregressive prior by taking into account temporal correlation when encoding and decoding the latent representation. The system provides performance comparable to H.265. 2.5 Preliminary Knowledge
[0136] Almost all natural images and / or videos are in digital format. A grayscale digital image can be represented as in is a set of pixel values, m is the image height, and n is the image width. For example, is an example setup, in this case Therefore, a pixel can be represented by an 8-bit integer. An uncompressed grayscale digital image has 8 bits per pixel (bpp), while compressed bits are certainly less.
[0137] Color images are usually represented in multiple channels to record color information. For example, in the RGB color space, an image can be represented by Representation, where three separate channels store red, green, and blue information. Similar to an 8-bit grayscale image, an uncompressed 8-bit RGB image has 24bpp. Digital images / videos can be represented in different color spaces. Neural network-based video compression schemes are mostly developed in the RGB color space, while video codecs typically use the YUV color space to represent video sequences. In the YUV color space, the image is decomposed into three channels, namely luminance (Y), blue difference chrominance (Cb), and red difference chrominance (Cr). Y is the luminance component, and Cb and Cr are the chrominance components. Since Cb and Cr are usually downsampled to achieve pre-compression, the compression benefit of YUV occurs because the human visual system is less sensitive to the chrominance components.
[0138] A color video sequence consists of multiple color images (also called frames) to record scenes at different time stamps. For example, in the RGB color space, a color video can be represented as X = {x0, x1, ..., x t ,…,x T-1}, where T is the number of frames in the video sequence and If m=1080, n=1920, And the video has 50 frames per second (fps), then the data rate of the uncompressed video is 1920×1080×8×3×50=2,488,320,000 bits per second (bps). This results in about 2.32 gigabits per second (Gbps), which uses a lot of storage and should be compressed before transmission over the Internet.
[0139] Typically, lossless methods can achieve compression ratios of approximately 1.5 to 3 for natural images, which is significantly lower than the streaming requirement. Therefore, lossy compression is employed to achieve better compression ratios, but at the expense of distortion. Distortion can be measured by calculating the mean squared difference between the original image and the reconstructed image, for example based on the mean squared error (MSE). For grayscale images, the MSE can be calculated using the following equation.
[0140] Therefore, the quality of the reconstructed image compared to the original image can be measured by the Peak Signal-to-Noise Ratio (PSNR): in yes The maximum value in , for example, for 8-bit grayscale images is 255. There are other quality assessment metrics such as structural similarity (SSIM) and multi-scale SSIM (MS-SSIM).
[0141] To compare different lossless compression schemes, a given ratio of the results can be used to compare the compression ratio, or vice versa. However, to compare different lossy compression methods, the comparison must take into account both bit rate and reconstruction quality. For example, this can be achieved by calculating the relative ratios at several different quality levels and then averaging the rates. The average relative ratio is called the Bjontegaard delta ratio (BD ratio). Other aspects of evaluating image and / or video codecs include encoding / decoding complexity, scalability, robustness, etc. 3. Technical Problems Solved by the Disclosed Technical Solution 3.1. Core Issues
[0142] Existing image compression networks typically include a synthetic transform module, such as a convolutional neural network with an attention module, to reconstruct the image from the latent codec. Deep synthetic transform modules can facilitate the recovery of visual signals. However, finding the optimal trade-off between complexity and reconstruction capability is challenging. Heavy designs are often accompanied by high computational complexity, which hinders their practical applications. Conversely, overly simple synthetic transform modules may have limited ability to reconstruct the image, which degrades the quality of the decompressed image. Existing methods rarely consider simplifying the design of the synthetic module while maintaining its reconstruction capability. 3.2 Background and details of the problem
[0143] Figure 13 An example of a synthetic transformation module is shown. Figure 13 In
[15] , an example of the synthetic transform module used in the JPEG-AI validation model is detailed. Given the luminance potential codec (i.e. input), brightness reconstruction (e.g. output) can be obtained through the synthesis transformation module. The synthesis transformation module may contain residual blocks, deconvolution layers, nonlinear activations, cropping layers and attention modules, convolution layers, and post-processing modules. The computational complexity of the attention module and post-processing is prominent, accounting for more than 80% of the overall synthesis complexity. Figure 13 In the example shown, a total of 4 upscaling layers (e.g., deconvolution layers) are involved, corresponding to the 4 steps of the reconstruction process, where each step may contain a deconvolution layer, a cropping layer, and a nonlinear activation layer. The upscaling (or upsampling or deconvolution) layer is depicted in the figure using convolution Cpx3x3 / 2↑. In addition, Figure 10 In the example of
[15] , the attention module (RNAB) is depicted. In addition to the block diagram of the synthetic transformation, the block diagrams of RNAB, ResAU and residual block (RB) are also depicted.
[0144] The above upsampling (e.g. deconvolution or upscaling) layers increase the width and height of the input tensor. For example, an upsampling layer can double the width and height of the input tensor.
[0145] The attention module may include at least two branches. The output of the first branch and the output of the second branch are multiplied with each other.
[0146] The attention module can include 3 branches. The output of the first branch and the output of the second branch are multiplied with each other, and their output is added to the third branch. One of these branches can be an identity branch, which means that the identity branch does not perform any operation on the input. 4. Detailed solutions
[0147] The detailed solutions below should be considered as examples to explain the general concept. These solutions should not be interpreted in a narrow sense. In addition, these solutions can be combined in any way. 4.1 Core of the Solution
[0148] The goal of the solution is to improve the reconstruction capability of the synthetic transform module under the constraints of computing resources. The core of the present disclosure is to simplify the synthetic transform module while maintaining the reconstruction capability. The structure of the attention module and the position of the attention module can be modified. The convolutional layers used by the attention module and the synthetic transform can be modified to further reduce the model complexity, which can be further applied / extended to other parts of the codec, such as the analysis transform, the entropy codec module, the context modeling module, the prediction fusion module, the codec module using super prior information, etc. 4.2 Solution Details 1. Depthwise separable convolution and / or pixel-wise convolution can be included in the synthesis transform. a) In one example, all existing convolutional layers (or deconvolutional layers) are replaced by depthwise separable convolutions. i. In one example, the number of groups of depthwise separable convolutional layers is consistent with the depth of the feature (i.e., the number of channels). ii. In one example, the number of groups of separable convolutional layers in the depthwise direction is set to F. 1. In one example, F is equal to 1. 2. In one example, F may be a power of 2. iii. The kernel size of the depthwise separable convolution can be K×P. 1. In one example, K and P are equal to 1. 2. In one example, K and P are equal to 2n+1, where n can be a positive integer. 3. In one example, the values of K and P may not be consistent. b) In one example, depthwise separable convolution (or deconvolution), pointwise convolution (or deconvolution), and regular convolution (or deconvolution) may be combined and used in the Nth step of the composite transform. i. In one example, depthwise separable convolution (or deconvolution) and pointwise convolution (or deconvolution) may be sequentially combined, with depthwise separable convolution (or deconvolution) being used first. 1. Alternatively, point-wise convolution (or deconvolution) is used first. ii. In one example, depthwise separable convolution (or deconvolution) and pointwise convolution (or deconvolution) are used as separate branches. iii. In one example, the order of depthwise separable convolution (or deconvolution), point-wise convolution (or deconvolution), and regular convolution (or deconvolution) can be arbitrarily changed. 2. The attention module can involve depthwise separable convolution or deconvolution layers. a) In one example, whether depthwise separable convolution is used in the attention module or where the depthwise separable convolution layer is placed can be determined by available computational resources. b) In one example, only the upscaling layer is involved in the synthesis transformation. The attention module is excluded from the synthesis transformation to adapt to extremely low-complexity decoding scenarios. Part or all of the upscaling layer is implemented as a depthwise separable convolutional layer. c) In one example, the attention module is placed at the Nth step of the synthesis transformation. i. In one example, N is equal to 0, corresponding to the position before the first deconvolution layer (e.g., the upsampling layer of the upscaling layer). In this example, the input First processed by the attention module, and then processed by one or more deconvolution (upsampling or upscaling) layers. Here, the deconvolution layer can be one or more depth-wise separable deconvolution layers and / or point-wise deconvolution layers. ii. In one example, N is equal to 1, corresponding to the position after the first deconvolution layer. In this example, the input It is first processed by a deconvolution (upsampling or upscaling) layer and then by an attention module. Here, the deconvolution layer can be one or more depth-wise separable deconvolution layers and / or point-wise deconvolution layers. iii. In one example, N is equal to 2, corresponding to the position after the second deconvolution layer. Here, the deconvolution layer can be one or more depthwise separable deconvolution layers and / or point-wise deconvolution layers. iv. In one example, N is equal to 3, corresponding to the position after the third deconvolution layer. Here, the deconvolution layer can be one or more depthwise separable deconvolution layers and / or point-wise deconvolution layers. v. In one example, N is equal to 4, corresponding to the position after the fourth deconvolution layer. Here, the deconvolution layer can be one or more depthwise separable deconvolution layers and / or point-wise deconvolution layers. vi. Alternatively, for example, the attention module can be placed before or after the cropping layer. Here, the deconvolution layer can be one or more depth-wise separable deconvolution layers and / or point-wise deconvolution layers. vii. Alternatively, for example, the attention module can be placed before or after the non-linear activation layer. viii. Alternatively, for example, the attention module can be placed before the depthwise separable convolutional layer or after. 3. The synthetic transformation can contain upscaling layers and / or attention modules. a) In one example, whether an attention module is involved or where the attention module is placed may be determined by available computing resources. b) In one example, only the upscaling layer is involved in the synthesis transformation. The attention module is excluded from the synthesis transformation to adapt to extremely low-complexity decoding scenarios. c) In one example, the attention module is placed at the Nth step of the synthesis transformation. i. In one example, N is equal to 0, corresponding to the position before the first deconvolution layer (e.g., the upsampling layer of the upscaling layer). In this example, the input It is first processed by an attention module and then by one or more deconvolution (upsampling or upscaling) layers. ii. In one example, N is equal to 1, corresponding to the position after the first deconvolution layer. In this example, the input It is first processed by a deconvolution (upsampling or upscaling) layer and then by an attention module. iii. In one example, N is equal to 2, corresponding to the position after the second deconvolution layer. iv. In one example, N is equal to 3, corresponding to the position after the third deconvolution layer. v. In one example, N is equal to 4, corresponding to the position after the fourth deconvolution layer. vi. Alternatively, for example, the attention module can be placed before or after the clipping layer. vii. Alternatively, for example, the attention module can be placed before or after the non-linear activation layer. 4. The synthetic transformation can contain upscaling layers and / or attention modules. a) In one example, whether an attention module is involved or where the attention module is located can be signaled. b) In one example, a flag is used to signal whether the attention module is involved in synthetic transformations. c) In one example, a syntax is used to indicate how many attention modules are involved in the synthesis transformation. d) In one example, a set of attention modules is provided. The use of attention modules can be determined based on the expected complexity or performance. i. In one example, a flag is signaled to indicate the priority of complexity. When low complexity is more desired, the lightweight attention module can be enabled. ii. In one example, a flag is signaled to indicate the priority of reconstruction quality. When higher performance is required, a complex attention module can be used. iii. In another example, a flag is included in the bitstream to indicate whether the first attention module is used or the second attention module is used. e) In one example, multiple flags can be used to indicate the Nth synthesis step (ie, deconvolution layer, Whether to include attention modules (cropping layers and non-linear deconvolution layers). f) In one example, multiple types of attention modules with different or equal complexity are provided. We further explore how to combine these attention modules with the synthesis step through signaling. 5. Downscaling and upscaling layers are involved to save computational complexity. Multiple depthwise separable deconvolution layers and / or point-wise deconvolution layers are involved as part of the downscaling and upscaling layers. a) In one example, downscaling and upscaling are operated in the spatial domain level. i. In one example, a downscaling layer is applied at the beginning of the main backbone and the mask backbone to reduce the spatial resolution of the feature map. After the combination of the mask backbone and the main backbone, an upscaling layer is applied to restore the spatial resolution of the feature map. ii. In one example, a downscaling layer is applied at the beginning of the mask backbone to reduce the spatial resolution of the feature map. Before the end of the mask backbone, an upscaling layer is applied to restore the spatial resolution of the feature map. iii. In one example, a downscaling layer is applied at the beginning of the main backbone to reduce the spatial resolution of the feature map. Before the end of the main backbone, an upscaling layer is applied to restore the spatial resolution of the feature map. iv. In one example, a combination of depthwise separable deconvolution and / or point-wise deconvolution layers constitutes upscaling and downscaling layers, changing the spatial resolution of the feature map. b) Alternatively, for example, the downscaling layer and the upscaling layer can operate at the channel level, and the channel level adjusts the features Number of channels. 6. The attention module can include the main backbone, skip backbone and mask backbone. The three backbones receive the same input. a) Alternatively, the input of the mask backbone may be, for example, an intermediate output of the main backbone. 7. The residual Swin transformer block can be used in the synthesis transform. Depthwise separable deconvolution layers and / or point-wise deconvolution layers are involved as part of the residual Swin transformer block. a) In one example, both the residual Swin transformer block and the attention module can be used in the synthetic transform. b) In one example, the residual Swin transformer block is used in the synthetic transform and the attention module is removed from the synthetic transform. c) In one example, the residual Swin transformer block can be part of the attention module. i. In one example, only the Swin transformer layer is included in the attention module. ii. In one example, the residual Swin transformer can act as a new branch and perform addition / subtraction / concatenation / fusion / multiplication / non-linear activation, the associated output can be combined with the output of the attention module. d) In one example, the residual Swin transformer block is a residual block having a Swin transformer layer and a depthwise separable convolutional layer. e) In one example, the residual Swin transformer block is a residual block having a Swin transformer layer and a depthwise separable deconvolution layer (e.g., an upsampling layer of an upscaling layer). f) In one example, the residual Swin transformer block is a modified residual block where one branch contains Swin transformer layers, and one branch contains depthwise separable convolutional and / or deconvolutional layers. g) In one example, a Swin transformer layer may be configured with a head H, a depth D, a window size W, and a patch size P. i. In one example, W=1, P=1, D=2, H=2. ii. In one example, W=1, P=1, D=4, H=4. iii. In one example, W=4, P=2. iv. In one example, H is the same as D. v. In one example, H is equal to 2, and D is equal to 2. vi. In one example, the setting of the Swin transformer layer may depend on the location of the residual Swin transformer block. 1. In one example, two residual Swin transformer blocks are placed in the Nth At step 1 and step M, the settings of the Swin transformer layers can be the same or different. 2. In one example, multiple residual Swin transformer blocks are placed within each step of the synthetic transform. The settings of the Swin transformer layers can be the same or different. h) In one example, a Swin Transformer layer consists of a multi-head self-attention layer, a multi-layer perceptron, and layer normalization. 8. In one example, an attention module is used in the synthesis transformation, wherein the attention module includes at least two branches. a) The first branch first includes a downscaling layer, then at least one convolutional layer, or a depthwise separable convolutional layer or residual block, and finally an upscaling layer. i. Alternatively or additionally, the first branch may include an activation layer, such as a Sigmoid layer, a hyperbolic tangent function, a ReLU layer, etc., which is applied after applying the upscaling layer. ii. Alternatively or additionally, the first branch may include a depthwise separable convolution layer on which the scaling layer is applied. iii. Upscaling layers can increase the size of the input in height or width or channel dimension. 1. In one example, the upscaling layer may increase the width and height of the input by a factor of ×2. 2. In one example, the upscaling layer may increase the number of channels of the input (eg, the number of channel maps) by a factor of ×2. iv. Downscaling layers can reduce the size of the input in height or width or channel dimensions. 1. In one example, the downscaling layer may reduce the width and height of the input by a factor of ×2. 2. In one example, the downscaling layer may reduce the number of channels of the input (eg, the number of channel maps) by a factor of ×2. v. The number of input samples can be reduced by the lower scaling layer and increased by the upper scaling layer. vi. Since the first branch starts from the lower zoom layer, all operations performed until the upper zoom layer All operations are performed using a reduced number of input samples, thereby reducing complexity. b) The second branch can include a depthwise separable convolutional layer or a residual block. c) The second branch may be an identity branch, which means that the input of the second branch is output without any modification. d) The output of the first branch and the output of the second branch may be multiplied with each other. i. Alternatively or additionally, the multiplication result (of the first branch and the second branch) may be added to the third The output of the branch. e) In one example, the attention module includes 3 branches. i. The first branch includes the first downscaling layer, 3 residual blocks, upscaling layer, convolution layer, and activation layer in the following order: Depthwise separable convolution layers can be used in the downscaling layer, residual blocks, and upscaling layer. ii. The second branch includes two residual blocks (RBs). The depthwise separable convolutional layer can involve residual blocks. iii. The output of the first branch and the output of the second branch are multiplied with each other. iv. The multiplication result is added to the input (ie the third branch is the identity branch). f) In one example, the attention module includes 3 branches. i. The first branch includes a first downscaling layer, one or more residual blocks, an upscaling layer, a convolutional layer, and an activation layer in the following order: Depthwise separable convolutional layers can be used in the downscaling layer, the residual block, and the upscaling layer. ii. The second branch includes one or more residual blocks (RBs). iii. The output of the first branch and the output of the second branch are multiplied with each other. iv. The multiplication result is added to the input (ie the third branch is the identity branch). g) In one example, the attention module includes 3 branches. i. The first branch includes a first downscaling layer, one or more residual blocks, an upscaling layer, a convolutional layer, and an activation layer in the following order: Depthwise separable convolutional layers can be used in the downscaling layer, the residual block, and the upscaling layer. ii. The second branch is the identity branch. iii. The output of the first branch and the output of the second branch are multiplied with each other. iv. The multiplication result is added to the input (ie the third branch is the identity branch). h) In one example, the attention module includes 2 branches. i. The first branch includes the first down-scaling layer, one or more residual blocks, up-scaling layer, Convolutional layers and activation layers. Depthwise separable convolutional layers can be used in downscaling layers, residual blocks, and upscaling layers. ii. The second branch is the identity branch. iii. The output of the first branch and the output of the second branch are multiplied with each other, and this multiplication is the output of the attention module. i) In one example, the attention module includes at least 2 branches. i. The first branch includes the first down-scaling layer, one or more residual blocks, up-scaling layer, Convolutional layers and activation layers. No residual blocks are applied before the first downscaling layer and after the upscaling layer (thus, computationally complex residual block operations are performed on a reduced number of samples). Depthwise separable convolutional layers can be used in both downscaling and upscaling layers. j) In one example (e.g., Figure 11 ), the attention module consists of 2 branches. i. The first branch is the identity branch. ii. The second branch first includes one or more residual blocks (in other words, the input of the second branch is first processed using one or more residual blocks), and then the second branch is divided into 2 branches (branch 3 and branch 4): 1. The third branch is the identity branch. 2. The fourth branch includes one or more residual blocks, and it may include an activation layer (e.g., Sigmoid layer or hyperbolic tangent function, etc.). 3. The output of the third branch and the output of the fourth branch are multiplied with each other to form the output of the second branch. General aspects 9. Whether and / or how to apply the method disclosed above can be transmitted through signals at the block level / sequence level / picture group level / picture level / slice level / slice group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. 10. Whether and / or how to apply the method disclosed above may depend on the coded information, such as block size, color format, single / dual tree partitioning, color component, slice / picture type. 11. The proposed method disclosed in this document can be used in other codecs that require chroma fusion. 12. The syntax elements disclosed above can be binarized as flag, fixed length codec, EG(x) codec, unary codec, truncated unary codec, truncated binary codec, etc. It can be signed or unsigned. 13. The syntax elements disclosed above can be encoded or decoded using at least one context model, or they can be bypassed. 14. The syntax elements disclosed above may be signaled in a conditional manner. a. SE is signaled if and only if the corresponding function applies. b. SE is signaled if and only if the dimensions (width and / or height) of the block meet the conditions. 15. The syntax elements disclosed above can be transmitted by signal at the block level / sequence level / picture group level / picture level / slice level / slice group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. 4.3 Benefits of the Disclosure
[0149] According to the present disclosure, the synthesis transformation subnetwork is modified. This disclosure involves reducing computational complexity by simplifying the attention module, depthwise convolutional layer, and pointwise convolutional layer by downscaling the feature map earlier and then upscaling it. In addition, the position of the simplified attention module and depthwise convolutional layer can be adjusted according to the computational resource budget. 5. Examples
[0150] The residual calibrated attention module is used in the synthetic transformation, where the structure of the residual calibrated attention module is as follows Figure 14 shown. Figure 14 The residual calibration module is shown.
[0151] A simplified attention module is used in the synthesis transformation, where the structure is as follows Figure 15 shown. Figure 15 A simplified attention module is shown.
[0152] The residual Swin transformer block is used in the synthesis transform, where the structure of the Swin transformer layer is as follows: Figure 16 shown. Figure 16 The Swin transformer layer is shown.
[0153] The simplified attention module is used in the synthetic transformation, where the structure of the residual attention module is as follows Figure 17 shown. Figure 17 A simplified attention module is shown.
[0154] A simplified attention module with deep convolutional layers for upscaling and downscaling is used in the synthetic transformation, where the structure is as follows Figure 18 shown. Figure 18A simplified attention module with depthwise separated convolutional layers for downscaling and upscaling is shown.
[0155] A simplified attention module with depthwise convolutional layers and pointwise convolutional layers is used in the residual block of the synthetic transformation, where the structure is as follows Figure 19 shown. Figure 19 A simplified attention module with depthwise separated convolutional layers and pointwise convolutional layers in the residual block is shown.
[0156] As used herein, the term "video unit" or "video block" may be a sequence, a picture, a slice, a tile, a sub-picture, a codec tree unit (CTU) / codec tree block (CTB), a CTU / CTB row, one or more codec units (CU) / codec blocks (CB), one or more CTU / CTBs, one or more virtual pipeline data units (VPDU), or a sub-region within a picture / slice / slice / tile.
[0157] Figure 20 FIG2 is a flow chart of a method 2000 for video processing according to an embodiment of the present disclosure. The method 2000 is implemented during conversion between a video unit of a video and a bitstream of the video.
[0158] At box 2010, for conversion between a video unit of a video and a bitstream of the video, it is determined to apply a neural network-based image compression network to the video unit, wherein the neural network-based image compression network includes a synthetic transformation module, and the synthetic transformation module includes at least one of the following: one or more upscaling layers and one or more attention modules.
[0159] At box 2020, the conversion is performed according to the neural network-based image compression network. In some embodiments, the conversion may include encoding the video unit into a bitstream. Alternatively or additionally, the conversion may include decoding the video unit from the bitstream. According to an embodiment of the present disclosure, the synthetic transformation subnetwork is modified. Specifically, it involves a simplified attention module, a depth convolution layer, and a point-by-point convolution layer that downscales the feature map earlier and then upscales the feature map, so that the computational complexity can be reduced. In addition, the position of the simplified attention module and the depth convolution layer can also be adjusted according to the budget of computing resources.
[0160] In some embodiments, the synthesis transform module includes only one or more upscaling layers, and one or more attention modules are excluded from the synthesis transform module. For example, one or more attention modules are excluded from the synthesis transform to accommodate very low complexity decoding scenarios.
[0161] In some embodiments, one or more attention modules are placed at step N of the synthetic transformation module, where N is an integer. For example, N is equal to 2, and the one or more attention modules are placed after the second deconvolution layer. As another example, N is equal to 3, and the one or more attention modules are placed after the third deconvolution layer. As another example, N is equal to 4, and the one or more attention modules are placed after the fourth deconvolution layer.
[0162] In some embodiments, N is equal to 0, and one or more attention modules are placed before the first deconvolution layer (e.g., an upsampling layer before an upscaling layer). In this case, the input can be processed by one or more attention modules and then by one or more deconvolution (e.g., upsampling or upscaling) layers.
[0163] In some embodiments, N is equal to 1, and one or more attention modules are placed after the first deconvolution layer. In this case, the input is processed by the deconvolution layer and then by one or more attention modules. In other words, the input It is first processed by a deconvolution (upsampling or upscaling) layer and then by an attention module.
[0164] In some embodiments, one or more attention modules are placed before the clipping layer. Alternatively, one or more attention modules are placed after the clipping layer.
[0165] In some embodiments, one or more attention modules are placed before the non-linear activation layer. Alternatively, one or more attention modules are placed after the non-linear activation layer.
[0166] In some embodiments, whether one or more modules are excluded from the composite transform module is based on currently available computing resources. In some other embodiments, where one or more modules are placed in the composite transform module is based on currently available computing resources.
[0167] In some embodiments, multiple types of attention modules are provided with different complexities. Alternatively, multiple types of attention modules are provided with the same complexity. In some embodiments, the manner in which the multiple types of attention modules are combined with the synthesis step can be indicated.
[0168] In some embodiments, whether one or more modules are related to a composite transform module is indicated. Alternatively or additionally, where in the composite transform module the one or more modules are located is indicated.
[0169] In some embodiments, a flag is used to indicate whether one or more modules are involved in a synthetic transform module. In some other embodiments, a syntax element is used to indicate the number of attention modules involved in a synthetic transform module.
[0170] In some embodiments, a set of attention modules is provided, in which case the use of the set of attention modules can be determined based on the expected complexity or performance.
[0171] In some embodiments, a flag is signaled to indicate a priority of complexity. In this case, if low complexity is desired, the lightweight attention module can be enabled.
[0172] Alternatively, a flag is signaled to indicate the priority of reconstruction quality. In this case, where high performance is desired, a complex attention module can be used. In some other embodiments, a flag is included in the bitstream to indicate whether the first attention module or the second attention module is used.
[0173] In some embodiments, multiple flags are used to indicate whether the Nth synthesis step includes an attention module, where N is an integer. For example, the Nth synthesis step may include a deconvolution layer, a cropping layer, and a nonlinear deconvolution layer.
[0174] In some embodiments, the downscaling layer and the upscaling layer involve a synthetic transform module. In some embodiments, at least one of a depthwise separable deconvolution layer or a point-wise deconvolution layer is involved as part of the downscaling layer and the upscaling layer.
[0175] In some embodiments, the down-scaling layer and the up-scaling layer are operated at a channel level, which adjusts the number of feature channels. Alternatively, the down-scaling layer and the up-scaling layer are operated at a spatial level.
[0176] For example, the downscaling layer is applied at the beginning of the mask stem, and the upscaling layer is applied before the end of the mask stem. For example, the downscaling layer is applied at the beginning of the main stem and the mask stem, and the upscaling layer is applied after the combination of the mask stem and the main stem. As another example, the downscaling layer is applied at the beginning of the main stem, and the upscaling layer is applied before the end of the main stem. As another example, a combination of a depthwise separable deconvolution layer and a point-by-point deconvolution layer acts as an upscaling and downscaling layer, and changes the spatial resolution of the feature map. In some embodiments, the downscaling layer reduces the spatial ratio of the feature map, and the upscaling layer restores the spatial resolution of the feature map.
[0177] In some embodiments, the attention module in one or more attention modules includes a main trunk, a skip trunk, and a mask trunk. In some embodiments, the main trunk, the skip trunk, and the mask trunk receive consistent inputs. Alternatively, the input of the mask trunk is an intermediate output of the main trunk.
[0178] In some embodiments, the attention module used in the synthetic transformation module includes at least two branches. In some embodiments, the first branch in the attention module includes: a downscaling layer, at least one convolutional layer or residual block, and an upscaling layer. In some embodiments, the first branch includes a depthwise separable convolutional layer, which is applied after the upscaling layer. Alternatively or additionally, the first branch includes an activation layer, which is applied after the upscaling layer. The activation layer may include one or more of the following: a sigmoid layer, a hyperbolic tangent function, or a ReLU layer.
[0179] In some embodiments, the upscaling layer increases the size or channel dimension of the input. In some embodiments, the upscaling layer increases the width and height of the input by a factor of 2. In some other embodiments, the upscaling layer increases the number of channels (e.g., the number of channel maps) of the input by a factor of 2.
[0180] In some embodiments, the downscaling layer reduces the size or channel dimension of the input. In some embodiments, the downscaling layer reduces the width and height of the input by a factor of 2. In some other embodiments, the downscaling layer reduces the number of channels (e.g., the number of channel maps) of the input by a factor of 2.
[0181] In some embodiments, the number of input samples is reduced by the lower scaling layer and increased by the upper scaling layer. In some other embodiments, all operations performed up to the upper scaling layer are performed using the reduced number of input samples. In this case, since the first branch starts at the lower scaling layer, it can reduce complexity.
[0182] In some embodiments, the second branch in the attention module comprises: a depthwise separable convolutional layer or a residual block. In some other embodiments, the second branch in the attention module comprises an identity branch, and the input of the second branch is output without modification. In this case, it means that the output of the first branch and the output of the second branch are multiplied with each other.
[0183] In some embodiments, the output of the first branch and the output of the second branch can be multiplied with each other. In some other embodiments, the multiplication result of the first branch and the second branch is added to the output of the third branch in the attention module.
[0184] In some embodiments, the attention module includes three branches. For example, the first of the three branches includes: a first downscaling layer, one or more residual blocks, an upscaling layer, a convolutional layer, and an activation layer in the following order. In some embodiments, a depthwise separable convolutional layer is used in at least one of the following: a first downscaling layer, one or more residual blocks, and an upscaling layer. Alternatively or additionally, the second of the three branches includes one or more residual blocks (RBs).
[0185] In some embodiments, the first branch of the three branches includes: a first down-scaling layer, three residual blocks, an up-scaling layer, a convolution layer, and an activation layer in the following order. In some embodiments, a depthwise separable convolution layer is used in at least one of the following: a first down-scaling layer, three residual blocks, and an up-scaling layer. In some embodiments, the second branch of the three branches includes two residual blocks (RBs), and the depthwise separable convolution layer involves two RBs. In some other embodiments, the second branch of the three branches includes an identity branch. In some embodiments, the output of the first branch and the output of the second branch of the three branches are multiplied with each other. In some other embodiments, the multiplication result of the first branch and the second branch is added to the input of the third branch, and the input of the third branch is the identity branch.
[0186] In some embodiments, the attention module includes two branches. In some embodiments, the first of the two branches includes: a first down-scaling layer, one or more residual blocks, an up-scaling layer, a convolutional layer, and an activation layer in the following order. In some embodiments, a depthwise separable convolutional layer is used in at least one of the following: the first down-scaling layer, one or more residual blocks, and the up-scaling layer. In some other embodiments, a depthwise separable convolutional layer is used in at least one of the following: the first down-scaling layer and the up-scaling layer. In some embodiments, no residual block is applied before the first down-scaling layer and after the up-scaling layer.
[0187] In some embodiments, the second branch of the two branches comprises an identity branch.In some embodiments, the output of the first branch and the output of the second branch of the two branches are multiplied with each other, and the multiplication is the output of the attention module.
[0188] In some embodiments, as Figure 14 As shown, the first branch of the two branches is an identity branch. In some embodiments, the second branch of the two branches includes one or more residual blocks, and the second branch is divided into two branches after the one or more residual blocks, and the two branches include a third branch and a fourth branch. In some embodiments, the third branch is an identity branch. In some embodiments, the fourth branch includes one or more other residual blocks. The fourth branch may also include an activation layer. For example, the fourth branch may include a sigmoid layer or a hyperbolic tangent function. In some embodiments, the output of the third branch and the output of the fourth branch are multiplied with each other to form the output of the second branch.
[0189] In some embodiments, a residual Swin transformer block is used in a synthetic transform module. For example, at least one of a depthwise separable deconvolution layer or a point-wise deconvolution layer is involved as part of the residual Swin transformer block.
[0190] In some embodiments, the residual Swin transformer block and one or more attention modules are used in the synthetic transform module. In some other embodiments, the residual Swin transformer block is used in the synthetic transform module and one or more attention modules are removed from the synthetic transform module.
[0191] In some embodiments, the residual Swin transformer block is part of one or more attention modules. For example, only the Swin transformer layer is included in one or more attention modules. As another example, the residual Swin transformer block acts as a new branch, and the associated output is combined with the output of one or more attention modules by one of the following: addition, subtraction, concatenation, fusion, multiplication, or nonlinear activation.
[0192] In some embodiments, the residual Swin transformer block is a residual block having a Swin transformer layer and a depthwise separable convolution layer. In some other embodiments, the residual Swin transformer block is a residual block having a Swin transformer layer and a depthwise separable deconvolution layer (e.g., an upsampling layer of an upscaling layer). In some other embodiments, the residual Swin transformer block is a modified residual block in which one branch includes a Swin transformer layer and one branch includes at least one of the following: a depthwise separable convolution layer or a depthwise separable deconvolution layer.
[0193] In some embodiments, the Swin transformer layer is configured with a head size, a depth size, a window size, and a tile size. In some embodiments, the window size is equal to 1, the path size is equal to 1, the depth size is equal to 2, and the head size is equal to 2. Alternatively, the window size is equal to 1, the path size is equal to 1, the depth size is equal to 4, and the head size is equal to 4. In some other embodiments, the window size is equal to 4, and the tile size is equal to 2.
[0194] In some embodiments, the head size is the same as the depth size. Alternatively, the head size is equal to 2 and the depth size is equal to 2.
[0195] In some embodiments, the placement of the Swin transformer layers depends on the location of the residual Swin transformer blocks. For example, one residual Swin transformer block is placed at step N of the composite transform module, and another residual Swin transformer block is placed at step M of the composite transform module, where N and M are integers.
[0196] In some embodiments, the settings of the residual Swin transformer block are the same as those of another residual Swin transformer block. Alternatively, the settings of the residual Swin transformer block are different from those of another residual Swin transformer block.
[0197] In some embodiments, multiple residual Swin transformer blocks are placed within each step of the composite transform module. In some embodiments, the configuration of the Swin transformer layers is the same. Alternatively, the configuration of the Swin transformer layers is different. In some embodiments, the Swin transformer layers include a multi-head self-attention layer, a multi-layer perceptron, and layer normalization.
[0198] In some embodiments, at least one of a depthwise separable convolution layer or a pixel-wise convolution layer is included in the synthesis transformation module. For example, all existing convolution layers or deconvolution layers are replaced by depthwise separable convolution layers.
[0199] In some embodiments, the number of groups of depthwise separable convolutional layers is the same as the depth of the feature. For example, the depth of the feature can be the number of channels.
[0200] In some embodiments, the number of groups of depthwise separable convolutional layers is set to a predetermined number. For example, the predetermined number is equal to 1. As another example, the predetermined number is a power of 2.
[0201] In some embodiments, the kernel size of the depthwise separable convolutional layer is K×P, where K and P are integers. For example, K and P are both equal to 1. For example, K and P are both equal to 2n+1, where n is a positive integer. As another example, the values of K and P are different.
[0202] In some embodiments, depthwise separable convolution layers, pointwise convolution layers, and regular convolution layers are combined and used in the Nth step of the synthesis transformation module. Alternatively or additionally, depthwise separable deconvolution layers, pointwise deconvolution layers, and regular deconvolution layers are combined and used in the Nth step of the synthesis transformation module. In this case, N can be an integer.
[0203] In some embodiments, the depthwise separable convolution layer and the pointwise convolution layer are sequentially combined, and the depthwise separable convolution layer is used first. Alternatively or additionally, the depthwise separable deconvolution layer and the pointwise deconvolution layer are sequentially combined, and the depthwise separable convolution layer is used first.
[0204] In some embodiments, the depthwise separable convolution layer and the pointwise convolution layer are sequentially combined, and the pointwise convolution layer is used first. Alternatively or additionally, the depthwise separable deconvolution layer and the pointwise deconvolution layer are sequentially combined, and the pointwise convolution layer is used first.
[0205] In some embodiments, depthwise separable convolutional layers and pointwise convolutional layers are used as separate branches. Alternatively or additionally, depthwise separable deconvolutional layers and pointwise deconvolutional layers are used as separate branches.
[0206] In some embodiments, the order of depthwise separable convolutional layers, pointwise convolutional layers, and regular convolutional layers is arbitrarily changed. Alternatively or additionally, the order of depthwise separable deconvolutional layers, pointwise deconvolutional layers, and regular deconvolutional layers is arbitrarily changed.
[0207] In some embodiments, one or more attention modules include at least one of a depthwise separable convolutional layer or a depthwise separable deconvolutional layer. In some embodiments, whether a depthwise separable convolutional layer is used in one or more attention modules is determined based on available computational resources. Alternatively or additionally, where the depthwise separable convolutional layer is placed is determined based on available computational resources.
[0208] In some embodiments, only the upscaling layer involves the synthesis transformation module, and one or more attention modules are excluded from the synthesis transformation module. In this case, part or all of the upscaling layer is implemented as a depthwise separable convolution layer.
[0209] In some embodiments, one or more attention modules are placed at step N of the synthesis transformation module, where N is an integer. For example, if N is 0, the one or more attention modules are placed before the first deconvolution layer. In some embodiments, the input is processed by the one or more attention modules and then processed by the one or more deconvolution layers.
[0210] In some embodiments, N is equal to 1, and one or more attention modules are placed at a position after the first deconvolution layer. In some embodiments, the input is processed by the deconvolution layer, and then the input is processed by the one or more attention modules. In some embodiments, N is equal to 2, and one or more attention modules are placed at a position after the second deconvolution layer. In some other embodiments, N is equal to 3, and one or more attention modules are placed at a position after the third deconvolution layer. In some further embodiments, N is equal to 4, and one or more attention modules are placed at a position after the fourth deconvolution layer. In some embodiments, the deconvolution layer includes one or more depthwise separable deconvolution layers or one or more point-by-point deconvolution layers.
[0211] In some embodiments, one or more attention modules are placed before the clipping layer. Alternatively, one or more attention modules are placed after the clipping layer.
[0212] In some embodiments, one or more attention modules are placed before the non-linear activation layer. Alternatively, one or more attention modules are placed after the non-linear activation layer.
[0213] In some embodiments, one or more attention modules are placed before the depthwise separable convolutional layer. Alternatively, one or more attention modules are placed after the depthwise separable convolutional layer.
[0214] In some embodiments, an indication of whether and / or how to determine whether to apply a neural network-based image compression network to a video unit is indicated at one of the following: sequence level, group of picture level, picture level, slice level, or slice group level. In some embodiments, an indication of whether and / or how to determine whether to apply a neural network-based image compression network to a video unit is indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header. In some embodiments, an indication of whether and / or how to determine whether to apply a neural network-based image compression network to a video unit is indicated in one of the following: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a codec tree block (CTB), or a codec tree unit (CTU).
[0215] In some embodiments, method 2000 further includes determining whether and / or how to apply a neural network-based image compression network to the video unit based on coded information of the video unit, the coded information including at least one of: block size, color format, single-tree partitioning and / or dual-tree partitioning, color component, slice type, or picture type. In some embodiments, the video unit is applied using a codec that requires chroma fusion.
[0216] In some embodiments, the SE is binarized as one of a sign, a fixed-length codec, an EG(x) codec, a unary codec, a truncated unary codec, or a truncated binary codec. In some embodiments, the SE is signed or unsigned. In some embodiments, the SE is encoded using at least one context model. Alternatively, the SE is bypassed. In some embodiments, the SE is signaled conditionally. In some embodiments, the SE is signaled only if and only if the corresponding function applies. Alternatively, the SE is signaled only if and only if the dimensions of the video unit meet a condition.
[0217] In some embodiments, SE is indicated at one of the following: sequence level, group of pictures level, picture level, slice level, or slice group level. In some embodiments, SE is indicated at one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), codec tree block (CTB), or codec tree unit (CTU).
[0218] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided, storing a bitstream of a video, the bitstream of the video being generated by a method performed by an apparatus for video processing. The method includes: determining to apply a neural network-based image compression network to a video unit of the video, wherein the neural network-based image compression network includes a synthetic transformation module, the synthetic transformation module including at least one of the following: one or more upscaling layers and one or more attention modules; and generating a bitstream based on the neural network-based image compression network.
[0219] According to yet other embodiments of the present disclosure, a method for storing a bitstream of a video is provided. The method includes: determining to apply a neural network-based image compression network to a video unit of the video, wherein the neural network-based image compression network includes a synthetic transformation module, and the synthetic transformation module includes at least one of the following: one or more upscaling layers and one or more attention modules; generating a bitstream based on the neural network-based image compression network; and storing the bitstream in a non-transitory computer-readable medium.
[0220] Figure 21 An example of a convolution-based attention block is shown in Figure 2. The convolution-based attention block receives a tensor input of size [C, h, w] and performs Figure 21 The sequence of steps depicted in , outputs a tensor output of the same size. Specifically, the process can be divided into two branches, operating at different spatial resolutions. The first branch may include two residual blocks. The second branch may start with a downsampling convolution with stride = 2, followed by two residual blocks and a transposed convolution with stride = 2. The second branch may end with a sigmoid. The two branches may be combined in an element-wise multiplication ⊙. The resulting tensor may then be multiplied by a parameter α and added to the input tensor. When α = 0, all operations in CAB are essentially bypassed.
[0221] The embodiments of the present disclosure may be described according to the following items, the features of which may be combined in any reasonable way.
[0222] Item 1. A method for video processing, comprising: for conversion between a video unit of a video and a bitstream of the video, determining to apply a neural network-based image compression network to the video unit, wherein the neural network-based image compression network includes a synthetic transformation module, the synthetic transformation module including at least one of the following: one or more upscaling layers and one or more attention modules; and performing conversion according to the neural network-based image compression network.
[0223] Item 2. A method according to Item 1, wherein the synthetic transform module comprises only one or more upscaling layers, and one or more attention modules are excluded from the synthetic transform module.
[0224] Item 3. The method of Item 1, wherein one or more attention modules are placed at step N of the synthetic transformation module, where N is an integer.
[0225] Item 4. The method of Item 3, wherein N is equal to 2 and one or more attention modules are placed after the second deconvolution layer.
[0226] Item 5. The method of Item 3, wherein N is equal to 3 and one or more attention modules are placed after the third deconvolution layer.
[0227] Item 6. The method of Item 3, wherein N is equal to 4 and one or more attention modules are placed after the fourth deconvolution layer.
[0228] Item 7. The method of Item 3, wherein N is equal to 0 and one or more attention modules are placed before the first deconvolution layer.
[0229] Item 8. The method of item 7, wherein the input is processed by one or more attention modules and then the input is processed by one or more deconvolution layers.
[0230] Item 9. The method of Item 3, wherein N is equal to 1 and one or more attention modules are placed after the first deconvolution layer.
[0231] Item 10. The method of Item 9, wherein the input is processed by a deconvolution layer and then the input is processed by one or more attention modules.
[0232] Item 11. The method of Item 3, wherein one or more attention modules are placed before the cropping layer, or wherein one or more attention modules are placed after the cropping layer.
[0233] Item 12. The method of Item 3, wherein the one or more attention modules are placed before the non-linear activation layer, or wherein the one or more attention modules are placed after the non-linear activation layer.
[0234] Item 13. The method of Item 1, wherein whether one or more modules are excluded from the synthetic transformation module is based on currently available computing resources.
[0235] Item 14. The method of Item 1, wherein where the one or more modules are placed in the composite transform module is based on currently available computing resources.
[0236] Item 15. A method according to Item 1, wherein the multiple types of attention modules are provided with different complexities, or wherein the multiple types of attention modules are provided with the same complexity.
[0237] Item 16. A method according to Item 15, wherein a manner of combining multiple types of attention modules with a synthesis step is indicated.
[0238] Item 17. The method of Item 1, wherein whether the one or more modules relate to a composite transform module is indicated, and / or wherein where in the composite transform module the one or more modules are located is indicated.
[0239] Clause 18. The method of clause 1, wherein a flag is used to indicate whether one or more modules relate to a synthetic transform module.
[0240] Item 19. The method of Item 1, wherein a syntax element is used to indicate the number of attention modules involved in the synthesis transformation module.
[0241] Item 20. The method of Item 1, wherein a set of attention modules is provided and use of the set of attention modules is determined based on expected complexity or performance.
[0242] Item 21. The method of Item 20, wherein a flag is signaled to indicate a priority of complexity, and wherein the lightweight attention module is enabled if low complexity is desired.
[0243] Item 22. The method of Item 20, wherein a flag is signaled to indicate the priority of reconstruction quality, and wherein a complex attention module is used if high performance is desired.
[0244] Item 23. The method of Item 20, wherein a flag is included in the bitstream to indicate whether the first attention module is used or the second attention module is used.
[0245] Item 24. The method of Item 1, wherein a plurality of flags are used to indicate whether the Nth synthesis step includes an attention module, where N is an integer.
[0246] Item 25. The method of Item 1, wherein the lower scaling layer and the upper scaling layer involve a synthesis transform module.
[0247] Item 26. The method of Item 25, wherein at least one of a depthwise separable deconvolution layer or a point-wise deconvolution layer is involved as part of the downscaling layer and the upscaling layer.
[0248] Item 27. The method of Item 25 or 26, wherein the downscaling layer and the upscaling layer are operated in a channel-wise manner, and the channel-wise manner adjusts the number of feature channels.
[0249] Item 28. The method of Item 25 or 26, wherein the lower scaling layer and the upper scaling layer are operated in a spatial domain level.
[0250] Item 29. The method of Item 28, wherein the lower scaling layer is applied at the beginning of the mask stem and the upper scaling layer is applied before the end of the mask stem.
[0251] Item 30. The method of Item 28, wherein the lower scaling layer is applied at the beginning of the main stem and at the beginning of the mask stem, and the upper scaling layer is applied after combining the mask stem with the main stem.
[0252] Item 31. The method of Item 28, wherein the lower scaling layer is applied at the beginning of the main stem and the upper scaling layer is applied before the end of the main stem.
[0253] Item 32. The method of any one of Items 29 to 31, wherein the downscaling layer reduces the spatial ratio of the feature map and the upscaling layer restores the spatial resolution of the feature map.
[0254] Item 33. The method of Item 28, wherein a combination of depthwise separable deconvolution layers and point-wise deconvolution layers acts as upscaling and downscaling layers and changes the spatial resolution of the feature map.
[0255] Item 34. The method of Item 1, wherein the attention module in the one or more attention modules comprises a main backbone, a skip backbone, and a mask backbone.
[0256] Item 35. The method of Item 34, wherein the main backbone, the skip backbone, and the mask backbone receive consistent input.
[0257] Item 36. The method of Item 34, wherein the input of the mask backbone is an intermediate output of the main backbone.
[0258] Item 37. The method of Item 1, wherein the attention module used in the synthetic transformation module comprises at least two branches.
[0259] Item 38. The method of Item 37, wherein the first branch in the attention module comprises: a downscaling layer, at least one convolutional layer or residual block, and an upscaling layer.
[0260] Item 39. The method of Item 38, wherein the first branch comprises an activation layer, the activation layer being applied after the upscaling layer.
[0261] Item 40. The method of Item 39, wherein the activation layer comprises at least one of: a sigmoid layer, a hyperbolic tangent function, or a ReLU layer.
[0262] Item 41. The method of Item 38, wherein the upscaling layer increases the size or channel dimension of the input.
[0263] Item 42. The method of Item 41, wherein the upscaling layer increases the width and height of the input by a factor of 2.
[0264] Item 43. The method of Item 41, wherein the upscaling layer increases the number of channels of the input by a factor of 2.
[0265] Item 44. The method of Item 38, wherein the downscaling layer reduces the size or channel dimension of the input.
[0266] Item 45. The method of Item 44, wherein the downscaling layer reduces the width and height of the input by a factor of 2.
[0267] Item 46. The method of Item 44, wherein the downscaling layer reduces the number of channels of the input by a factor of 2.
[0268] Item 47. The method of Item 38, wherein the number of samples of the input is reduced by a lower scaling layer and increased by an upper scaling layer.
[0269] Item 48. The method of Item 38, wherein all operations performed up to the upper scaling layer are performed using a reduced number of input samples.
[0270] Item 49. The method of Item 38, wherein the first branch comprises a depthwise separable convolution layer, the depthwise separable convolution layer being applied after the upscaling layer.
[0271] Item 50. The method of Item 37, wherein the second branch in the attention module comprises: a depthwise separable convolutional layer or a residual block.
[0272] Item 51. The method of Item 37, wherein the second branch in the attention module comprises an identity branch, and the input of the second branch is output without modification.
[0273] Item 52. The method of Item 37, wherein the output of the first branch and the output of the second branch are multiplied with each other.
[0274] Item 53. The method of Item 52, wherein the multiplication result of the first branch and the second branch is added to the output of the third branch in the attention module.
[0275] Item 54. The method of Item 37, wherein the attention module comprises three branches.
[0276] Item 55. The method of Item 54, wherein a first branch of the three branches comprises: a first downscaling layer, one or more residual blocks, an upscaling layer, a convolutional layer, and an activation layer in the following order.
[0277] Item 56. The method of Item 55, wherein a depthwise separable convolutional layer is used in at least one of: a first downscaling layer, one or more residual blocks, and an upscaling layer.
[0278] Item 57. The method of Item 54, wherein a second branch of the three branches comprises one or more residual blocks (RBs).
[0279] Item 58. The method of Item 54, wherein the first of the three branches comprises: a first downscaling layer, three residual blocks, an upscaling layer, a convolutional layer, and an activation layer in the following order.
[0280] Item 59. The method of Item 58, wherein a depthwise separable convolutional layer is used in at least one of: a first downscaling layer, three residual blocks, and an upscaling layer.
[0281] Item 60. The method of Item 54, wherein the second branch of the three branches comprises two residual blocks (RBs), and the depthwise separable convolutional layer involves the two RBs.
[0282] Item 61. The method of Item 54, wherein the second of the three branches comprises an identity branch.
[0283] Item 62. The method according to any one of Items 54 to 61, wherein the output of a first branch and the output of a second branch of the three branches are multiplied with each other.
[0284] Item 63. The method of Item 62, wherein the multiplication result of the first branch and the second branch is added to the input of the third branch, the input of the third branch being the identity branch.
[0285] Item 64. The method of Item 37, wherein the attention module comprises two branches.
[0286] Item 65. The method of Item 37 or 64, wherein the first branch of the two branches comprises: a first downscaling layer, one or more residual blocks, an upscaling layer, a convolutional layer, and an activation layer in the following order.
[0287] Item 66. The method of Item 65, wherein a depthwise separable convolutional layer is used in at least one of: a first downscaling layer, one or more residual blocks, and an upscaling layer.
[0288] Item 67. The method of Item 65, wherein depthwise separable convolutional layers are used in at least one of: a first downscaling layer and an upscaling layer.
[0289] Item 68. The method of Item 65, wherein no residual block is applied before the first downscaling layer and after the upscaling layer.
[0290] Item 69. The method of Item 65, wherein the second of the two branches comprises an identity branch.
[0291] Item 70. A method according to Item 65, wherein the output of the first branch and the output of the second branch of the two branches are multiplied with each other, and the multiplication is the output of the attention module.
[0292] Item 71. The method of Item 64, wherein the first of the two branches is an identity branch.
[0293] Item 72. The method of Item 64, wherein a second branch of the two branches comprises one or more residual blocks, and the second branch is split into two branches after the one or more residual blocks, the two branches comprising a third branch and a fourth branch.
[0294] Item 73. The method of Item 72, wherein the third branch is an identity branch.
[0295] Item 74. The method of Item 72, wherein the fourth branch comprises one or more other residual blocks.
[0296] Item 75. The method of Item 74, wherein the fourth branch further comprises an activation layer.
[0297] Item 76. The method of Item 72, wherein the output of the third branch and the output of the fourth branch are multiplied with each other to form the output of the second branch.
[0298] Item 77. The method of Item 1, wherein a residual Swin transformer block is used in the synthesis transformer module.
[0299] Item 78. The method of Item 77, wherein at least one of a depthwise separable deconvolution layer or a point-wise deconvolution layer is involved as part of a residual Swin transformer block.
[0300] Item 79. The method of Item 77, wherein a residual Swin transformer block and one or more attention modules are used in the synthetic transformer module.
[0301] Item 80. The method of Item 77, wherein a residual Swin transformer block is used in a synthetic transform module and one or more attention modules are removed from the synthetic transform module.
[0302] Item 81. The method of Item 77, wherein the residual Swin transformer block is part of one or more attention modules.
[0303] Item 82. The method of Item 81, wherein only Swin transformer layers are included in one or more attention modules.
[0304] Item 83. The method of Item 81, wherein the residual Swin transformer block acts as a new branch and the associated output is combined with the output of one or more attention modules by one of the following: addition, subtraction, concatenation, fusion, multiplication, or nonlinear activation.
[0305] Item 84. The method of Item 77, wherein the residual Swin transformer block is a residual block having a Swin transformer layer and a depthwise separable convolutional layer.
[0306] Item 85. The method of Item 77, wherein the residual Swin transformer block is a residual block having a Swin transformer layer and a depthwise separable deconvolution layer.
[0307] Item 86. The method of Item 77, wherein the residual Swin transformer block is a modified residual block, wherein one branch comprises a Swin transformer layer, and one branch comprises at least one of: a depthwise separable convolutional layer or a depthwise separable deconvolutional layer.
[0308] Item 87. The method of Item 77, wherein the Swin transformer layer is configured with a head size, a depth size, a window size, and a tile size.
[0309] Item 88. The method of Item 87, wherein the window size is equal to 1, the path size is equal to 1, the depth size is equal to 2, and the head size is equal to 2, or wherein the window size is equal to 1, the path size is equal to 1, the depth size is equal to 4, and the head size is equal to 4, or wherein the window size is equal to 4 and the tile size is equal to 2.
[0310] Item 89. The method of Item 87, wherein the head dimension is the same as the depth dimension, or wherein the head dimension is equal to 2 and the depth dimension is equal to 2.
[0311] Item 90. The method of Item 87, wherein the setting of the Swin transformer layer depends on the position of the residual Swin transformer block.
[0312] Item 91. The method of Item 90, wherein one residual Swin transformer block is placed at the Nth step of the synthesis transform module and the other residual Swin transformer block is placed at the Mth step of the synthesis transform module, where N and M are integers.
[0313] Item 92. The method of Item 91, wherein the settings of the residual Swin transformer block are the same as the settings of another residual Swin transformer block, or wherein the settings of the residual Swin transformer block are different from the settings of another residual Swin transformer block.
[0314] Item 93. The method of Item 90, wherein multiple residual Swin transformer blocks are placed within each step of the synthetic transform module.
[0315] Item 94. The method of Item 93, wherein the configuration of the Swin transformer layers is the same, or wherein the configuration of the Swin transformer layers is different.
[0316] Item 95. The method of Item 77, wherein the Swin transformer layer comprises a multi-head self-attention layer, a multilayer perceptron, and layer normalization.
[0317] Item 96. The method of Item 1, wherein at least one of a depthwise separable convolutional layer or a pixel-wise convolutional layer is included in the synthesis transform module.
[0318] Item 97. The method of Item 96, wherein all existing convolutional layers or deconvolutional layers are replaced by depthwise separable convolutional layers.
[0319] Item 98. The method of Item 97, wherein the number of groups of depthwise separable convolutional layers is consistent with the depth of the features.
[0320] Item 99. The method of Item 97, wherein the number of groups of depthwise separable convolutional layers is set to a predetermined number.
[0321] Item 100. The method of Item 99, wherein the predetermined number is equal to 1, or wherein the predetermined number is a power of 2.
[0322] Item 101. The method of Item 97, wherein the kernel size of the depthwise separable convolutional layer is K×P, where K and P are integers.
[0323] Item 102. The method of Item 101, wherein K and P are both equal to 1, or wherein K and P are both equal to 2n+1, where n is a positive integer, or wherein the values of K and P are inconsistent.
[0324] Item 103. A method according to Item 96, wherein depthwise separable convolution layers, pointwise convolution layers and regular convolution layers are combined and used in the Nth step of the synthesis transformation module, and / or wherein depthwise separable deconvolution layers, pointwise deconvolution layers and regular deconvolution layers are combined and used in the Nth step of the synthesis transformation module, and wherein N is an integer.
[0325] Item 104. A method according to Item 103, wherein depthwise separable convolution layers and pointwise convolution layers are sequentially combined and the depthwise separable convolution layers are used first, and / or wherein depthwise separable deconvolution layers and pointwise deconvolution layers are sequentially combined and the depthwise separable convolution layers are used first.
[0326] Item 105. A method according to Item 103, wherein depthwise separable convolution layers and pointwise convolution layers are sequentially combined and the pointwise convolution layers are used first, and / or wherein depthwise separable deconvolution layers and pointwise deconvolution layers are sequentially combined and the pointwise convolution layers are used first.
[0327] Item 106. The method of Item 103, wherein depthwise separable convolutional layers and pointwise convolutional layers are used as separate branches, and / or wherein depthwise separable deconvolutional layers and pointwise deconvolutional layers are used as separate branches.
[0328] Item 107. A method according to Item 103, wherein the order of depthwise separable convolution layers, pointwise convolution layers and regular convolution layers is arbitrarily changed, and / or wherein the order of depthwise separable deconvolution layers, pointwise deconvolution layers and regular deconvolution layers is arbitrarily changed.
[0329] Item 108. The method of Item 1, wherein the one or more attention modules comprise at least one of a depthwise separable convolutional layer or a depthwise separable deconvolutional layer.
[0330] Item 109. A method according to Item 108, wherein whether a depthwise separable convolutional layer is used in one or more attention modules is determined based on available computational resources, and / or wherein where the depthwise separable convolutional layer is placed is determined based on available computational resources.
[0331] Item 110. The method of Item 108, wherein only the upscaling layer involves the synthesis transform module and one or more attention modules are excluded from the synthesis transform module.
[0332] Item 111. The method of Item 110, wherein part or all of the upscaling layer is implemented as a depthwise separable convolutional layer.
[0333] Item 112. The method of Item 108, wherein one or more attention modules are placed at step N of the synthetic transformation module, where N is an integer.
[0334] Item 113. The method of Item 112, wherein N is equal to 0 and one or more attention modules are placed before the first deconvolution layer.
[0335] Item 114. The method of Item 113, wherein the input is processed by one or more attention modules and then the input is processed by one or more deconvolution layers.
[0336] Item 115. The method of Item 112, wherein N is equal to 1 and one or more attention modules are placed after the first deconvolution layer.
[0337] Item 116. The method of Item 115, wherein the input is processed by a deconvolution layer and then the input is processed by one or more attention modules.
[0338] Item 117. The method of Item 112, wherein N is equal to 2 and the one or more attention modules are placed after the second deconvolution layer.
[0339] Item 118. The method of Item 112, wherein N is equal to 3 and the one or more attention modules are placed after the third deconvolution layer.
[0340] Item 119. The method of Item 112, wherein N is equal to 4 and the one or more attention modules are placed after the fourth deconvolution layer.
[0341] Item 120. The method of any one of Items 113 to 116, wherein the deconvolution layer comprises one or more depthwise separable deconvolution layers or one or more point-wise deconvolution layers.
[0342] Item 121. The method of Item 112, wherein one or more attention modules are placed before the cropping layer, or wherein one or more attention modules are placed after the cropping layer.
[0343] Item 122. The method of Item 112, wherein the one or more attention modules are placed before the non-linear activation layer, or wherein the one or more attention modules are placed after the non-linear activation layer.
[0344] Item 123. The method of Item 112, wherein the one or more attention modules are placed before the depthwise separable convolutional layer, or wherein the one or more attention modules are placed after the depthwise separable convolutional layer.
[0345] Item 124. A method according to any of Items 1 to 123, wherein an indication of whether and / or how to determine to apply a neural network-based image compression network to a video unit is indicated at one of: a sequence level, a group of pictures level, a picture level, a slice level, or a slice group level.
[0346] Item 125. A method according to any of Items 1-123, wherein an indication of whether and / or how to determine whether to apply a neural network-based image compression network to a video unit is indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.
[0347] Item 126. A method according to any of items 1-123, wherein an indication of whether and / or how to determine whether to apply a neural network-based image compression network to a video unit is indicated in one of the following items: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a codec tree block (CTB) or a codec tree unit (CTU).
[0348] Item 127. The method according to any one of Items 1 to 123 further includes: determining whether and / or how to determine whether to apply a neural network-based image compression network to the video unit based on encoded and decoded information of the video unit, the encoded and decoded information including at least one of the following: block size, color format, single tree partitioning and / or double tree partitioning, color component, slice type or picture type.
[0349] Item 128. The method of any one of Items 1 to 127, wherein the video unit is applied using a codec that requires chroma fusion.
[0350] Item 129. The method of any one of items 1 to 128, wherein the SE is binarized as one of a flag, a fixed length codec, an EG(x) codec, a unary codec, a truncated unary codec, or a truncated binary codec.
[0351] Item 130. The method of Item 129, wherein SE is signed or unsigned.
[0352] Item 131. A method according to any one of items 1 to 130, wherein the SE is encoded using at least one context model, or wherein the SE is bypass encoded.
[0353] Item 132. The method of any one of Items 1 to 131, wherein the SE is signaled in a conditional manner.
[0354] Item 133. The method of Item 132, wherein the SE is signaled if and only if the corresponding function applies, or wherein the SE is signaled if and only if the dimensions of the video unit satisfy a condition.
[0355] Item 134. The method of any of Items 1 to 133, wherein SE is indicated at one of: sequence level, group of pictures level, picture level, slice level, or slice group level.
[0356] Item 135. A method according to any one of items 1 to 133, wherein the SE is indicated at one of the following items: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a codec tree block (CTB) or a codec tree unit (CTU).
[0357] Item 136. The method of any one of Items 1 to 135, wherein converting comprises encoding the video unit into a bitstream.
[0358] Item 137. The method of any one of Items 1 to 135, wherein converting comprises decoding the video unit from a bitstream.
[0359] Item 138. An apparatus for video processing, comprising a processor and non-transitory memory having instructions, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of Items 1 to 137.
[0360] Item 139. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of Items 1 to 137.
[0361] Item 140. A non-transitory computer-readable recording medium storing a bitstream of a video, the bitstream of the video being generated by a method performed by an apparatus for video processing, wherein the method comprises: determining to apply a neural network-based image compression network to a video unit of the video, wherein the neural network-based image compression network comprises a synthetic transform module, the synthetic transform module comprising at least one of the following: one or more upscaling layers and one or more attention modules; and generating a bitstream based on the neural network-based image compression network.
[0362] Item 141. A method for storing a bitstream of a video, comprising: determining to apply a neural network-based image compression network to a video unit of the video, wherein the neural network-based image compression network includes a synthetic transform module, the synthetic transform module including at least one of the following: one or more upscaling layers and one or more attention modules; generating a bitstream based on the neural network-based image compression network; and storing the bitstream in a non-transitory computer-readable medium. Example device
[0363] Figure 22 A block diagram of a computing device 2100 in which various embodiments of the present disclosure may be implemented is shown. The computing device 2100 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0364] It should be understood that Figure 22 The computing device 2100 shown in FIG. 2 is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the embodiments of the present disclosure.
[0365] like Figure 22 As shown, computing device 2100 comprises a general computing device 2100. Computing device 2100 may include at least one or more processors or processing units 2110, memory 2120, storage unit 2130, one or more communication units 2140, one or more input devices 2150, and one or more output devices 2160.
[0366] In some embodiments, the computing device 2100 can be implemented as any user terminal or server terminal with computing power. The server terminal can be a server, a large computing device, etc. provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 2100 can support any type of interface to the user (such as a "wearable" circuit device, etc.).
[0367] The processing unit 2110 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 2120. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capability of the computing device 2100. The processing unit 2110 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0368] The computing device 2100 typically includes various computer storage media. Such media can be any media accessible by the computing device 2100, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 2120 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM) or flash memory) or any combination thereof. The storage unit 2130 can be any removable or non-removable medium and can include machine-readable media, such as memory, flash drive, disk or other media that can be used to store information and / or data and can be accessed in the computing device 2100.
[0369] The computing device 2100 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Figure 22 Although not shown, a magnetic disk drive for reading from and / or writing to a removable nonvolatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data medium interfaces.
[0370] The communication unit 2140 communicates with another computing device via a communication medium. In addition, the functions of the components in the computing device 2100 can be implemented by a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 2100 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0371] The input device 2150 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, and the like. The output device 2160 may be one or more of various output devices, such as a display, a speaker, a printer, and the like. With the aid of the communication unit 2140, the computing device 2100 may also communicate with one or more external devices (not shown), such as storage devices and display devices, and may also communicate with one or more devices that enable a user to interact with the computing device 2100, or, if desired, any device that enables the computing device 2100 to communicate with one or more other computing devices (e.g., a network card, a modem, and the like). Such communication may be performed via an input / output (I / O) interface (not shown).
[0372] In some embodiments, some or all components of the computing device 2100 may also be arranged in a cloud computing architecture rather than being integrated into a single device. In a cloud computing architecture, components can be provided remotely and work together to implement the functionality described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring the end user to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides an application via a wide area network that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on servers in a remote location. Computing resources in a cloud computing environment can be consolidated or distributed across remote data centers. Cloud computing infrastructure can provide services through shared data centers, although to users, they appear as a single access point. Therefore, cloud computing architecture can be used to provide the components and functionality described herein from a service provider in a remote location. Alternatively, the components and functionality described herein can be provided by a conventional server or installed directly or otherwise on a client device.
[0373] In embodiments of the present disclosure, the computing device 2100 may be used to implement video encoding / decoding. The memory 2120 may include one or more video encoding / decoding modules 2125 having one or more program instructions. These modules are accessible and executable by the processing unit 2110 to perform the functions of the various embodiments described herein.
[0374] In an example embodiment performing video encoding, an input device 2150 may receive video data as input to be encoded 2170. The video data may be processed, for example, by a video codec module 2125 to generate an encoded bitstream. The encoded bitstream may be provided as output 2180 via an output device 2160.
[0375] In an example embodiment performing video decoding, input device 2150 may receive an encoded bitstream as input 2170. The encoded bitstream may be processed, for example, by video codec module 2125 to generate decoded video data. The decoded video data may be provided as output 2180 via output device 2160.
[0376] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present application as defined by the appended claims. Such variations are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for video processing, comprising: For conversion between a video unit of a video and a bitstream of the video, determining to apply a neural network-based image compression network to the video unit, wherein the neural network-based image compression network includes a synthetic transformation module, the synthetic transformation module including at least one of the following: one or more upscaling layers and one or more attention modules; as well as The conversion is performed according to the neural network based image compression network.
2. The method of claim 1, wherein the synthetic transform module includes only the one or more upscaling layers, and the one or more attention modules are excluded from the synthetic transform module.
3. The method of claim 1, wherein the one or more attention modules are placed at the Nth step of the synthetic transformation module, where N is an integer.
4. The method of claim 3, wherein N is equal to 2 and the one or more attention modules are placed after the second deconvolution layer.
5. The method of claim 3, wherein N is equal to 3 and the one or more attention modules are placed after the third deconvolution layer.
6. The method of claim 3, wherein N is equal to 4 and the one or more attention modules are placed after the fourth deconvolution layer.
7. The method of claim 3, wherein N is equal to 0 and the one or more attention modules are placed before the first deconvolution layer.
8. The method of claim 7, wherein the input is processed by the one or more attention modules and then the input is processed by one or more deconvolution layers.
9. The method of claim 3, wherein N is equal to 1, and the one or more attention modules are placed after the first deconvolution layer.
10. The method of claim 9, wherein the input is processed by a deconvolution layer and then the input is processed by the one or more attention modules.
11. The method of claim 3, wherein the one or more attention modules are placed before the clipping layer, or The one or more attention modules are placed after the cropping layer.
12. The method of claim 3, wherein the one or more attention modules are placed before a non-linear activation layer, or Wherein the one or more attention modules are placed after the non-linear activation layer.
13. The method of claim 1, wherein whether the one or more modules are excluded from the synthetic transformation module is based on currently available computing resources.
14. The method of claim 1, wherein where the one or more modules are placed in the composite transform module is based on currently available computing resources.
15. The method of claim 1, wherein multiple types of attention modules are provided with different complexities, or The multiple types of attention modules are provided with the same complexity.
16. A method according to claim 15, wherein the manner of combining the multiple types of attention modules with the synthesis step is indicated.
17. The method according to claim 1, wherein whether the one or more modules are involved in the synthetic transformation module is indicated, and / or Where the one or more modules are placed in the composite transformation module is indicated. The method according to claim 1 , wherein a flag is used to indicate whether the one or more modules involve the synthetic transformation module.
19. The method of claim 1, wherein a syntax element is used to indicate the number of the attention modules involved in the synthesis transformation module.
20. The method of claim 1, wherein a set of attention modules is provided and usage of the set of attention modules is determined based on expected complexity or performance.
21. The method of claim 20, wherein a flag is signaled to indicate a priority of complexity, and If low complexity is desired, the lightweight attention module is enabled.
22. The method of claim 20, wherein a flag is signaled to indicate a priority of reconstruction quality, and If high performance is desired, complex attention modules are used.
23. The method of claim 20, wherein a flag is included in the bitstream to indicate whether the first attention module is used or the second attention module is used.
24. The method of claim 1, wherein a plurality of flags are used to indicate whether the Nth synthesis step includes an attention module, where N is an integer.
25. The method of claim 1, wherein a lower scaling layer and an upper scaling layer involve the synthetic transform module.
26. The method of claim 25, wherein at least one of a depthwise separable deconvolution layer or a point-wise deconvolution layer is involved as part of the downscaling layer and the upscaling layer.
27. The method according to claim 25 or 26, wherein the down-scaling layer and the up-scaling layer are operated in a channel level, and the channel level adjusts the number of feature channels.
28. The method according to claim 25 or 26, wherein the lower scaling layer and the upper scaling layer are operated in a spatial domain level.
29. The method of claim 28, wherein the lower scaling layer is applied at the beginning of a mask stem and the upper scaling layer is applied before the end of the mask stem.
30. The method of claim 28, wherein the lower scaling layer is applied at the beginning of a main stem and at the beginning of a mask stem, and the upper scaling layer is applied after combining the mask stem with the main stem.
31. The method of claim 28, wherein the lower scaling layer is applied at the beginning of a main stem and the upper scaling layer is applied before the end of the main stem.
32. The method according to any one of claims 29 to 31, wherein the down-scaling layer reduces the spatial ratio of the feature map, and the up-scaling layer restores the spatial resolution of the feature map.
33. The method according to claim 28, wherein the combination of the depthwise separable deconvolution layer and the point-wise deconvolution layer acts as the upscaling layer and the downscaling layer and changes the spatial resolution of the feature map.
34. The method of claim 1, wherein the attention modules in the one or more attention modules include a main backbone, a skip backbone, and a mask backbone.
35. The method of claim 34, wherein the main stem, the skip stem, and the mask stem receive consistent input.
36. The method of claim 34, wherein the input of the masking backbone is an intermediate output of the main backbone.
37. The method according to claim 1, wherein the attention module used in the synthetic transformation module includes at least two branches.
38. The method of claim 37, wherein the first branch in the attention module comprises: A downscaling layer, at least one convolutional layer or residual block, and an upscaling layer.
39. The method of claim 38, wherein the first branch comprises an activation layer, the activation layer being applied after the upper scaling layer.
40. The method of claim 39, wherein the activation layer comprises at least one of the following: a sigmoid layer, a hyperbolic tangent function, or a ReLU layer.
41. The method of claim 38, wherein the upscaling layer increases the size or channel dimension of the input.
42. The method of claim 41, wherein the upper scaling layer increases the width and height of the input by a factor of 2. The method of claim 41 , wherein the upscaling layer increases the number of channels of the input by a factor of 2.
44. The method of claim 38, wherein the downscaling layer reduces the size or channel dimension of the input. The method of claim 44 , wherein the lower scaling layer reduces the width and height of the input by a factor of 2.
46. The method of claim 44, wherein the downscaling layer reduces the number of channels of the input by a factor of 2.
47. The method of claim 38, wherein the number of samples of an input is reduced by the lower scaling layer and increased by the upper scaling layer.
48. The method of claim 38, wherein all operations performed up to the upper scaling layer are performed using a reduced number of input samples.
49. The method of claim 38, wherein the first branch comprises a depthwise separable convolution layer, the depthwise separable convolution layer being applied after the upscaling layer.
50. The method of claim 37, wherein the second branch in the attention module comprises: Depthwise separable convolutional layers or residual blocks.
51. The method of claim 37, wherein the second branch in the attention module comprises an identity branch, and the input of the second branch is output without modification.
52. The method of claim 37, wherein the output of the first branch and the output of the second branch are multiplied with each other.
53. The method of claim 52, wherein the multiplication result of the first branch and the second branch is added to the output of a third branch in the attention module.
54. The method of claim 37, wherein the attention module comprises three branches.
55. The method of claim 54, wherein a first of the three branches comprises: First downscaling layer, one or more residual blocks, upscaling layer, convolutional layer, and activation layer in the following order.
56. The method of claim 55, wherein depthwise separable convolutional layers are used in at least one of: the first downscaling layer, the one or more residual blocks, and the upscaling layer.
57. The method of claim 54, wherein a second branch of the three branches comprises one or more residual blocks (RBs).
58. The method of claim 54, wherein a first of the three branches comprises: First downscaling layer, three residual blocks, upscaling layer, convolutional layer, and activation layer in the following order.
59. The method of claim 58, wherein a depthwise separable convolutional layer is used in at least one of: the first downscaling layer, the three residual blocks, and the upscaling layer.
60. The method of claim 54, wherein a second branch of the three branches comprises two residual blocks (RBs), and a depthwise separable convolutional layer involves the two RBs.
61. The method of claim 54, wherein a second of the three branches comprises an identity branch.
62. The method according to any one of claims 54 to 61, wherein the output of a first branch and the output of a second branch of the three branches are multiplied with each other.
63. The method of claim 62, wherein a result of the multiplication of the first branch and the second branch is added to an input of a third branch, the input of the third branch being an identity branch.
64. The method of claim 37, wherein the attention module comprises two branches.
65. The method of claim 37 or 64, wherein a first of the two branches comprises: First downscaling layer, one or more residual blocks, upscaling layer, convolutional layer, and activation layer in the following order.
66. The method of claim 65, wherein depthwise separable convolutional layers are used in at least one of: the first downscaling layer, the one or more residual blocks, and the upscaling layer.
67. The method of claim 65, wherein depthwise separable convolutional layers are used in at least one of: the first downscaling layer and the upscaling layer.
68. The method of claim 65, wherein no residual block is applied before the first lower scaling layer and after the upper scaling layer.
69. The method of claim 65, wherein the second of the two branches comprises an identity branch.
70. The method according to claim 65, wherein the output of the first branch and the output of the second branch of the two branches are multiplied with each other, and the multiplication is the output of the attention module.
71. The method of claim 64, wherein a first of the two branches is an identity branch.
72. The method of claim 64, wherein a second branch of the two branches comprises one or more residual blocks, and the second branch is divided into two branches after the one or more residual blocks, the two branches comprising a third branch and a fourth branch.
73. The method of claim 72, wherein the third branch is an identity branch.
74. The method of claim 72, wherein the fourth branch comprises one or more other residual blocks.
75. The method of claim 74, wherein the fourth branch further comprises an activation layer.
76. The method of claim 72, wherein the output of the third branch and the output of the fourth branch are multiplied with each other to form the output of the second branch.
77. The method of claim 1, wherein a residual Swin transformer block is used in the synthetic transform module.
78. The method of claim 77, wherein at least one of a depthwise separable deconvolution layer or a point-wise deconvolution layer is involved as part of the residual Swin transformer block.
79. The method of claim 77, wherein the residual Swin transformer block and the one or more attention modules are used in the synthetic transform module.
80. The method of claim 77, wherein the residual Swin transformer block is used in the synthetic transform module and the one or more attention modules are removed from the synthetic transform module.
81. The method of claim 77, wherein the residual Swin transformer block is part of the one or more attention modules.
82. The method of claim 81 , wherein only Swin transformer layers are included in the one or more attention modules.
83. A method according to claim 81, wherein the residual Swin transformer block acts as a new branch and the associated output is combined with the output of the one or more attention modules by one of the following: addition, subtraction, concatenation, fusion, multiplication or non-linear activation.
84. The method of claim 77, wherein the residual Swin transformer block is a residual block having a Swin transformer layer and a depthwise separable convolutional layer.
85. The method of claim 77, wherein the residual Swin transformer block is a residual block having a Swin transformer layer and a depthwise separable deconvolution layer.
86. The method of claim 77, wherein the residual Swin transformer block is a modified residual block wherein one branch comprises a Swin transformer layer, and one branch comprises at least one of: a depthwise separable convolutional layer or a depthwise separable deconvolutional layer.
87. The method of claim 77, wherein a Swin transformer layer is configured with a head size, a depth size, a window size, and a tile size.
88. The method of claim 87, wherein the window size is equal to 1, the path size is equal to 1, the depth size is equal to 2, and the head size is equal to 2, or wherein the window size is equal to 1, the path size is equal to 1, the depth size is equal to 4, and the head size is equal to 4, or Wherein the window size is equal to 4, and the patch size is equal to 2.
89. The method of claim 87, wherein the head dimension is the same as the depth dimension, or Wherein the head dimension is equal to 2, and the depth dimension is equal to 2.
90. The method of claim 87, wherein the setting of the Swin transformer layer depends on the location of the residual Swin transformer block.
91. The method of claim 90, wherein one residual Swin transformer block is placed at the Nth step of the synthetic transform module, and another residual Swin transformer block is placed at the Mth step of the synthetic transform module, where N and M are integers.
92. The method of claim 91 , wherein the settings of the residual Swin transformer block are the same as the settings of the other residual Swin transformer block, or The setting of the residual Swin transformer block is different from the setting of the another residual Swin transformer block.
93. The method of claim 90, wherein a plurality of residual Swin transformer blocks are placed within each step of the synthetic transform module.
94. The method of claim 93, wherein the Swin transformer layers are arranged identically, or The configuration of the Swin converter layer is different.
95. The method of claim 77, wherein the Swin transformer layer comprises a multi-head self-attention layer, a multi-layer perceptron, and layer normalization.
96. The method of claim 1, wherein at least one of a depthwise separable convolutional layer or a pixel-level convolutional layer is included in the synthetic transform module.
97. The method of claim 96, wherein all existing convolutional layers or deconvolutional layers are replaced by the depthwise separable convolutional layers.
98. The method of claim 97, wherein the number of groups of separable convolutional layers in the depth direction is consistent with the depth of the feature.
99. The method of claim 97, wherein the number of groups of the depth-wise separable convolutional layers is set to a predetermined number.
100. The method of claim 99, wherein the predetermined number is equal to 1, or Wherein the predetermined number is a power of 2.
101. The method of claim 97, wherein the kernel size of the depthwise separable convolutional layer is K×P, where K and P are integers.
102. The method of claim 101, wherein K and P are both equal to 1, or where K and P are both equal to 2n+1, where n is a positive integer, or The values of K and P are inconsistent.
103. The method of claim 96, wherein the depthwise separable convolutional layers, point-wise convolutional layers, and regular convolutional layers are combined and used in the Nth step of the synthetic transform module, and / or wherein depthwise separable deconvolution layers, point-wise deconvolution layers, and regular deconvolution layers are combined and used in the Nth step of the synthetic transform module, and Where N is an integer.
104. The method of claim 103, wherein a depthwise separable convolution layer and a pointwise convolution layer are sequentially combined, and the depthwise separable convolution layer is used first, and / or The depthwise separable deconvolution layer and the pointwise deconvolution layer are sequentially combined, and the depthwise separable convolution layer is used first.
105. The method of claim 103, wherein a depthwise separable convolution layer and a point-wise convolution layer are sequentially combined, and the point-wise convolution layer is used first, and / or The depthwise separable deconvolution layer and the point-by-point deconvolution layer are sequentially combined, and the point-by-point deconvolution layer is used first.
106. The method of claim 103, wherein depthwise separable convolutional layers and pointwise convolutional layers are used as separate branches, and / or The depthwise separable deconvolution layer and the point-by-point deconvolution layer are used as separate branches.
107. The method of claim 103, wherein the order of the depthwise separable convolutional layers, the point-wise convolutional layers, and the regular convolutional layers is arbitrarily changed, and / or The order of depthwise separable deconvolution layers, point-wise deconvolution layers, and regular deconvolution layers is arbitrarily changed.
108. The method of claim 1, wherein the one or more attention modules comprise at least one of a depthwise separable convolutional layer or a depthwise separable deconvolutional layer.
109. The method of claim 108, wherein whether the depthwise separable convolutional layer is used in the one or more attention modules is determined based on available computational resources, and / or Where the depthwise separable convolutional layer is placed is determined based on the available computing resources.
110. The method of claim 108, wherein only an upper scaling layer is involved in the synthetic transform module, and the one or more attention modules are excluded from the synthetic transform module.
111. The method of claim 110, wherein part or all of the upscaling layer is implemented as a depthwise separable convolutional layer.
112. The method of claim 108, wherein the one or more attention modules are placed at the Nth step of the synthetic transformation module, where N is an integer.
113. The method of claim 112, wherein N is equal to 0 and the one or more attention modules are placed before the first deconvolution layer.
114. The method of claim 113, wherein the input is processed by the one or more attention modules and then the input is processed by one or more deconvolution layers.
115. The method of claim 112, wherein N is equal to 1 and the one or more attention modules are placed after the first deconvolution layer.
116. The method of claim 115, wherein the input is processed by a deconvolution layer and then the input is processed by the one or more attention modules.
117. The method of claim 112, wherein N is equal to 2 and the one or more attention modules are placed after the second deconvolution layer.
118. The method of claim 112, wherein N is equal to 3 and the one or more attention modules are placed after the third deconvolution layer.
119. The method of claim 112, wherein N is equal to 4 and the one or more attention modules are placed after the fourth deconvolution layer.
120. The method according to any one of claims 113 to 116, wherein the deconvolution layer comprises one or more depthwise separable deconvolution layers or one or more point-wise deconvolution layers.
121. The method of claim 112, wherein the one or more attention modules are placed before the clipping layer, or The one or more attention modules are placed after the cropping layer.
122. The method of claim 112, wherein the one or more attention modules are placed before a non-linear activation layer, or Wherein the one or more attention modules are placed after the non-linear activation layer.
123. The method of claim 112, wherein the one or more attention modules are placed before a depthwise separable convolutional layer, or The one or more attention modules are placed after the depthwise separable convolutional layer.
124. The method of any one of claims 1 to 123, wherein an indication of whether and / or how to determine to apply the neural network-based image compression network to the video unit is indicated at one of: Sequence level, Picture group level, Picture level, Stripe level, or Film group level.
125. The method of any one of claims 1-123, wherein an indication of whether and / or how to determine to apply the neural network-based image compression network to the video unit is indicated in one of: Sequence header, Picture header, Sequence Parameter Set (SPS), Video Parameter Set (VPS), Dependent Parameter Set (DPS), Decoding Capability Information (DCI), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Strip header, or Film group header.
126. The method of any one of claims 1-123, wherein an indication of whether and / or how to determine to apply the neural network-based image compression network to the video unit is indicated in one of: Prediction Block (PB), Transform Block (TB), Codec Block (CB), Prediction Unit (PU), Transformation Unit (TU), Codec Unit (CU), Codec Tree Block (CTB) or Codec Tree Unit (CTU).
127. The method of any one of claims 1 to 123, further comprising: Determining whether and / or how to apply the neural network-based image compression network to the video unit based on coded information of the video unit, the coded information comprising at least one of the following: Block size, Color format, Single-tree partitioning and / or dual-tree partitioning, Color component, Strip type or Image type.
128. The method of any one of claims 1 to 127, wherein the video unit is applied using a codec that requires chroma fusion.
129. The method according to any one of claims 1 to 128, wherein the SE is binarized as one of a flag, a fixed length codec, an EG(x) codec, a unary codec, a truncated unary codec, or a truncated binary codec.
130. The method of claim 129, wherein the SE is signed or unsigned.
131. The method according to any one of claims 1 to 130, wherein the SE is encoded or decoded using at least one context model, or The SE is bypassed for coding.
132. The method of any one of claims 1 to 131, wherein the SE is signaled in a conditional manner.
133. A method according to claim 132, wherein the SE is signalled if and only if the corresponding function applies, or The SE is transmitted via a signal if and only if the dimension of the video unit meets a condition.
134. The method of any one of claims 1 to 133, wherein the SE is indicated at one of: Sequence level, Picture group level, Picture level, Stripe level or Film group level.
135. The method of any one of claims 1 to 133, wherein the SE is indicated at one of: Prediction Block (PB), Transform Block (TB), Codec Block (CB), Prediction Unit (PU), Transformation Unit (TU), Codec Unit (CU), Codec Tree Block (CTB) or Codec Tree Unit (CTU).
136. The method of any one of claims 1 to 135, wherein the converting comprises encoding the video unit into the bitstream.
137. The method of any one of claims 1 to 135, wherein the converting comprises decoding the video unit from the bitstream.
138. An apparatus for video processing, comprising a processor and a non-volatile memory having instructions, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 137.
139. A non-transitory computer-readable storage medium storing instructions for causing a processor to execute the method according to any one of claims 1 to 137.
140. A non-transitory computer-readable recording medium storing a bitstream of a video, wherein the bitstream of the video is generated by a method performed by an apparatus for video processing, wherein the method comprises: determining to apply a neural network-based image compression network to a video unit of the video, wherein the neural network-based image compression network comprises a synthetic transform module, the synthetic transform module comprising at least one of: one or more upscaling layers and one or more attention modules; as well as The bitstream is generated according to the neural network-based image compression network.
141. A method for storing a bitstream of a video, comprising: determining to apply a neural network-based image compression network to a video unit of the video, wherein the neural network-based image compression network comprises a synthetic transform module, the synthetic transform module comprising at least one of: one or more upscaling layers and one or more attention modules; generating the bitstream according to the neural network-based image compression network; as well as The bitstream is stored in a non-transitory computer-readable medium.