Method and device for video processing and medium
By applying a neural network filtering model in the RDO process of video encoding and decoding, processing the video unit and considering the distortion impact caused by the filter, the problem of poor encoding and decoding performance in the prior art is solved, and more efficient encoding and decoding performance is achieved.
Patent Information
- Application Number
- CN202380072321.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-13
- Filing Date
- 2023-10-12
- Publication Date
- 2025-05-30
AI Technical Summary
The existing video encoding and decoding technology fails to fully consider the effect of reducing distortion caused by neural network filters during rate distortion optimization (RDO), resulting in poor encoding and decoding performance.
During the RDO process, by determining whether the video unit applies a neural network filtering model, the model is applied to process the video unit, and the distortion effect caused by the filter is considered during the conversion process, thereby improving the encoding and decoding performance.
By considering the impact of neural network filters in the RDO process, the best encoding and decoding mode can be selected to improve the performance and efficiency of video encoding and decoding.
Smart Images

Figure CN120077644A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to video processing technologies, and more particularly to rate distortion optimization based on neural network loop filtering for image / video coding and decoding. Background Art
[0002] Nowadays, digital video capabilities are being applied to all aspects of people's lives. For video encoding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard. However, there is an overall expectation to further improve the coding and decoding efficiency of video coding and decoding technologies. Summary of the Invention
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method includes: determining whether to apply at least one NN model for neural network (NN) filtering during the process of a video unit for the conversion between the video unit of a video and the bitstream of the video unit; based on the determination, processing the video unit by applying the process to the video unit; and performing the conversion based on the processed video unit. In this way, the influence of reducing distortion caused by the NN filter is considered during the RDO process, thereby improving the coding and decoding performance.
[0005] In a second aspect, a device for video processing is proposed. The device includes a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to execute the method according to the first aspect of the present disclosure.
[0006] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores the bitstream generated by the method executed by a device for video processing for a video. The method includes: determining whether to apply at least one NN model for neural network (NN) filtering during the process of a video unit of the video; based on the determination, processing the video unit by applying the process to the video unit; and generating a bitstream based on the processed video unit.
[0008] In a fifth aspect, a method for storing a bitstream of a video is provided. The method includes: determining whether to apply at least one neural network (NN) model for NN filtering during a process of a video unit of the video; based on the determination, processing the video unit by applying the process to the video unit; generating a bitstream based on the processed video unit; and storing the bitstream in a non-transitory computer-readable recording medium.
[0009] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other objects, features, and advantages of the example embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0011] Figure 1 FIG. illustrates a block diagram showing an example video codec system according to some embodiments of the present disclosure;
[0012] Figure 2 FIG. illustrates a block diagram showing a first example video encoder according to some embodiments of the present disclosure;
[0013] Figure 3 FIG. illustrates a block diagram showing an example video decoder according to some embodiments of the present disclosure;
[0014] Figure 4 FIG. illustrates an example diagram showing an example of raster scan strip partitioning of a picture;
[0015] Figure 5 FIG. illustrates an example diagram showing an example of rectangular strip partitioning of a picture;
[0016] Figure 6 FIG. illustrates an example diagram showing an example of a picture segmented into slices, tiles, and rectangular strips;
[0017] Figure 7A FIG. illustrates an example diagram showing an example of a coding tree block (CTB) across the bottom picture boundary;
[0018] Figure 7B FIG. illustrates an example diagram showing an example of a CTB across the right picture boundary;
[0019] Figure 7C FIG. illustrates an example diagram showing an example of a CTB across the bottom-right picture boundary;
[0020] Figure 8An example diagram showing an encoder block diagram is illustrated;
[0021] Figure 9 An example diagram showing picture samples on an 8×8 grid, as well as horizontal and vertical block boundaries and non-overlapping blocks of 8×8 samples is illustrated;
[0022] Figure 10 An example diagram showing pixels involved in filter on / off decision and strong / weak filter selection is illustrated;
[0023] Figures 11A to 11D An example diagram showing four one-dimensional direction diagrams for EO sample classification is illustrated;
[0024] Figures 12A to 12C An example diagram showing an example of the GALF filter shape is illustrated;
[0025] Figures 13A to 13C An example diagram showing an example of the relative coordinates for 5×5 rhombus filter support is illustrated;
[0026] Figure 14 An example diagram showing an example of the relative coordinates for 5×5 rhombus filter support is illustrated;
[0027] Figure 15A An example diagram showing the architecture of the proposed CNN filter is illustrated;
[0028] Figure 15B An example diagram showing the construction of the ResBlock (residual block) in the CNN filter is illustrated;
[0029] Figure 16A An example diagram showing the architecture of the proposed CNN filter is illustrated;
[0030] Figure 16B An example diagram showing the construction of the attention residual block in FIG. 16 is illustrated;
[0031] Figure 17 A flowchart of a method for video processing according to an embodiment of the present disclosure is illustrated; and
[0032] Figure 18 A block diagram of a computing device in which various embodiments of the present disclosure can be implemented is illustrated.
[0033] Throughout the drawings, the same or similar reference numerals generally refer to the same or similar elements. Detailed Description
[0034] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is for illustrative purposes only and to assist those skilled in the art in understanding and implementing the present disclosure, and does not imply any limitation on the scope of the present disclosure. The disclosure described herein can be implemented in various ways other than those described below.
[0035] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0036] References in this disclosure to "one embodiment", "an embodiment", "example embodiment", etc., indicate that the embodiment described may include a particular feature, structure, or characteristic, but not every embodiment must include that particular feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is contended that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art in relation to other embodiments.
[0037] It should be understood that although terms such as "first" and "second" may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0038] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the example embodiments. As used herein, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "comprises", "comprising", "has", "having", "includes", and / or "including" when used herein specify the presence of the stated features, elements, and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Example Environment
[0039] Figure 1FIG. 0 is a block diagram illustrating an example video coding and decoding system 100 that can utilize the techniques of the present disclosure. As shown, video coding and decoding system 100 can include a source device 110 and a destination device 120. Source device 110 may also be referred to as a video encoding device, and destination device 120 may also be referred to as a video decoding device. In operation, source device 110 may be configured to generate encoded video data, and destination device 120 may be configured to decode the encoded video data generated by source device 110. Source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0040] Video source 112 can include sources such as video capture devices. Examples of video capture devices include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.
[0041] The video data can include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream can include a sequence of bits that forms an encoded representation of the video data. The bitstream can include encoded pictures and associated data. An encoded picture is an encoded representation of a picture. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 can include a modulator / demodulator and / or a transmitter. The encoded video data can be directly transmitted to destination device 120 via I / O interface 116 over network 130A. The encoded video data can also be stored on storage medium / server 130B for access by destination device 120.
[0042] Destination device 120 can include an I / O interface 126, a video decoder 124, and a display device 122. I / O interface 126 can include a receiver and / or a modulator. I / O interface 126 can obtain the encoded video data from source device 110 or storage medium / server 130B. Video decoder 124 can decode the encoded video data. Display device 122 can display the decoded video data to a user. Display device 122 can be integrated with destination device 120 or can be external to destination device 120, which is configured to interface with an external display device.
[0043] Video encoder 114 and video decoder 124 can operate according to video compression standards such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other existing and / or future standards.
[0044] Figure 2is a block diagram showing an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 can be an Figure 1 example of the video encoder 114 in the system 100 shown.
[0045] The video encoder 200 can be configured to implement any or all of the techniques of the present disclosure. In Figure 2 an example, the video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure can be shared among the various components of the video encoder 200. In some examples, a processor can be configured to execute any or all of the techniques described in the present disclosure.
[0046] In some embodiments, the video encoder 200 can include a splitting unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214. The prediction unit 202 can include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206.
[0047] In other examples, the video encoder 200 can include more, fewer, or different functional components. In one example, the prediction unit 202 can include an Intra Block Copy (IBC) unit. The IBC unit can perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0048] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) can be integrated, for purposes of explanation, these components are shown separately in Figure 2 an example.
[0049] The splitting unit 201 can split a picture into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0050] The mode selection unit 203 can select, for example, one coding mode (intra coding or inter coding) from a plurality of coding modes based on an error result, and provide the resulting intra-coded block or inter-coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a Combined Intra and Inter Prediction (CIIP) mode in which the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 can also select a resolution for the motion vector for the block (e.g., sub-pixel accuracy or integer pixel accuracy).
[0051] To perform inter prediction on a current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from the cache 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and the decoded samples of pictures from the cache 213 other than the picture associated with the current video block.
[0052] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, e.g., depending on whether the current video block is in an I-slice, a P-slice, or a B-slice. As used herein, an "I-slice" may refer to a part of a picture composed of macroblocks, all of which are based on macroblocks within the same picture. Additionally, as used herein, in some aspects, a "P-slice" and a "B-slice" may refer to parts of a picture composed of macroblocks independent of macroblocks in the same picture.
[0053] In some examples, the motion estimation unit 204 may perform uni-directional prediction on the current video block, and the motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. The motion estimation unit 204 may then generate a reference index and a motion vector, the reference index indicating the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0054] Alternatively, in other examples, the motion estimation unit 204 may perform bi-directional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 to find one reference video block for the current video block, and may also search the reference pictures in list 1 to find another reference video block for the current video block. The motion estimation unit 204 may then generate a plurality of reference indices and a plurality of motion vectors, the plurality of reference indices indicating the plurality of reference pictures in list 0 and list 1 that contain the plurality of reference video blocks, and the plurality of motion vectors indicating the plurality of spatial displacements between the plurality of reference video blocks and the current video block. The motion estimation unit 204 may output the plurality of reference indices and the plurality of motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the plurality of reference video blocks indicated by the motion information of the current video block.
[0055] In some examples, the motion estimation unit 204 may output a complete set of motion information for use in the decoding process of the decoder. Alternatively, in some embodiments, the motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is similar enough to the motion information of neighboring video blocks.
[0056] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block, which value indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0057] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0058] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0059] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0060] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the (multiple) predicted video blocks of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0061] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[0062] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0063] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0064] The inverse quantization unit 210 and the inverse transform unit 211 may respectively apply inverse quantization and inverse transform to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples of one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the cache 213.
[0065] After the reconstruction unit 212 reconstructs the video block, a loop filter operation may be performed to reduce video block effect artifacts in the video block.
[0066] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0067] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 an example of the video decoder 124 in the system 100 shown.
[0068] The video decoder 300 may be configured to perform any or all of the techniques of the present disclosure. In Figure 3 an example, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0069] In Figure 3 an example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a cache 307. In some examples, the video decoder 300 may perform a decoding process generally opposite to the encoding process described with respect to the video encoder 200.
[0070] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream can include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, the motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, including deriving several most likely candidates based on data from neighboring PBs and reference pictures. Motion information generally includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indices, and in the case of the prediction region in a B slice, also includes an indication of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" can refer to deriving motion information from spatially neighboring blocks or temporally neighboring blocks.
[0071] The motion compensation unit 302 can generate a motion-compensated block, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter used at sub-pixel precision can be included in the syntax element.
[0072] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of a video block to calculate interpolated values for sub-integer pixels of a reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 according to received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate a prediction block.
[0073] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks for encoding the (multiple) frames and / or (multiple) slices of the encoded video sequence, the partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame encoded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A slice can be the entire picture, or it can also be a region of the picture.
[0074] The intra prediction unit 303 can use, for example, the intra prediction mode received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0075] The reconstruction unit 306 can obtain the decoded block, for example, by adding a residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block effect artifacts. The decoded video block is then stored in the cache 307, and the cache 307 provides a reference block for subsequent motion compensation / intra prediction, and the cache 307 also generates the decoded video for presentation on a display device.
[0076] Some exemplary embodiments of the present disclosure will be described in detail below. It should be noted that the use of section headings in this document is for ease of understanding and does not limit the embodiments disclosed in the section to that section. In addition, although some embodiments are described with reference to multi-functional video coding or other specific video codecs, the disclosed techniques are also applicable to other video coding techniques. In addition, although some embodiments describe the video coding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by a decoder. In addition, the term video processing includes video coding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at different compression bit rates. 1. Brief Overview The present disclosure relates to video coding and decoding techniques. Specifically, it relates to loop filters in image / video coding and decoding. It can be applied to existing video coding and decoding standards, such as High Efficiency Video Coding (HEVC), Multi-functional Video Coding (VVC), or standards to be completed (e.g., AVS3). It is also applicable to future video coding and decoding standards or video codecs, or used as a post-processing method outside the encoding / decoding process. 2. Introduction Video coding standards have evolved mainly through the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Due to H.262, video coding standards are based on a hybrid video coding structure that utilizes temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was created to work on the VVC standard targeting a 50% bitrate reduction compared to HEVC. VVC Version 1 was completed in July 2020. The latest version of the VVC draft, i.e., Versatile Video Coding (Draft 10) can be found at: http: / / phenix.it-sudparis.eu / jvet / doc_end_user / current_document.php?id=10399. The latest reference software for VVC (named VTM) can be found at: https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / - / tags / VTM-10.0. 2.1 Color Spaces and Chroma Subsampling A color space (also called a color model (or color system)) is an abstract mathematical model that simply describes a color range as a digital tuple, typically 3 or 4 values or color components (such as RGB). Basically, a color space is a refinement of a coordinate system and subspace. For video compression, the most frequently used color spaces are YCbCr and RGB. YCbCr, Y′CbCr or Y Pb / Cb Pr / Cr (also written as YCBCR or Y′CBCR) is a family of color spaces used as part of the color image pipeline in video and digital photography systems. Y′ is the luminance component, and CB and CR are the blue-difference and red-difference chrominance components. Y′ (with a superscript symbol) is distinguished from Y, which is luminance, meaning that the light intensity is non-linearly encoded based on gamma-corrected RGB primaries. Chroma subsampling is the practice of encoding an image by achieving a lower resolution for chrominance information than for luma information, taking advantage of the fact that the human visual system is less sensitive to color differences than to luminance. 2.1.1 4:4:4 Each of the three Y'CbCr components has the same sampling rate, so there is no chroma subsampling. This scheme is sometimes used in high-end film scanners and film post-production. 2.1.2 4:2:2 The two chrominance components are sampled at half the sampling rate of luma: the horizontal chrominance resolution is halved. This reduces the bandwidth of the uncompressed video signal by one-third with little visual difference. 2.1.3 4:2:0 In 4:2:0, the horizontal sampling is doubled compared to 4:1:1, but since the Cb and Cr channels are only sampled on every other line in this scheme, the vertical resolution is halved. Thus, the data rate is the same. Cb and Cr are each subsampled horizontally and vertically by a factor of 2. There are three variants of the 4:2:0 scheme, which have different horizontal and vertical positioning. · In MPEG-2, Cb and Cr are horizontally co-located. Cb and Cr are positioned between pixels in the vertical direction (positioned with gaps). · In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are positioned with gaps, in the middle between alternating luma samples. · In 4:2:0 DV, Cb and Cr are horizontally co-located. In the vertical direction, they are co-located on alternating lines. 2.2 Definition of Video Units A picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of CTUs that cover a rectangular region of the picture. A tile is divided into one or more bricks, each brick consisting of multiple CTU rows within the tile. A tile that is not divided into multiple bricks is also called a brick. However, a brick that is a true subset of a tile is not called a tile. A strip contains multiple tiles of a picture or multiple bricks of a tile. Two modes of strips are supported, namely the raster scan strip mode and the rectangular strip mode. In the raster scan strip mode, a strip contains a sequence of tiles in the tile raster scan of the picture. In the rectangular strip mode, a strip contains multiple bricks of the picture that together form a rectangular region of the picture. The bricks within a rectangular strip are in the order of the brick raster scan of the strip. Figure 4 Example diagram 400 illustrates an example showing the raster scan strip partitioning of a picture. In Figure 4In it, the picture is divided into 12 slices and 3 raster scan stripes. Figure 4 In it, a picture with 18×12 luma CTUs is divided into 12 slices and 3 raster scan stripes (informative). Figure 5 Fig. 500 shows an example diagram illustrating an example of the rectangular stripe division of a picture. In Figure 5 it, the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular stripes. Figure 5 In it, a picture with 18×12 luma CTUs is divided into 24 slices and 9 rectangular stripes (informative). Figure 6 Fig. 600 shows an example diagram illustrating an example of a picture divided into slices, tiles, and rectangular stripes. In Figure 6 it, the picture is divided into 4 slices (2 slice columns and 2 slice rows), 11 tiles (the upper left slice contains 1 tile, the upper right slice contains 5 tiles, the lower left slice contains 2 tiles, and the lower right slice contains 3 tiles), and 4 rectangular stripes. Figure 6 The picture in it is divided into 4 slices, 11 tiles, and 4 rectangular stripes (informative). 2.2.1 CTU / CTB Sizes In VVC, the CTU size signaled by the syntax element log2_ctu_size_minus2 in the SPS can be as small as 4x4. 7.3.2.3 Sequence Parameter Set RBSP Syntax log2_ctu_size_minus2 plus 2 specifies the luma coding tree block size of each CTU. log2_min_luma_coding_block_size_minus2 plus 2 specifies the minimum luma coding block size. Derive the following variables: CtbLog2SizeY, CtbSizeY, MinCbLog2SizeY, MinCbSizeY, MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, MaxTbSizeY, PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC as follows: CtbLog2SizeY = log2_ctu_size_minus2 + 2 (7-9) CtbSizeY = 1 << CtbLog2SizeY (7-10) MinCbLog2SizeY = log2_min_luma_coding_block_size_minus2 + 2 (7-11) MinCbSizeY = 1 << MinCbLog2SizeY (7-12) MinTbLog2SizeY = 2 (7-13) MaxTbLog2SizeY = 6 (7-14) MinTbSizeY = 1 << MinTbLog2SizeY (7-15) MaxTbSizeY = 1 << MaxTbLog2SizeY (7-16) PicWidthInCtbsY = Ceil(pic_width_in_luma_samples ÷ CtbSizeY) (7-17) PicHeightInCtbsY = Ceil(pic_height_in_luma_samples ÷ CtbSizeY) (7-18) PicSizeInCtbsY = PicWidthInCtbsY * PicHeightInCtbsY (7-19) PicWidthInMinCbsY = pic_width_in_luma_samples / MinCbSizeY (7-20) PicHeightInMinCbsY = pic_height_in_luma_samples / MinCbSizeY (7-21) PicSizeInMinCbsY = PicWidthInMinCbsY * PicHeightInMinCbsY (7-22) PicSizeInSamplesY = pic_width_in_luma_samples * pic_height_in_luma_samples (7-23) PicWidthInSamplesC = pic_width_in_luma_samples / SubWidthC (7-24) PicHeightInSamplesC = pic_height_in_luma_samples / SubHeightC (7-25). 2.2.2 CTUs in the Picture Assume the CTB / LCU size indicated by M×N (usually M equals N, as defined in HEVC / VVC), and for the CTBs located at the picture (or slice or strip or other kind of type, picture boundary as an example) boundary, K×L samples are within the picture boundary, where K < M or L < N. Figure 7A Fig. 700 shows an example diagram illustrating a CTB spanning the bottom picture boundary, where K = M, L < N. Figure 7B Fig. 720 shows an example diagram illustrating a CTB spanning the right picture boundary, where K < M, L = N. Figure 7C Fig. 740 shows an example diagram illustrating a CTB spanning the bottom right picture boundary, where K < M, L < N. For those CTBs depicted as Figures 7A to 7C above, the CTB size is still equal to MxN. However, the bottom boundary / right boundary of the CTB is outside the picture. 2.3 Coding and Decoding Processes of Typical Video Codecs Figure 8FIG. 800 shows an example diagram illustrating an encoder block diagram of VVC, which includes three loop filter blocks: Deblocking Filter (DF) 805, Sample Adaptive Offset (SAO) 806, and ALF 807. Different from DF 805 that uses predefined filters, SAO 806 and ALF 807 utilize the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding offsets and by applying finite impulse response (FIR) filters respectively, where the transcoded side information signals the offsets and filter coefficients. ALF 807 is located at the last processing stage of each picture and can be regarded as a tool to attempt to capture and fix the artifacts created by the previous stages. 2.4 Deblocking Filter (DB) The input of the DB is the reconstructed samples before the loop filter. The vertical edges in the picture are filtered first. Then, the samples modified by the vertical edge filtering process are used as the input to filter the horizontal edges in the picture. The vertical and horizontal edges in the coding tree block (CTB) of each coding tree unit (CTU) are processed separately based on the coding unit. The vertical edges of the coding block in the coding unit are filtered in their geometric order starting from the left - hand side edge of the coding block towards the right - hand side. The horizontal edges of the coding block in the coding unit are filtered in their geometric order starting from the top - hand side edge of the coding block towards the bottom. Figure 9 FIG. 900 shows an example diagram illustrating picture samples on an 8×8 grid, as well as horizontal and vertical block boundaries and non - overlapping blocks of 8×8 samples (which can be deblocked in parallel). 2.4.1 Boundary Decision Filtering is applied to the 8x8 block boundaries. Additionally, it must be a transform block boundary or a coding sub - block boundary (e.g., due to the use of affine motion prediction ATMVP). For those codings that are not such boundaries, the filter is disabled. 2.4.2 Boundary Strength Calculation For a transform block boundary / coding sub - block boundary, if it is located in the 8x8 grid, it can be filtered, and the setting of bS[xD i [yD j (where [xD i [yD j represents coordinates) is defined in Tables 1 and 2 respectively. Table 1 Boundary Strength (when SPS IBC is disabled) Table 2 Boundary Strength (when SPS IBC is enabled) 2.4.3 Deblocking Decision for Luminance Component The deblocking decision process is described in this subsection. Figure 10 Example diagram 1000 showing the pixels involved in filter on / off decision and strong / weak filter selection is illustrated. A stronger luminance filter is used only when Condition1, Condition2, and Condition3 are all true. Condition 1 is the "large block condition". This condition detects whether the samples on the P side and Q side belong to large blocks, represented by variables bSidePisLargeBlk and bSideQisLargeBlk respectively. bSidePisLargeBlk and bSideQisLargeBlk are defined as follows. bSidePisLargeBlk = ((the side type is vertical and p 0 belongs to a CU with width >= 32) || (the side type is horizontal and p 0 belongs to a CU with height >= 32))? true : false bSideQisLargeBlk = ((the side type is vertical and q 0 belongs to a CU with width >= 32) || (the side type is horizontal and q 0 belongs to a CU with height >= 32))? true : false Based on bSidePisLargeBlk and bSideQisLargeBlk, Condition 1 is defined as follows. Condition1 = (bSidePisLargeBlk || bSidePisLargeBlk)? true : false Next, if Condition1 is true, Condition 2 will be further checked. First, the following variables are derived: – dp0, dp3, dq0, dq3 are first derived as in HEVC. – If (the p side is greater than or equal to 32) dp0 = (dp0 + Abs(p5 0 - 2 * p4 0 + p3 0 )) + 1) >> 1 dp3 = (dp3 + Abs(p5 3 - 2 * p4 3 + p3 3 )) + 1) >> 1 – If (the q side is greater than or equal to 32) dq0 = (dq0 + Abs(q5 0 - 2 * q4 0 + q30 ) + 1) >> 1 dq3 = (dq3 + Abs(q5 3 - 2 * q4 3 + q3 3 ) + 1) >> 1 Condition2 = (d < β)? true : false where d = dp0 + dq0 + dp3 + dq3. If Condition1 and Condition2 are valid, further check whether any block uses a sub - block: Finally, if both Condition1 and Condition2 are valid, the proposed de - blocking method will check Condition3 (the large - block strong filter condition), which is defined as follows. In Condition3 StrongFilterCondition, the following variables are derived: As in HEVC, StrongFilterCondition = (dpq less than (β >> 2), sp 3 + sq 3 less than (3 * β >> 5), and Abs(p 0 - q 0 ) less than (5 * t C + 1) >> 1)? true : false. 2.4.4 Stronger Deblocking Filter for Luminance (Designed for Larger Blocks) When the samples on either side of the boundary belong to a large block, a bilinear filter is used. Samples belonging to a large block are defined as those for the vertical edge when the width >= 32, and for the horizontal edge when the height >= 32. The bilinear filter is listed below. Block - boundary sample p i (i = 0 to Sp - 1) and q i (j = 0 to Sq - 1) (in the above HEVC de - blocking, pi and qi are the i - th sample in the row for filtering the vertical edge, or the i - th sample in the column for filtering the horizontal edge) are then linearly interpolated and replaced as follows: — p i ' = (fi * Middle s,t +(64 - f i ) * P s + 32) >> 6), clamped to p i ± tcPDi —q j ′=(g j *Middle s,t +(64 - g j )*Q s + 32) >> 6), clamped to q j ±tcPD j where the terms tcPD i and tcPD j are position - dependent clamps described in Section 2.4.7, and g j , f i , Middle s,t , P s and Q s are given below. 2.4.5 Chroma Deblocking Control Chroma strong filters are applied to both sides of the block boundary. Here, the chroma filter is selected when both sides of the chroma edge are greater than or equal to 8 (chroma position), and the following decision with three conditions is satisfied: The first condition is for the decision of boundary strength and large blocks. When the block width or height orthogonally crossing the block edge is equal to or greater than 8 in the chroma sample domain, the proposed filter can be applied. The second and third conditions are basically the same as the HEVC luma deblocking decisions, which are the on - off decision and the strong filter decision respectively. In the first decision, the boundary strength (bS) is modified for chroma filtering, and the conditions are checked sequentially. If the condition is satisfied, the remaining conditions with lower priority are skipped. When bS is equal to 2, chroma deblocking is performed, or when a large block boundary is detected, bS is equal to 1. The second and third conditions are basically the same as the HEVC luma strong filter decisions as follows. Under the second condition: Subsequently, d is derived as in HEVC luma deblocking. The second condition will be true when d is less than β. Under the third condition, StrongFilterCondition is derived as follows: dpq is derived as in HEVC. sp 3 = Abs(p 3 - p 0 ), derived as in HEVC. sq 3 = Abs(q 0 - q 3 ), derived as in HEVC. In the HEVC design, StrongFilterCondition = (dpq is less than (β >> 2), sp 3 + sq 3 is less than (β >> 3), and Abs(p 0 - q 0 ) is less than (5 * t C + 1) >> 1). 2.4.6 Strong Deblocking Filter for Chrominance The strong deblocking filter for chrominance is defined as follows: p 2 ′ = (3 * p 3 + 2 * p 2 + p 1 + p 0 + q 0 + 4) >> 3 p 1 ′ = (2 * p 3 + p 2 + 2 * p 1 + p 0 + q 0 + q 1 + 4) >> 3 p 0 ′ = (p 3 + p 2 + p 1 + 2 * p 0 + q 0 + q 1 + q 2 + 4) >> 3. The proposed chrominance filter performs deblocking on a 4x4 chrominance sample grid. 2.4.7 Position - Dependent Clipping The position - dependent clipping tcPD is applied to the output samples of the luminance filtering process, which involves strong and long filters that modify 7, 5, and 3 samples at the boundaries. Assuming a quantization error distribution, it is proposed to increase the clipping value for samples expected to have a higher quantization noise, and thus a higher deviation of the reconstructed sample value from the true sample value is expected. For each P or Q boundary filtered with an asymmetric filter, depending on the result of the decision - making process in Section 2.4.2, the position - dependent threshold table is selected from two tables (i.e., Tc7 and Tc3 below) provided to the decoder as side information: Tc7 = {6, 5, 4, 3, 2, 1, 1}; Tc3 = {6, 4, 2}; tcPD = (Sp == 3)? Tc3 : Tc7; tcQD = (Sq == 3)? Tc3 : Tc7; For P or Q boundaries filtered using short symmetric filters, lower magnitude position-dependent thresholds are applied: Tc3 = {3, 2, 1}; After defining the thresholds, the filtered p' i and q' i sample values are clipped according to the tcP and tcQ clipping limits: p" i = Clip3(p' i + tcP i , p' i – tcP i , p' i ); q" j = Clip3(q' j + tcQ j , q' j – tcQ j , q' j ); where p' i and q' i are the filtered sample values, p" i and q" j are the output sample values after clipping, and tcP i tcP i is the clipping threshold derived from the VVC tc parameters and tcPD and tcQD. The function Clip3 is the clipping function as specified in VVC. 2.4.8 Sub-block Deblocking Adjustment To achieve parallel-friendly deblocking using long filters and sub-block deblocking, the long filter is restricted to modifying at most 5 samples on the side where sub-block deblocking (AFFINE or ATMVP or DMVR) is used, as shown for the luminance control of the long filter. Additionally, the sub-block deblocking is adjusted such that the sub-block boundaries on the 8x8 grid near the CU or implicit TU boundary are restricted to modifying at most two samples on each side. The following applies to sub-block boundaries that are not aligned with the CU boundary. where the side equal to 0 corresponds to the CU boundary, and the sides equal to 2 or equal to orghogonalLength - 2 correspond to the sub-block boundaries 8 samples from the CU boundary, etc. Where if the implicit partitioning of the TU is used, implicitTU is true. 2.5 SAO The input to SAO is the reconstructed samples after DB. The SAO concept reduces the average sample distortion of a region by first classifying the region samples into multiple classes with a selected classifier, obtaining an offset for each class, and then adding the offset to each sample of the class, where the classifier index and the offset of the region are encoded and decoded in the bitstream. In HEVC and VVC, a region (the unit for SAO parameter signaling) is defined as a CTU. Two SAO types that meet the low complexity requirements are adopted in HEVC. These two types are Edge Offset (EO) and Band Offset (BO), which will be discussed in further detail below. The index of the SAO type is encoded and decoded (which is in the range of [0, 2]). For EO, sample classification is based on the comparison between the current sample and its neighboring samples according to a one-dimensional direction pattern: horizontal, vertical, 135° diagonal, and 45° diagonal. Figure 11A Fig. 1100 is an example diagram showing a one-dimensional direction pattern for EO sample classification with horizontal (EO class = 0). Figure 11B Fig. 1120 is an example diagram showing a one-dimensional orientation pattern for EO sample classification with vertical (EO class = 1). Figure 11C Fig. 1140 is an example diagram showing a one-dimensional direction pattern for EO sample classification with 135° diagonal (EO class = 2). Figure 11D Fig. 1160 is an example diagram showing a one-dimensional direction pattern for EO sample classification with 45° diagonal (EO class = 3). For a given EO class, each sample within a CTB is classified into one of five classes. The current sample value labeled "c" is compared with its two neighbors along the selected one-dimensional pattern. The classification rules for each sample are summarized in Table 1. Classes 1 and 4 are associated with local valleys and local peaks along the selected one-dimensional pattern respectively. Classes 2 and 3 are associated with concave corners and convex corners along the selected one-dimensional pattern respectively. If the current sample does not belong to EO classes 1 to 4, it is class 0 and SAO is not applied. Table 3: Sample Classification Rules for Edge Offset Category Condition 1 c < a and c < b 2 (c < a && c == b) || (c == a && c < b) 3 (c > a && c == b) || (c == a && c > b) 4 c > a && c > b 5 None of the above 2.6 Geometric-Transformation-Based Adaptive Loop Filter in JEM The input to DB is the reconstructed samples after DB and SAO. The sample classification and filtering process are based on the reconstructed samples after DB and SAO. In JEM, a Geometric-Transformation-Based Adaptive Loop Filter (GALF) with block-based filter adaptation is applied. For the luminance component, one of 25 filters is selected for each 2×2 block based on the direction and activity of the local gradient. 2.6 Filter Shape Figure 12A Example diagram 1200 shows an example of the shape of a GALF filter having a 5×5 rhombus. Figure 12B Example diagram 1220 shows an example of the shape of a GALF filter having a 7×7 rhombus. Figure 12C Example diagram 1240 shows an example of the shape of a GALF filter having a 9×9 rhombus. In JEM, up to three rhombus filter shapes can be selected for the luminance component (as Figures 12A to 12C shown). The filter shape for the luminance component is indicated by a signal transmission index at the picture level. Each square represents a sample point, and Ci (i ranges from 0 to 6 (left), 0 to 12 (middle), 0 to 20 (right)) represents the coefficient to be applied to the sample point. For the chrominance component in the picture, a 5×5 rhombus shape is always used. 2.6.1.1 Block Classification Each 2×2 block is classified into one of 25 classes. The classification index C is derived based on its directionality D and the quantized value of activity as follows: To calculate D and First, the gradients in the horizontal, vertical, and two diagonal directions are calculated using a 1-D Laplacian operator: The indices i and j refer to the coordinates of the upper left sample point in the 2×2 block, and R(i,j) indicates the reconstructed sample point at the coordinates (i,j). Then the maximum and minimum values of the gradients in the horizontal and vertical directions are set to: And the maximum and minimum values of the gradients in the two diagonal directions are set to: To derive the value of the directionality D, these values are compared with each other and with two thresholds t 1 and t 2 as follows: Step 1. If and are both true, then D is set to 0. Step 2. If then proceed to Step 3; otherwise, proceed to Step 4. Step 3. If then D is set to 2; otherwise D is set to 1. Step 4. If then D is set to 4; otherwise D is set to 3. The activity value A is calculated as: A is further quantized to the range of 0 to 4 (including the boundary values), and the quantization value is represented as For the two chrominance components in the picture, the classification method is not applied, that is, a single set of ALF coefficients is applied to each chrominance component. 2.6.1.2 Geometric Transformations of Filter Coefficients Figure 13A Fig. 1300 shows an example diagram showing the relative coordinates for a 5×5 rhombic filter support (diagonal). Figure 13B Fig. 1320 shows an example diagram showing the relative coordinates for a 5×5 rhombic filter support (vertically flipped). Figure 13C Fig. 1340 shows an example diagram showing the relative coordinates for a 5×5 rhombic filter support (rotated). Before filtering each 2×2 block, geometric transformations such as rotation or diagonal and vertical flipping are applied to the filter coefficients f(l,k) associated with the coordinates (k, l) according to the gradient value calculated for the block. This is equivalent to applying these transformations to the samples in the filter support region. The idea is to make the different blocks to which ALF is applied more similar by aligning the directions. Three geometric transformations are introduced, including diagonal, vertical flipping, and rotation: Diagonal: f D (k,l) = f(l,k), Vertical flipping: f V (k,l) = f(k,K-l-1), (9) Rotation: f R (k,l) = f(K-l-1,k) where K is the size of the filter, and 0 ≤ k,l ≤ K-1 are the coefficient coordinates, such that the position (0,0) is in the upper left corner and the position (K-1,K-1) is in the lower right corner. The transformation is applied to the filter coefficients f(k,l) according to the gradient value calculated for the block. The relationship between the transformation and the four gradients in the four directions is summarized in Table 4. Figures 12A to 12C The transformed coefficients for each position based on the 5x5 rhombus are shown. Table 4 Mapping of Gradients Calculated for a Block and a Transformation Gradient value Transformation <![CDATA[g d2 <g d1 and g h <g v > No transformation <![CDATA[g d2 <g d1 and g v <g h > Diagonal <![CDATA[g d1 <g d2 and g h <g v > Vertical flip <![CDATA[g d1 <g d2 and g v <g h > Rotation 2.6.1.3 Filter Parameter Signaling In JEM, GALF filter parameters are signaled for the first CTU (i.e., after the slice header and before the SAO parameters of the first CTU). Up to 25 sets of luminance filter coefficients can be signaled. To reduce the bit overhead, filter coefficients of different classifications can be merged. Additionally, the GALF coefficients of the reference picture are stored and are allowed to be used as the GALF coefficients of the current picture. The current picture can choose to use the stored GALF coefficients for the reference picture and bypass the GALF coefficient signaling. In this case, only the index of one of the reference pictures in the reference pictures is signaled, and the stored GALF coefficients of the indicated reference picture are inherited for the current picture. To support GALF temporal prediction, a candidate list of GALF filter sets is maintained. At the start of decoding a new sequence, the candidate list is empty. After decoding a picture, the corresponding filter set can be added to the candidate list. Once the size of the candidate list reaches the maximum allowed value (i.e., 6 in the current JEM), the new filter set overwrites the oldest set in decoding order, i.e., the first-in-first-out (FIFO) rule is applied to update the candidate list. To avoid duplicates, a set can be added to the list only when the corresponding picture does not use GALF temporal prediction. To support temporal scalability, there are multiple candidate lists of filter sets, and each candidate list is associated with a temporal layer. More specifically, each array assigned by the temporal layer index (TempIdx) can consist of filter sets of previously decoded pictures with a lower TempIdx. For example, the k-th array is assigned to be associated with a TempIdx equal to k, and it contains only filter sets from pictures with a TempIdx less than or equal to k. After encoding / decoding a certain picture, the filter sets associated with the picture are used to update those arrays associated with an equal or higher TempIdx. Temporal prediction of GALF coefficients is used for inter-coded frames to minimize the signaling overhead. For intra frames, temporal prediction is not available, and a set of 16 fixed filters is assigned to each class. To indicate the use of the fixed filters, a flag for each class is signaled, and if needed, the index of the selected fixed filter is signaled. Even when a fixed filter is selected for a given class, the coefficients f(k,l) of the adaptive filter can still be sent for that class, in which case the coefficients of the filter applied to the reconstructed image are the sum of the two sets of coefficients. The filtering process of the luminance component can be controlled at the CU level. A flag is signaled to indicate whether GALF is applied to the luminance component of the CU. For the chrominance components, whether GALF is applied is indicated only at the picture level. 2.6.1.4 Filtering Process At the decoder side, when GALF is enabled for a block, each sample R(i,j) within the block is filtered to produce a sample value R′(i,j) as shown below, where L represents the filter length, f m,n represents the filter coefficients, and f(k,l) represents the decoded filter coefficients. Figure 14 FIG. 1400 is an example diagram showing an example of relative coordinates supported for a 5×5 diamond filter. Figure 14 An example of relative coordinates supported for a 5x5 diamond filter is shown assuming the coordinates (i, j) of the current sample are (0,0). Samples in different coordinates filled with the same color are multiplied by the same filter coefficient. 2.7 Geometric-Transform-Based Adaptive Loop Filter (GALF) in VVC 2.7.1 GALF in VTM-4 In VTM4.0, the filtering process of the adaptive loop filter is performed as follows: O(x, y) = ∑ (i,j) w(i, j).I(x + i, y + j), (11) where the sample I(x + i, y + j) is the input sample, O(x, y) is the filtered output sample (i.e., the filter result), and w(i, j) represents the filter coefficients. In fact, in VTM4.0, integer operations are used to achieve fixed-point precision calculations: where L represents the filter length, and where w(i, j) are the filter coefficients in fixed-point precision. Compared to JEM, the current design of GALF in VVC has the following main changes: 1) The adaptive filter shape is removed. Only 7x7 filter shapes are allowed for the luma component, and only 5x5 filter shapes are allowed for the chroma component. 2) The signaling of ALF parameters is removed from the slice / picture level to the CTU level. 3) The calculation of the class index is performed at the 4x4 level instead of the 2x2 level. Additionally, a subsampled Laplacian calculation method is used for ALF classification. More specifically, it is not necessary to calculate the horizontal / vertical / 45-degree diagonal / 135-degree gradients for each sample within a block. Instead, 1:2 subsampling is used. 2.8 Nonlinear ALF in Current VVC 2.8.1 Filtering Reformulation Equation (11) can be reformulated in the following expression without affecting the coding / decoding efficiency: O(x,y) = I(x,y) + ∑ (i,j)≠(0,0) w(i,j)·(I(x + i,y + j) - I(x,y)), (13) where w(i,j) are the same filter coefficients as in equation (11) [except that w(0,0) in equation (13) is equal to 1, which is equal to 1 - ∑ (i,j)≠(0,0) w(i,j)] in equation (11). Using the above filter formula (13), VVC introduces non - linearity to make the ALF more effective by using a simple clipping function, thus reducing the influence of neighboring sample values (I(x + i,y + j)) that are very different from the current sample value (I(x,y)) being filtered. More specifically, the ALF filter is modified as follows: O′(x,y) = I(x,y) + ∑ (i,j)≠(0,0) w(i,j)·K(I(x + i,y + j) - I(x,y),k(i,j)), (14) where K(d,b) = min(b,max( - b,d)) is the clipping function and k(i,j) is the clipping parameter, which depends on the (i,j) filter coefficients. The encoder performs optimization to find the best k(i,j). In some implementations, the clipping parameter k(i,j) is specified for each ALF filter, and a clipping value is signaled according to each filter coefficient. This means that up to 12 clipping values can be signaled in the bit - stream for each luminance filter and up to 6 clipping values for the chrominance filter. To limit the signaling cost and encoder complexity, only 4 fixed values that are the same for both inter - frame and intra - frame stripes are used. Since the variance of the local differences for luminance is generally higher than that for chrominance, two different sets are applied for the luminance and chrominance filters. The maximum sample value in each set (here 1024 for a 10 - bit bit - depth) is also introduced so that the clipping can be disabled if not needed. The set of clipping values is provided in Table 5. The 4 values have been selected by roughly equally partitioning the full range of sample values for luminance (encoded and decoded on 10 bits) in the logarithmic domain and the range from 4 to 1024 for chrominance. More precisely, the luminance table of clipping values has been obtained by the following formula: Similarly, the chrominance table of clipping values is obtained according to the following formula: Table 5 Authorized clipping values The selected clipping value is encoded and decoded in the "alf_data" syntax element by using the Golomb coding scheme corresponding to the index of the clipping value in the above Table 5. This coding scheme is the same as the coding scheme used for the filter index. 2.9 Convolutional Neural Network-Based Loop Filters for Video Codecs 2.9.1 Convolutional Neural Networks In deep learning, convolutional neural networks (CNN or ConvNet) are a class of deep neural networks that are most commonly used to analyze visual images. They have very successful applications in image and video recognition / processing, recommendation systems, image classification, medical image analysis, and natural language processing. CNNs are regularized versions of multilayer perceptrons. Multilayer perceptrons usually refer to fully connected networks, i.e., each neuron in one layer is connected to all neurons in the next layer. The "full connectivity" of these networks makes them prone to overfitting the data. Typical ways of regularization involve adding some form of magnitude measure of the weights to the loss function. CNNs take a different approach toward regularization: they exploit hierarchical patterns in the data and use smaller and simpler patterns to assemble more complex patterns. Therefore, on the scale of connectivity and complexity, CNNs are at the lower extreme. Compared to other image classification / processing algorithms, CNN uses relatively less preprocessing. This means that the network learns filters that are manually engineered in traditional algorithms. Independence from existing knowledge and human effort in feature design is a major advantage. 2.9.2 Deep Learning for Image / Video Encoding and Decoding Image / video compression based on deep learning generally has two meanings: end-to-end compression based purely on neural networks and traditional frameworks enhanced by neural networks. The first type generally adopts an autoencoder-like structure implemented by a convolutional neural network or a recurrent neural network. Although purely relying on neural networks for image / video compression can avoid any manual optimization or manual design, the compression efficiency may not be satisfactory. Therefore, works distributed in the second type use neural networks as an aid to enhance the traditional compression framework by replacing or enhancing some modules. In this way, they can inherit the advantages of highly optimized traditional frameworks. For example, a fully connected network for intra-frame prediction is proposed. In addition to intra-frame prediction, deep learning is also used to enhance other modules. For example, the loop filter of HEVC with a convolutional neural network is replaced, and promising results are achieved. Neural networks are applied to improve arithmetic codec engines. 2.9.3 Loop Filtering Based on Convolutional Neural Network In lossy image / video compression, the reconstructed frame is an approximation of the original frame because the quantization process is irreversible, thus resulting in distortion of the reconstructed frame. To mitigate this distortion, a convolutional neural network can be trained to learn the mapping from the distorted frame to the original frame. In practice, the training must be performed before deploying the CNN-based loop filter. 2.9.3.1 Training The purpose of the training process is to find the optimal values of the parameters including weights and biases. First, a codec (such as HM, JEM, VTM, etc.) is used to compress the training dataset to generate distorted reconstructed frames. Then, the reconstructed frames are fed into the CNN, and the cost is calculated using the output of the CNN and the ground truth frame (original frame). Commonly used cost functions include SAD (Sum of Absolute Differences) and MSE (Mean Squared Error). Next, the gradient of the cost with respect to each parameter is derived by the backpropagation algorithm. Using the gradient, the values of the parameters can be updated. The above process is repeated until the convergence criterion is met. After the training is completed, the derived optimal parameters are saved for the inference phase. 2.9.3.2 Convolution Process During convolution, the filter moves across the image from left to right and top to bottom, where there is a 1-pixel column change for the horizontal movement and then a 1-pixel row change for the vertical movement. The amount of movement between the application of the filter to the input image is called the stride, and it is almost always symmetric in the height and width dimensions. For the height and width movements, the default stride of one or more in both dimensions is (1,1). Figure 15A Fig. 1500 shows an example diagram illustrating the architecture of the proposed CNN filter. Figure 15B Fig. 1550 shows an example diagram illustrating the construction of a ResBlock (residual block) in the CNN filter. In most deep convolutional neural networks, the residual block is used as a basic module and is stacked several times to build the final network, where in one example, the residual block is obtained by combining a convolutional layer, a ReLU / PReLU activation function, and a convolutional layer, as Figure 15B shown. 2.9.3.3 Inference During the inference phase, the distorted reconstructed frames are fed into the CNN and processed by the CNN model whose parameters have been determined in the training phase. The input samples to the CNN can be the reconstructed samples before or after the DB, or before or after the SAO, or before or after the ALF. 3. Problems The current NN filters have the following problems: 1. Prior art designs that apply the NN filter only after the reconstruction of all blocks before the loop filtering process within a stripe. Thus, the impact of the reduced distortion caused by the NN filter is not considered during the rate-distortion optimization (RDO) process (such as intra-mode selection, partition selection, intra-mode selection, inter-mode selection, transform kernel selection, etc.). Considering the following, the codec performance is sub-optimal: a. The best mode (e.g., codec method / partition size) of the current block selected during the RDO process may be incorrect because the distortion is calculated without the NN filter being applied. b. The reconstruction of the current block and the associated coded / decoded information have a large impact on the coding / decoding of subsequent blocks (e.g., due to intra prediction or motion prediction). If the best mode is not selected for the current block, the codec performance of the subsequence blocks will also be sub-optimal. 4. Detailed solutions The following detailed embodiments should be considered as examples for explaining the general concept. These embodiments should not be interpreted in a narrow sense. Additionally, these embodiments can be combined in any way. To solve the above problems, it is proposed to consider the NN filter during the rate-distortion optimization (RDO) process. This disclosure details how to use the NN filter model to expand the RDO range, how to use the NN filter model to select modes (e.g., intra-mode, partition mode, inter-mode, or transform kernel), and how to control the use of the NN filter model. In this disclosure, the NN filter can be any kind of NN filter, such as a convolutional neural network (CNN) filter; alternatively, it can also be applied to non-NN-based filters. In the following discussion, the NN filter can also be referred to as a CNN filter. In the following discussion, a video unit can be a sequence, picture, stripe, slice, tile, sub-picture, CTU / CTB, CTU / CTB row, one or more CUs / CBs, one or more CTUs / CTBs, one or more VPDUs (virtual pipeline data units), a sub-region within a picture / stripe / slice / tile. A parent video unit represents a unit larger than the video unit. Generally, the parent unit will contain several video units. For example, when the video unit is a CTU, the parent unit can be a stripe, CTU row, multiple CTUs, etc. Matrix-based intra prediction is denoted as MIP. Intra sub-partitioning is denoted as ISP. Multiple reference rows are denoted as MRL. Merge with MVD is denoted as MMVD. Combined intra and inter prediction is denoted as CIIP. Geometric partitioning mode is denoted as GPM. Quad-tree is denoted as QT. Binary tree is denoted as BT. Ternary tree is denoted as TT. Cross-component SAO is denoted as CCSAO. Cross-component ALF is denoted as CCALF. The width and height of the video unit are represented as W and H respectively. Regarding the integration of the NN filter model during the RDO process 1. It is proposed that at least one NN model for NN filtering can be included in the encoder. a. In one example, the NN model can be used in the RDO process. b. In one example, the NN model may not be included in the compatible decoder. c. In one example, the NN model can be simpler than another NN model used for NN filtering in the compatible decoder. 2. The NN filter model can be combined with other filter models in the encoder. a. In one example, the NN filter model can be different from the NN filter. b. In one example, the NN model can be applied before other filter models. c. In one example, the NN model can be applied after other filter models. d. In one example, other filter models can be CNN filter models, deblocking, SAO, ALF, CCSAO, CCALF. e. In one example, the NN model and / or other filter models can be applied according to a specific order or an adaptive order. i. In one example, deblocking, CNN filter, SAO, and ALF are applied in sequence. f. In one example, the order of applying the NN model and / or other filter models can depend on the coding / decoding mode / statistics of the video unit (e.g., prediction mode, qp, temporal layer, slice type, etc.). g. In one example, whether and / or how to utilize the NN model and / or other filter models can depend on the coding / decoding mode / statistics of the video unit (e.g., prediction mode, qp, temporal layer, slice type, etc.). 3. The mode decision process (e.g., RDO process) can depend on the NN filter model, for example, according to the filtered reconstruction information attributed to the NN filter model. a. In one example, the NN filter model can be used when determining the best intra prediction mode (e.g., together with the RDO of intra mode selection). b. In one example, the NN filter model can be used when determining the best coding / decoding intra method (e.g., whether to apply MIP, ISP, MRL). c. In one example, the NN filter model can be used together with the RDO of inter mode selection (e.g., whether to use AMVP or skip mode or Merge mode). d. In one example, the NN model can be used when determining the best intra- and inter-frame coding methods (e.g., whether to use an affine motion model or a translational motion model to code / decode blocks, whether to apply MMVD, CIIP, GPM, etc.). e. In one example, the NN filter model can be used together with RDO for segmentation mode selection (e.g., whether to apply QT, BT, TT, non-partitioning, etc.). f. In one example, the NN filter model can be used together with RDO for transform kernel selection. g. In one example, the NN filter model can be used when determining the best coding method including intra- and inter-frame methods (e.g., whether to apply intra-frame (MIP, ISP, etc.) or inter-frame (MMVD, AMVP, skip, etc.)). h. In one example, the NN filter model can be used whenever distortion is calculated. i. Alternatively, the NN filter model can be used whenever distortion is calculated. (For example, when calculating distortion using SSE / MSE / SSIM / MS-SSIM / IW-SSIM matrices.) ii. Alternatively, when distortion is calculated using a specific matrix (e.g., when distortion is calculated using SAD / SATD matrices), the NN filter model is not used. 4. The distortion or cost calculated during the mode decision process (e.g., the RDO process) can be modified such that the impact of the NN filtering process is taken into account. a. In one example, distortion or cost can be calculated according to a specific matrix (e.g., SSE / MSE / SSIM / MS-SSIM / IW-SSIM matrices). b. In one example, instead of using the distortion calculated between the reconstruction before the in-loop filtering method (represented by the unfiltered reconstruction) and the original samples, it is proposed to apply the NN filtering process to the reconstruction to obtain the NN-filtered reconstruction and calculate the distortion between the NN-filtered reconstruction and the original samples. c. In one example, two distortions are calculated, one between the unfiltered reconstruction and the original samples, and the other between the NN-filtered reconstruction and the original samples. i. Alternatively, in addition, a function of the two distortions is called, and the output of the function is set as the true distortion associated with the current mode to be examined during the RDO process. d. In one example, multiple distortions are calculated, one between the unfiltered reconstruction and the original samples, and other distortions between the filtered reconstructions and the original samples. i. In one example, the filtered reconstructed samples can be filtered by an NN filter model. ii. In one example, the filtered reconstructed samples can be filtered by other filter models. iii. In one example, the filtered reconstructed samples can be filtered by an NN filter model and / or other filter models. iv. Alternatively, in addition, a distortion function is called, and the output of the function is set to the true distortion associated with the current mode to be examined during the RDO process. e. In one example, two distortions are calculated, one between the filtered reconstruction and the original samples, and the other between the NN-filtered reconstruction and the original samples. The filtered reconstruction represents the reconstruction obtained using other filters but before the NN filter. i. Alternatively, in addition, a function of two distortions is called, and the output of the function is set to the true distortion associated with the current mode to be examined during the RDO process. f. In one example, the distortion is first calculated between the unfiltered reconstruction and the original samples and then scaled by a factor. i. In one example, the factor is a constant between 0 and 1.0. ii. In one example, the factor depends on the current mode to be examined during the RDO process. iii. In one example, the factor depends on lambda. iv. In one example, the factor depends on the color component. Regarding the simplification of the NN filter model in the RDO process 5. The filtering process applied to the reconstructed video unit during the RDO process can be different from the filtering process applied in the loop filtering process / post-processing process. a. In one example, the filter models can be different. b. In one example, the number of filter models can be different. c. In one example, the network structure can be different. d. In one example, the filtering process during the RDO process can be applied only to a specific sub-region of the video unit. i. In one example, it can be applied only to the boundary samples of the video unit. ii. In one example, it can be applied only to the internal samples of the video unit. e. In one example, the filtering process during the RDO process can be applied only to the down-sampled version of the video unit. 6. The NN filter model used in the RDO process can be the same as the NN filter model on the decoder side. a. In one example, the number of ResBlocks is the same as that of the decoder. 7. The NN filter model used in the RDO process can be a simplified version of the model used on the decoder side. a. In one example, the depth of the NN filter model can be different. i. In one example, the NN filter model used in the RDO process can have a shallower depth. b. In one example, the feature maps of the NN filter model can be different. i. In one example, the NN filter model used in the RDO process can have fewer feature maps. c. In one example, the number of ResBlocks of the NN filter model can be different. i. In one example, the number of ResBlocks of the NN filter model used in the RDO process can be smaller. ii. In one example, the number of ResBlocks is 1, 2, 3, 4, 5, 6. d. In one example, the convolutional kernels of the NN filter model can be different. Regarding the use of the NN filter model in the RDO process 8. Whether and / or how to use the NN filter model in the RDO process can depend on the coding / decoding mode / statistics of the video unit (e.g., prediction mode, qp, temporal layer, slice type, etc.). a. In one example, it can depend on the prediction mode, qp, temporal layer, slice type, etc. b. In one example, it can depend on the quantization step size. c. In one example, it can depend on the temporal layer. d. In one example, it can depend on the slice type. e. In one example, it can depend on the block size of the video unit. f. In one example, it can depend on the color component. g. In one example, it can depend on the rate-distortion cost without the NN filter. 5. Embodiments 5.1 Embodiment #1 In this implementation, the convolutional neural network-based loop filter with adaptive model selection (DAM) is extended to the rate-distortion optimization (RDO) process. And the number of residual blocks in DAM is reduced to 4. DAM is applied at the codec unit level to select the best segmentation structure based on the RDO criterion. The rate-distortion cost can be expressed as: J = D + lambda * R where D represents the minimum of the distortions with DAM and without DAM. Before applying DAM, check the cost J of segmentation mode A A and the cost J of segmentation mode B B . When the following conditions are met, skip DAM. J A > f 0 * J B || J B > f 1 * J A where f 0 and f 1 are parameters. 5.2 Example #2 5.2.1 Proposed method It is proposed to involve CNN-based filtering during segmentation mode selection. Specifically, the samples obtained after CNN-based filtering are compared with the original samples to calculate the distortion. Then the best segmentation mode is selected based on the refined rate-distortion (RD) cost. To reduce the complexity of applying CNN-based filtering in RDO, several fast algorithms are proposed. First, a simplified version of the CNN model as shown in Figure 16A and Figure 16B is additionally trained and used in the RDO stage, where the simplified model is implemented using fixed-point calculation-based SADL. Second, only one filter is included in the RDO process without considering filter selection. Finally, the proposed technique is only applied to codec units with a height and width not greater than 64. Figure 16A The figure illustrates an example diagram showing the architecture of the proposed CNN filter, where M represents the number of feature maps and N represents the number of samples in one dimension. Figure 16B The figure illustrates Figure 16A the construction of the attention residual block in The inference and training processes of the model are the same as those in JVET-AA0111. 6.2.2 Inference SADL is used to perform the inference of the proposed CNN filter in the RDO process. The network information in the inference stage is provided in Table 6. Table 6 Network information for testing NN-based video codec tools in the inference phase 6.2.3 Training PyTorch is used as the training platform. The DIV2K and BVI-DVC datasets are adopted to train the CNN filters for I slices and B slices respectively. The network information in the training phase is shown in Table 7. Table 7 Network information for testing NN-based video codec tools in the training phase
[0077] As used herein, the term "video unit" or "video block" can be a sequence, a picture, a slice, a tile, a brick, a sub-picture, a codec tree unit (CTU) / codec tree block (CTB), a CTU / CTB row, one or more codec units (CU) / codec blocks (CB), one or more CTUs / CTBs, one or more virtual pipeline data units (VPDU), a sub-region within a picture / slice / tile / brick. As used herein, the term "independent filter (ID) filter" can mean that the filter is not exactly the same as other filters, and some parts of the filter are different, such as the input of the filter, the structure of the filter, the parameters of the filter, the neural network model of the filter. In one example, the design of the ID filter is unique and different from the design of other filters. In one example, when the filters share a consistent structure or consistent parameters or consistent model of the neural network, the input of the ID filter is different. The ID filter can be any kind of filter, including filters without neural networks (non-NN filters) and filters with neural networks (NN filters). The non-NN filter can be one of a deblocking filter (DF), sample adaptive offset (SAO), adaptive loop filter (ALF), etc. The NN filter can be any kind of NN filter, such as a convolutional neural network (CNN) filter. In the following discussion, the NN filter can also be referred to as the CNN filter.
[0078] Figure 17 A flowchart of a method 1700 for video processing according to an embodiment of the present disclosure is shown. Method 1700 is implemented during the conversion between a target video block of a video and the bitstream of the video.
[0079] At block 1710, for the conversion between a video unit of a video and the bitstream of the video unit, it is determined whether to apply at least one NN model for neural network (NN) filtering during the process of the video unit. For example, the process can be an RDO process.
[0080] At block 1720, based on the determination, the video unit is processed by applying a process to the video unit.
[0081] At block 1730, a transformation is performed based on the processed video unit. Alternatively or additionally, the transformation may include decoding the video unit from a bitstream. In this way, the effect of reduced distortion due to the NN filter can be considered during the RDO process, thereby improving the coding and decoding performance.
[0082] In some embodiments, at least one NN model is included in the encoder. In some embodiments, the process includes a rate distortion optimization (RDO) process, and at least one NN model is used in the RDO process of the video unit. In some embodiments, at least one NN model is not included in a compliant decoder.
[0083] In some embodiments, at least one NN model is simpler than another NN filter model used for NN filtering in a compliant decoder. For example, at least one NN model may have fewer layers. Alternatively or additionally, at least one NN model may be less complex.
[0084] In some embodiments, at least one NN model is combined with another filter model in the encoder. In some embodiments, at least one NN model is different from the NN filter. In some embodiments, at least one NN model is applied before another filter model. Alternatively, at least one NN model is applied after another filter model.
[0085] In some embodiments, the another filter model includes at least one of the following: a convolutional neural network (CNN) filter model, a deblocking filter, a sample adaptive offset (SAO) filter, an adaptive loop filter (ALF), a cross-component SAO (CCSAO) filter, or a cross-component ALF (CCALF).
[0086] In some embodiments, at least one of at least one NN model or another filter model is applied according to a predefined order or an adaptive order. For example, the predefined order includes sequentially applying a deblocking filter, a CNN filter model, an SAO filter, and an ALF filter.
[0087] In some embodiments, the order of applying at least one of at least one NN model or another filter model depends on at least one of the following: the coding and decoding mode of the video unit, or the coding and decoding statistics of the video unit.
[0088] In some embodiments, whether to utilize at least one of the NN models or another filter model depends on at least one of the following: the coding / decoding mode of the video unit, or the coding / decoding statistics of the video unit. In some embodiments, the method of utilizing at least one of the NN models or another filter model depends on at least one of the following: the coding / decoding mode of the video unit, or the coding / decoding statistics of the video unit. For example, the coding / decoding statistics may include one or more of the following: prediction mode, QP, temporal layer, or slice type.
[0089] In some embodiments, the process includes a mode decision process, and the mode decision process depends on at least one NN filter model. For example, the mode decision process is based on the filtered reconstruction information attributed to at least one NN model.
[0090] In some embodiments, at least one NN model is used in determining the best intra prediction mode of a video unit. For example, the NN filter model can be used (together with the RDO for intra mode selection) in determining the best intra prediction mode.
[0091] In some embodiments, at least one NN model is used in determining the best coding / decoding intra method of a video unit. For example, the NN filter model can be used in determining the best coding / decoding intra method (e.g., whether to apply MIP, ISP, MRL).
[0092] In some embodiments, at least one NN model is used together with the RDO for inter mode selection. For example, the NN filter model can be used together with the RDO for inter mode selection (e.g., whether to use AMVP or skip mode or Merge mode).
[0093] In some embodiments, at least one NN model is used in determining the best coding / decoding inter method of a video unit. For example, the NN filter model can be used in determining the best coding / decoding inter method (e.g., whether to code a block using an affine motion model or a translational motion model, whether to apply MMVD, CIIP, GPM, etc.).
[0094] In some embodiments, at least one NN model is used together with the RDO for segmentation mode selection. For example, the NN filter model can be used together with the RDO for segmentation mode selection (e.g., whether to apply QT, BT, TT, non - partition, etc.).
[0095] In some embodiments, at least one NN model is used together with RDO for transform kernel selection. In some embodiments, at least one NN model is used in determining the best codec method for intra and inter methods including video units. For example, an NN filter model can be used in determining the best codec method for intra and inter methods (e.g., whether to apply intra (MIP, ISP, etc.) or inter (MMVD, AMVP, skip, etc.)).
[0096] In some embodiments, at least one NN model is used whenever distortion is calculated. For example, whenever distortion is calculated, an NN filter model is used.
[0097] In some embodiments, at least one NN model is used whenever distortion is calculated. For example, when calculating distortion using SSE / MSE / SSIM / MS - SSIM / IW - SSIM matrices.
[0098] In some embodiments, at least one NN model is not used when distortion is calculated using a matrix. For example, when distortion is calculated using SAD / SATD matrices.
[0099] In some embodiments, the process includes a mode decision process, and the distortion or cost calculated in the mode decision process is adjusted such that the impact of the NN filtering process is considered. For example, the distortion or cost calculated in the mode decision process (e.g., RDO process) can be corrected such that the impact of the NN filtering process is considered.
[0100] In some embodiments, the distortion or cost is calculated according to a matrix. For example, the matrix includes one or more of the following: SSE, MSE, SSIM, MS - SSIM, IW - SSIM matrices.
[0101] In some embodiments, the process includes an NN filtering process, and method 1700 further includes: applying the NN filtering process to the reconstruction to obtain an NN - filtered reconstruction; and calculating the distortion between the NN - filtered reconstruction and the original samples. For example, instead of using the distortion calculated between the reconstruction before the in - loop filtering method (represented by the unfiltered reconstruction) and the original samples, it is proposed to apply the NN filtering process to the reconstruction to obtain an NN - filtered reconstruction and calculate the distortion between the NN - filtered reconstruction and the original samples.
[0102] In some embodiments, two distortions can be calculated. In this case, one distortion can be between the unfiltered reconstruction and the original samples, and the other distortion can be between the NN - filtered reconstruction and the original samples.
[0103] In some embodiments, the process includes an RDO process, and two distortion functions are called, and the output of the function is set to the true distortion associated with the current mode to be examined during the RDO process.
[0104] In some embodiments, multiple distortions are calculated. In this case, one distortion is between the unfiltered reconstruction and the original samples, and the other distortions are between the filtered reconstruction and the original samples. In some embodiments, the filtered reconstruction samples are filtered by at least one NN model. In some embodiments, the filtered reconstruction samples are filtered by another filter model. In some embodiments, the filtered reconstruction samples are filtered by at least one of at least one NN model or another filter model. Alternatively, a distortion function is called, and the output of the function is set to the true distortion associated with the current mode to be examined during the RDO process.
[0105] In some embodiments, two distortions are calculated. In this case, one distortion is between the filtered reconstruction and the original samples, and the other distortion is between the NN-filtered reconstruction and the original samples. The filtered reconstruction is obtained using another filter, but before at least one NN model. That is, the filtered reconstruction represents the reconstruction obtained using other filters, but before the NN filter.
[0106] In some embodiments, the process includes an RDO process. For example, two distortion functions are called, and the output of the function is set to the true distortion associated with the current mode to be examined during the RDO process.
[0107] In some embodiments, the distortion is first calculated between the unfiltered reconstruction and the original samples and then scaled by a factor. In some embodiments, the factor is a constant between 0 and 1.0. Alternatively, the factor depends on the current mode to be examined during the RDO process. In some embodiments, the factor depends on lambda. In some embodiments, the factor depends on the color component.
[0108] In some embodiments, the process includes an RDO process, and the first filtering process applied to the reconstructed video unit during the RDO process is different from the second filtering process applied during the loop filtering process. Alternatively, the first filtering process is different from the second filtering process applied during the post-processing process.
[0109] In some embodiments, the first filter model in the first filtering process is different from the second filter model in the second filtering process. That is, the filter models can be different.
[0110] In some embodiments, the number of filter models in the first filtering process is different from the number of filter models in the second filtering process. In other words, the number of filter models can be different.
[0111] In some embodiments, the first network structure of the first filtering process is different from the second network structure of the second filtering process. For example, the network structures can be different.
[0112] In some embodiments, the first filtering process during the RDO process is only applied to a sub-region of a video unit. In one example, the filtering process during the RDO process can be only applied to a specific sub-region of a video unit. In some embodiments, the first filtering process is only applied to the boundary samples of a video unit. Alternatively, the first filtering process is only applied to the internal samples of a video unit. In some embodiments, the first filtering process during the RDO process is only applied to the downsampled version of a video unit.
[0113] In some embodiments, the process includes an RDO process, and at least one NN model used in the RDO process is the same as the NN filter model at the decoder. In some embodiments, the number of residual blocks is the same as that of the decoder.
[0114] In some embodiments, at least one NN model used in the RDO process is a simplified version of the NN model used at the decoder. In some embodiments, the first depth of at least one NN model is different from the second depth of the NN model used at the decoder. In one example, the depth of the NN filter model can be different. In some embodiments, the first depth is shallower than the second depth. For example, the NN filter model used in the RDO process can have a shallower depth.
[0115] In some embodiments, the first feature map of at least one NN model is different from the second feature map of the NN model used at the decoder. In some embodiments, at least one NN model in the RDO process has fewer feature maps than the NN model used at the decoder.
[0116] In some embodiments, the number of residual blocks of at least one NN model is different from the number of residual blocks of the NN model at the decoder. In some embodiments, the number of residual blocks of at least one NN model is less than the number of residual blocks of the NN model at the decoder. In some embodiments, the number of residual blocks of at least one NN model is one of the following: 1, 2, 3, 4, 5, 6. In some embodiments, the convolutional kernel of at least one NN model is different from the convolutional kernel of the NN model at the decoder.
[0117] In some embodiments, the process includes an RDO process, and whether and / or how to use at least one NN model in the RDO process depends on at least one of the following: the coding mode of the video unit, or the coding statistics of the video unit. In some embodiments, whether and / or how to use at least one NN model in the RDO process depends on at least one of the following: the prediction mode, quantization step size, temporal layer, slice type, block size of the video unit, color component, or the rate-distortion cost without at least one NN model.
[0118] According to other embodiments of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream generated by a method executed by a device for video processing of a video. The method includes: determining whether to apply at least one NN model for neural network (NN) filtering during a process of a video unit of the video; based on the determination, processing the video unit by applying the process to the video unit; and generating a bitstream based on the processed video unit.
[0119] According to other embodiments of the present disclosure, a method for storing a bitstream of a video is provided. The method includes: determining whether to apply at least one NN model for neural network (NN) filtering during a process of a video unit of the video; based on the determination, processing the video unit by applying the process to the video unit; generating a bitstream based on the processed video unit; and storing the bitstream in a non-transitory computer-readable recording medium.
[0120] The implementations of the present disclosure can be described according to the following items, and the features can be combined in any reasonable manner.
[0121] Item 1. A method for video processing, including: for the conversion between a video unit of a video and the bitstream of the video unit, determining whether to apply at least one neural network (NN) model for NN filtering during the process of the video unit; based on the determination, processing the video unit by applying the process to the video unit; and performing the conversion based on the processed video unit.
[0122] Item 2. The method according to Item 1, wherein the at least one NN model is included in an encoder.
[0123] Item 3. The method according to Item 1, wherein the process includes a rate-distortion optimization (RDO) process, and the at least one NN model is used in the RDO process of the video unit.
[0124] Item 4. The method according to Item 1, wherein the at least one NN model is not included in a compatible decoder.
[0125] Item 5. The method according to Item 1, wherein the at least one NN model is simpler than another NN filter model used for NN filtering in a compatible decoder.
[0126] Item 6. The method according to Item 1, wherein the at least one NN model is combined with another filter model in an encoder.
[0127] Item 7. The method according to Item 6, wherein the at least one NN model is different from an NN filter.
[0128] Item 8. The method according to Item 6, wherein the at least one NN model is applied before the other filter model, or wherein the at least one NN model is applied after the other filter model.
[0129] Item 12. The method according to Item 6, wherein the other filter model includes at least one of the following: a convolutional neural network (CNN) filter model, a deblocking filter, a sample adaptive offset (SAO) filter, an adaptive loop filter (ALF), a cross-component SAO (CCSAO) filter, or a cross-component ALF (CCALF).
[0130] Item 10. The method according to Item 6, wherein at least one of the at least one NN model or the other filter model is applied according to a predefined order or an adaptive order.
[0131] Item 11. The method according to Item 10, wherein the predefined order includes sequentially applying a deblocking filter, a CNN filter model, an SAO filter, and an ALF filter.
[0132] Item 12. The method according to Item 6, wherein the order of applying at least one of the at least one NN model or the other filter model depends on at least one of the following: the coding / decoding mode of the video unit, or the coding / decoding statistics of the video unit.
[0133] Item 13. The method according to Item 6, wherein whether to utilize at least one of the at least one NN model or the other filter model depends on at least one of the following: the coding / decoding mode of the video unit, or the coding / decoding statistics of the video unit.
[0134] Item 14. The method according to Item 6, wherein the method of utilizing at least one of the at least one NN model or the other filter model depends on at least one of the following: the coding / decoding mode of the video unit, or the coding / decoding statistics of the video unit.
[0135] Item 15. The method according to Item 1, wherein the process includes a mode decision process, and the mode decision process depends on the at least one NN filter model.
[0136] Item 16. The method according to Item 15, wherein the mode decision process is based on the filtered reconstruction information attributed to the at least one NN model.
[0137] Item 17. The method according to Item 15, wherein the at least one NN model is used in determining the best intra prediction mode of the video unit.
[0138] Item 18. The method according to Item 15, wherein the at least one NN model is used in determining the best codec intra method of the video unit.
[0139] Item 19. The method according to Item 15, wherein the at least one NN model is used together with the RDO for inter prediction mode selection.
[0140] Item 20. The method according to Item 15, wherein the at least one NN model is used in determining the best codec inter method of the video unit.
[0141] Item 21. The method according to Item 15, wherein the at least one NN model is used together with the RDO for segmentation mode selection.
[0142] Item 22. The method according to Item 15, wherein the at least one NN model is used together with the RDO for transform kernel selection.
[0143] Item 23. The method according to Item 15, wherein the at least one NN model is used in determining the best codec method including the intra and inter methods of the video unit.
[0144] Item 24. The method according to Item 15, wherein the at least one NN model is used whenever distortion is calculated.
[0145] Item 25. The method according to Item 15, wherein the at least one NN model is used whenever distortion is calculated.
[0146] Item 26. The method according to Item 15, wherein the at least one NN model is not used when the distortion is calculated using a matrix.
[0147] Item 27. The method according to Item 1, wherein the process includes a mode decision process, and the distortion or cost calculated in the mode decision process is adjusted such that the influence of the NN filtering process is considered.
[0148] Item 28. The method according to Item 27, wherein the distortion or cost is calculated according to a matrix.
[0149] Item 29. The method according to Item 27, wherein the process includes an NN filtering process, and the method further includes: applying the NN filtering process to the reconstruction to obtain an NN-filtered reconstruction; and calculating the distortion between the NN-filtered reconstruction and the original samples.
[0150] Item 30. The method according to Item 27, further including: calculating two distortions, one distortion being between the unfiltered reconstruction and the original samples, and the other distortion being between the NN-filtered reconstruction and the original samples.
[0151] Item 31. The method according to Item 27, wherein the process includes an RDO process, and a function of two distortions is called, and the output of the function is set to the true distortion associated with the current mode to be examined during the RDO process.
[0152] Item 32. The method according to Item 27, further including: calculating multiple distortions, one distortion being between the unfiltered reconstruction and the original samples, and other distortions being between the filtered reconstructions and the original samples.
[0153] Item 33. The method according to Item 32, wherein the filtered reconstruction samples are filtered by the at least one NN model.
[0154] Item 34. The method according to Item 32, wherein the filtered reconstruction samples are filtered by the other filter model.
[0155] Item 35. The method according to Item 32, wherein the filtered reconstruction samples are filtered by at least one of the at least one NN model or the other filter model.
[0156] Item 36. The method according to Item 32, wherein a function of the distortion is called, and the output of the function is set to the true distortion associated with the current mode to be examined during the RDO process.
[0157] Item 37. The method according to Item 27, further including: calculating two distortions, one distortion being between the filtered reconstruction and the original samples, and the other distortion being between the NN-filtered reconstruction and the original samples, wherein the filtered reconstruction is obtained by the other filter but before the at least one NN model.
[0158] Item 38. The method according to Item 27, wherein the process includes an RDO process, and the two distortion functions are called, and the output of the function is set to the true distortion associated with the current mode to be examined during the RDO process.
[0159] Item 39. The method according to Item 27, wherein the distortion is first calculated between the unfiltered reconstruction and the original samples and then scaled by a factor.
[0160] Item 40. The method according to Item 39, wherein the factor is a constant between 0 and 1.0, or wherein the factor depends on the current mode to be examined during the RDO process, or wherein the factor depends on lambda, or wherein the factor depends on the color component.
[0161] Item 41. The method according to Item 1, wherein the process includes an RDO process, and the first filtering process applied to the reconstructed video unit during the RDO process is different from the second filtering process applied during the loop filtering process, or wherein the first filtering process is different from the second filtering process applied during the post-processing process.
[0162] Item 42. The method according to Item 41, wherein the first filter model in the first filtering process is different from the second filter model in the second filtering process.
[0163] Item 43. The method according to Item 41, wherein the number of filter models in the first filtering process is different from the number of filter models in the second filtering process.
[0164] Item 44. The method according to Item 41, wherein the first network structure of the first filtering process is different from the second network structure of the second filtering process.
[0165] Item 45. The method according to Item 41, wherein the first filtering process during the RDO process is only applied to a sub-region of the video unit.
[0166] Item 46. The method according to Item 45, wherein the first filtering process is only applied to the boundary samples of the video unit.
[0167] Item 47. The method according to Item 45, wherein the first filtering process is only applied to the internal samples of the video unit.
[0168] Item 48. The method according to Item 41, wherein the first filtering process during the RDO process is only applied to the down-sampled version of the video unit.
[0169] Item 49. The method according to Item 1, wherein the process includes an RDO process, and the at least one NN model used in the RDO process is the same as the NN filter model at the decoder.
[0170] Item 50. The method according to Item 49, wherein the number of the residual blocks is the same as that of the decoder.
[0171] Item 51. The method according to Item 1, wherein the at least one NN model used in the RDO process is a simplified version of the NN model used at the decoder.
[0172] Item 52. The method according to Item 51, wherein a first depth of the at least one NN model is different from a second depth of the NN model used at the decoder.
[0173] Item 53. The method according to Item 52, wherein the first depth is shallower than the second depth.
[0174] Item 54. The method according to Item 51, wherein a first feature map of the at least one NN model is different from a second feature map of the NN model used at the decoder.
[0175] Item 55. The method according to Item 54, wherein the at least one NN model in the RDO process has fewer feature maps than the NN model used at the decoder.
[0176] Item 56. The method according to Item 51, wherein the number of the residual blocks of the at least one NN model is different from the number of the residual blocks of the NN model at the decoder.
[0177] Item 57. The method according to Item 56, wherein the number of the residual blocks of the at least one NN model is less than the number of the residual blocks of the NN model at the decoder.
[0178] Item 58. The method according to Item 56, wherein the number of the residual blocks of the at least one NN model is one of the following: 1, 2, 3, 4, 5, 6.
[0179] Item 59. The method according to Item 51, wherein the convolutional kernel of the at least one NN model is different from the convolutional kernel of the NN model at the decoder.
[0180] Item 60. The method according to Item 1, wherein the process includes an RDO process, and whether and / or how to use the at least one NN model in the RDO process depends on at least one of the following: the coding / decoding mode of the video unit, or the coding / decoding statistics of the video unit.
[0181] Item 61. The method according to Item 60, wherein whether and / or how to use the at least one NN model in the RDO process depends on at least one of the following: the prediction mode, the quantization step size, the temporal layer, the slice type, the block size of the video unit, the color component, or the rate-distortion cost without the at least one NN model.
[0182] Item 62. The method according to any one of Items 1 to 61, wherein the conversion includes encoding the video unit into the bitstream.
[0183] Item 63. The method according to any one of Items 1 to 61, wherein the conversion includes decoding the video unit from the bitstream.
[0184] Item 64. An apparatus for video processing, including a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to execute the method according to any one of Items 1 to 63.
[0185] Item 65. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of Items 1 to 63.
[0186] Item 66. A non-transitory computer-readable recording medium storing a bitstream generated by a method executed by an apparatus for video processing for a video, wherein the method includes: determining whether to apply at least one neural network (NN) model for NN filtering during a process of a video unit of the video; based on the determination, processing the video unit by applying the process to the video unit; and generating the bitstream based on the processed video unit.
[0187] Item 67. A method for storing a bitstream of a video, including: determining whether to apply at least one neural network (NN) model for NN filtering during a process of a video unit of the video; based on the determination, processing the video unit by applying the process to the video unit; generating the bitstream based on the processed video unit; and storing the bitstream in a non-transitory computer-readable recording medium. Example device
[0188] Figure 18FIG. shows a block diagram of a computing device 1800 in which various embodiments of the present disclosure may be implemented. The computing device 1800 may be implemented as the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300), or may be included in the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300).
[0189] It should be understood that Figure 18 the computing device 1800 shown in is for illustrative purposes only and does not imply any limitation to the functionality and scope of the embodiments of the present disclosure in any way.
[0190] As Figure 18 shown, the computing device 1800 includes a general-purpose computing device 1800. The computing device 1800 may include at least one or more processors or processing units 1810, a memory 1820, a storage unit 1830, one or more communication units 1840, one or more input devices 1850, and one or more output devices 1860.
[0191] In some embodiments, the computing device 1800 may be implemented as any user terminal or server terminal with computing capabilities. The server terminal may be a server provided by a service provider, a large computing device, etc. The user terminal may be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistant (PDA), audio / video players, digital cameras / cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, game devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is conceivable that the computing device 1800 may support any type of interface to the user (such as "wearable" circuitry, etc.).
[0192] The processing unit 1810 may be a physical processor or a virtual processor, and may implement various processes based on programs stored in the memory 1820. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 1800. The processing unit 1810 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0193] The computing device 1800 generally includes various computer storage media. Such media can be any media accessible to the computing device 1800, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 1820 can be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. The storage unit 1830 can be any removable or non-removable media and can include machine-readable media such as a memory, flash drive, magnetic disk, or other media that can be used to store information and / or data and can be accessed within the computing device 1800.
[0194] The computing device 1800 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although not shown in Figure 18 , a disk drive for reading from and / or writing to a removable non-volatile magnetic disk and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk may be provided. In such a case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0195] The communication unit 1840 communicates with another computing device via a communication medium. Additionally, the functions of the components in the computing device 1800 may be implemented by a single computing cluster or multiple computer machines, which may communicate via a communication connection. Thus, the computing device 1800 may operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs), or other general network nodes.
[0196] The input device 1850 can be one or more of various input devices such as a mouse, keyboard, trackball, voice input device, and so on. The output device 1860 can be one or more of various output devices such as a display, speaker, printer, and so on. With the aid of the communication unit 1840, the computing device 1800 can also communicate with one or more external devices (not shown), such as storage devices and display devices, the computing device 1800 can also communicate with one or more devices that enable a user to interact with the computing device 1800, or if needed, the computing device 1800 can also communicate with any device (e.g., network card, modem, etc.) that enables the computing device 1800 to communicate with one or more other computing devices. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0197] In some embodiments, some or all components of computing device 1800 may also be arranged in a cloud computing architecture rather than integrated in a single device. In a cloud computing architecture, components may be provided remotely and work together to implement the functions described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, which do not require an end user to be aware of the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing uses suitable protocols to provide services via a wide area network such as the Internet. For example, a cloud computing provider provides an application via a wide area network, and the application can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on a server at a remote location. Computing resources in a cloud computing environment may be consolidated or distributed at the locations of remote data centers. The cloud computing infrastructure may provide services through a shared data center, although to a user, they appear as a single access point. Thus, the cloud computing architecture may be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein may be provided by a conventional server or installed directly or otherwise on a client device.
[0198] In an embodiment of the present disclosure, computing device 1800 may be used to implement video encoding / decoding. Memory 1820 may include one or more video codec modules 1825 having one or more program instructions. These modules are accessible and executable by processing unit 1810 to perform the functions of the various embodiments described herein.
[0199] In an example embodiment of performing video encoding, input device 1850 may receive video data as input 1870 to be encoded. The video data may be processed, for example, by video codec module 1825 to generate an encoded bitstream. The encoded bitstream may be provided as output 1880 via output device 1860.
[0200] In an example embodiment of performing video decoding, input device 1850 may receive the encoded bitstream as input 1870. The encoded bitstream may be processed, for example, by video codec module 1825 to generate decoded video data. The decoded video data may be provided as output 1880 via output device 1860.
[0201] Although the present disclosure has been specifically shown and described with reference to preferred embodiments of the present disclosure, those skilled in the art will understand that various changes may be made in form and detail without departing from the spirit and scope of the present application as defined by the appended claims. These variations are intended to be covered by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for video processing, comprising: determining whether to apply at least one neural network (NN) model for NN filtering during a process of the video unit, for a conversion between a video unit of a video and a bitstream of the video unit; processing the video unit by applying the process to the video unit based on the determination; and performing the conversion based on the processed video unit.
2. The method according to claim 1, wherein the at least one NN model is included in an encoder.
3. The method according to claim 1, wherein the process includes a rate distortion optimization (RDO) process, and the at least one NN model is used in the RDO process of the video unit.
4. The method according to claim 1, wherein the at least one NN model is not included in a compliant decoder.
5. The method according to claim 1, wherein the at least one NN model is simpler than another NN filter model used for NN filtering in a compliant decoder.
6. The method according to claim 1, wherein the at least one NN model is combined with another filter model in an encoder.
7. The method according to claim 6, wherein the at least one NN model is different from an NN filter.
8. The method according to claim 6, wherein the at least one NN model is applied before the another filter model, or wherein the at least one NN model is applied after the another filter model.
9. The method according to claim 6, wherein the another filter model includes at least one of the following: a convolutional neural network (CNN) filter model, a deblocking filter, a sample adaptive offset (SAO) filter, an adaptive loop filter (ALF), a cross-component SAO (CCSAO) filter, or a cross-component ALF (CCALF).
10. The method according to claim 6, wherein at least one of the at least one NN model or the another filter model is applied according to a predefined order or an adaptive order.
11. The method according to claim 10, wherein the predefined order includes sequentially applying a deblocking filter, a CNN filter model, a SAO filter, and an ALF filter.
12. The method according to claim 6, wherein an order of applying at least one of the at least one NN model or the another filter model depends on at least one of the following: a coding / decoding mode of the video unit, or coding / decoding statistics of the video unit.
13. The method according to claim 6, wherein whether to utilize at least one of the at least one NN model or the another filter model depends on at least one of the following: a coding / decoding mode of the video unit, or coding / decoding statistics of the video unit.
14. The method according to claim 6, wherein a method of utilizing at least one of the at least one NN model or the another filter model depends on at least one of the following: a coding / decoding mode of the video unit, or coding / decoding statistics of the video unit.
15. The method according to claim 1, wherein the process includes a mode decision process, and the mode decision process depends on the at least one NN filter model.
16. The method according to claim 15, wherein the mode decision process is based on the filtered reconstruction information attributed to the at least one NN model.
17. The method according to claim 15, wherein the at least one NN model is used when determining the best intra prediction mode of the video unit.
18. The method according to claim 15, wherein the at least one NN model is used when determining the best codec intra method of the video unit.
19. The method according to claim 15, wherein the at least one NN model is used together with the RDO for inter mode selection.
20. The method according to claim 15, wherein the at least one NN model is used when determining the best codec inter method of the video unit.
21. The method according to claim 15, wherein the at least one NN model is used together with the RDO for segmentation mode selection.
22. The method according to claim 15, wherein the at least one NN model is used together with the RDO for transform kernel selection.
23. The method according to claim 15, wherein the at least one NN model is used when determining the best codec method including the intra and inter methods of the video unit.
24. The method according to claim 15, wherein the at least one NN model is used whenever distortion is calculated.
25. The method according to claim 15, wherein the at least one NN model is used whenever distortion is calculated.
26. The method according to claim 15, wherein the at least one NN model is not used when the distortion is calculated using a matrix.
27. The method according to claim 1, wherein the process includes a mode decision process, and the distortion or cost calculated in the mode decision process is adjusted such that the impact of the NN filtering process is considered.
28. The method according to claim 27, wherein the distortion or cost is calculated according to a matrix.
29. The method according to claim 27, wherein the process includes an NN filtering process, and the method further includes: applying the NN filtering process to the reconstruction to obtain an NN-filtered reconstruction; and calculating the distortion between the NN-filtered reconstruction and the original samples.
30. The method according to claim 27, further includes: calculating two distortions, one between the unfiltered reconstruction and the original samples, and the other between the NN-filtered reconstruction and the original samples.
31. The method according to claim 27, wherein the process includes an RDO process, and a function of two distortions is called, and the output of the function is set as the true distortion associated with the current mode to be examined during the RDO process.
32. The method according to claim 27, further includes: Calculate multiple distortions, where one distortion is between the unfiltered reconstruction and the original samples, and the other distortions are between the filtered reconstruction and the original samples.
33. The method according to claim 32, wherein the filtered reconstruction samples are filtered by the at least one NN model.
34. The method according to claim 32, wherein the filtered reconstruction samples are filtered by the other filter model.
35. The method according to claim 32, wherein the filtered reconstruction samples are filtered by at least one of the at least one NN model or the other filter model.
36. The method according to claim 32, wherein the function of the distortion is called, and the output of the function is set as the true distortion associated with the current mode to be examined during the RDO process.
37. The method according to claim 27, further comprising: Calculate two distortions, where one distortion is between the filtered reconstruction and the original samples, and the other distortion is between the NN-filtered reconstruction and the original samples, wherein the filtered reconstruction is obtained by the other filter, but before the at least one NN model.
38. The method according to claim 27, wherein the process includes an RDO process, and the function of the two distortions is called, and the output of the function is set as the true distortion associated with the current mode to be examined during the RDO process.
39. The method according to claim 27, wherein the distortion is first calculated between the unfiltered reconstruction and the original samples, and then scaled by a factor.
40. The method according to claim 39, wherein the factor is a constant between 0 and 1.0, or wherein the factor depends on the current mode to be examined during the RDO process, or wherein the factor depends on lambda, or wherein the factor depends on the color component.
41. The method according to claim 1, wherein the process includes an RDO process, and the first filtering process applied to the reconstructed video unit during the RDO process is different from the second filtering process applied during the loop filtering process, or wherein the first filtering process is different from the second filtering process applied during the post-processing process.
42. The method according to claim 41, wherein the first filter model in the first filtering process is different from the second filter model in the second filtering process.
43. The method according to claim 41, wherein the number of filter models in the first filtering process is different from the number of filter models in the second filtering process.
44. The method according to claim 41, wherein the first network structure of the first filtering process is different from the second network structure of the second filtering process.
45. The method according to claim 41, wherein the first filtering process during the RDO process is only applied to a sub-region of the video unit.
46. The method according to claim 45, wherein the first filtering process is only applied to the boundary samples of the video unit.
47. The method according to claim 45, wherein the first filtering process is only applied to the intra-samples of the video unit.
48. The method according to claim 41, wherein the first filtering process during the RDO process is only applied to the downsampled version of the video unit.
49. The method according to claim 1, wherein the process includes an RDO process, and the at least one NN model used in the RDO process is the same as the NN filter model at the decoder.
50. The method according to claim 49, wherein the number of residual blocks is the same as that of the decoder.
51. The method according to claim 1, wherein the at least one NN model used in the RDO process is a simplified version of the NN model used at the decoder.
52. The method according to claim 51, wherein a first depth of the at least one NN model is different from a second depth of the NN model used at the decoder.
53. The method according to claim 52, wherein the first depth is shallower than the second depth.
54. The method according to claim 51, wherein a first feature map of the at least one NN model is different from a second feature map of the NN model used at the decoder.
55. The method according to claim 54, wherein the at least one NN model in the RDO process has fewer feature maps than the NN model used at the decoder.
56. The method according to claim 51, wherein the number of residual blocks of the at least one NN model is different from the number of residual blocks of the NN model at the decoder.
57. The method according to claim 56, wherein the number of residual blocks of the at least one NN model is less than the number of residual blocks of the NN model at the decoder.
58. The method according to claim 56, wherein the number of the residual blocks of the at least one NN model is one of the following: 1, 2, 3, 4, 5, 6.
59. The method according to claim 51, wherein the convolution kernel of the at least one NN model is different from the convolution kernel of the NN model at the decoder.
60. The method according to claim 1, wherein the process includes an RDO process, and whether and / or how to use the at least one NN model in the RDO process depends on at least one of the following: the coding / decoding mode of the video unit, or the coding / decoding statistics of the video unit.
61. The method according to claim 60, wherein whether and / or how to use the at least one NN model in the RDO process depends on at least one of the following: prediction mode, quantization step size, temporal layer, slice type, the block size of the video unit, color component, or the rate-distortion cost without the at least one NN model.
62. The method according to any one of claims 1 to 61, wherein the transformation includes encoding the video unit into the bitstream.
63. The method according to any one of claims 1 to 61, wherein the conversion includes decoding the video unit from the bitstream.
64. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 63.
65. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 63.
66. A non-transitory computer-readable recording medium storing a bitstream generated by a method executed by an apparatus for video processing for a video, wherein the method comprises: determining whether to apply at least one NN model for neural network (NN) filtering during a process of video units of the video; processing the video unit by applying the process to the video unit based on the determination; and generating the bitstream based on the processed video unit.
67. A method for storing a bitstream of a video, comprising: determining whether to apply at least one NN model for neural network (NN) filtering during a process of video units of the video; processing the video unit by applying the process to the video unit based on the determination; generating the bitstream based on the processed video unit; and storing the bitstream in a non-transitory computer-readable recording medium.