Method and device for video processing and medium

By introducing a neural network-based loop filter into video encoding and decoding, the problem of low encoding and decoding efficiency in existing technologies is solved, and more efficient video encoding and decoding processing is achieved.

CN120937375APending Publication Date: 2025-11-11DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480025447.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-13
Filing Date
2024-04-12
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies have low encoding and decoding efficiency, making it difficult to meet the needs of high-efficiency encoding and decoding.

Method used

Video processing is performed using a loop filter based on a neural network (NN). By limiting the number of candidate parameters during parameter selection and feeding the residual samples between the current video unit and the bitstream into the loop filter, the filter is applied to the rate-distortion optimization process to filter different types of stripes and color components.

Benefits of technology

It improves the effectiveness and efficiency of video encoding and decoding, and enhances the performance of encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120937375A_ABST
    Figure CN120937375A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. In the method, a conversion between a current video unit of a video and a bitstream of the video is performed. A neural network (NN)-based loop filter is applied to the conversion. During the parameter selection process for the NN-based loop filter, the number of parameter candidates in the parameter candidate list for the NN-based loop filter is less than a threshold number.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure generally relate to video encoding and decoding technologies, and more specifically, to neural network (NN) based filters for video encoding and decoding. Background Technology

[0002] Today, digital video capabilities are being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Multi-Functional Video Codec (VVC) standard. However, the encoding and decoding efficiency of conventional video codec technologies is usually very low, which is undesirable. Summary of the Invention

[0003] Embodiments of this disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is proposed. The method includes: performing a conversion between a current video unit and a bitstream of the video, wherein a neural network (NN)-based loop filter is applied to the conversion, and wherein during the parameter selection process for the NN-based loop filter, the number of parameter candidates in the parameter candidate list for the NN-based loop filter is less than a threshold number. The method according to the first aspect of this disclosure improves encoding / decoding efficiency and encoding / decoding effectiveness.

[0005] In a second aspect, another method for video processing is proposed. This method includes: performing a conversion between a current video unit and a bitstream of the video, wherein a neural network (NN)-based loop filter is applied to the conversion, and wherein at least one residual sample between at least one reconstructed sample and at least one predicted sample of the current video unit is fed into the NN-based loop filter. The method according to the second aspect of this disclosure improves encoding / decoding efficiency and encoding / decoding effectiveness.

[0006] In a third aspect, another method for video processing is proposed. This method includes performing a conversion between the current video unit and the video bitstream, wherein a unified neural network (NN)-based filter is applied to a rate-distortion optimization (RDO) process, and wherein the unified NN-based filter is applied to different types of stripes and / or different color components. The method according to the third aspect of this disclosure improves encoding / decoding efficiency and encoding / decoding effectiveness.

[0007] In a fourth aspect, an apparatus for video processing is proposed. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform a method according to the first, second, or third aspect of this disclosure.

[0008] In a fifth aspect, a non-transitory computer-readable storage medium is provided. This non-transitory computer-readable storage medium stores instructions that cause a processor to execute a method according to the first, second, or third aspect of this disclosure.

[0009] In a sixth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: generating the bitstream of video, wherein a neural network (NN)-based loop filter is applied to the generation, and wherein during the parameter selection process for the NN-based loop filter, the number of parameter candidates in the parameter candidate list for the NN-based loop filter is less than a threshold number.

[0010] In a seventh aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: generating the bitstream of video, wherein a neural network (NN)-based loop filter is applied to the generation, and wherein during the parameter selection process for the NN-based loop filter, the number of parameter candidates in the parameter candidate list for the NN-based loop filter is less than a threshold number; and storing the bitstream in the non-transitory computer-readable recording medium.

[0011] In an eighth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: generating the bitstream of video, wherein a neural network (NN)-based loop filter is applied to the generation, and wherein at least one residual sample between at least one reconstructed sample and at least one predicted sample of a current video unit is fed into the NN-based loop filter.

[0012] In a ninth aspect, a method for storing a bitstream of video is proposed. The method includes: generating a bitstream of video, wherein a neural network (NN)-based loop filter is applied to the generation, and wherein at least one residual sample between at least one reconstructed sample and at least one predicted sample of a current video unit is fed into the NN-based loop filter; and storing the bitstream in a non-transitory computer-readable recording medium.

[0013] In a tenth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: generating the bitstream of video, wherein a unified neural network (NN)-based filter is applied to a rate-distortion optimization (RDO) process, and wherein the unified NN-based filter is applied to different types of stripes and / or different color components.

[0014] In the eleventh aspect, a method for storing a bitstream of video is proposed. The method includes: generating a bitstream of video, wherein a unified neural network (NN)-based filter is applied to a rate-distortion optimization (RDO) process, and wherein the unified NN-based filter is applied to different types of stripes and / or different color components; and storing the bitstream in a non-transitory computer-readable recording medium.

[0015] This summary aims to present, in a simplified form, the selected concepts further described below in the detailed embodiments. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description

[0016] The above and other objects, features and advantages of exemplary embodiments of the present disclosure will become clearer from the following detailed description with reference to the accompanying drawings, in which the same reference numerals generally refer to the same parts.

[0017] Figure 1 A block diagram of an example video codec system according to some embodiments of the present disclosure is shown; Figure 2 A block diagram of a first example video encoder according to some embodiments of the present disclosure is shown; Figure 3 A block diagram of an example video decoder according to some embodiments of the present disclosure is shown; Figure 4 A schematic diagram illustrating an example of raster scan strip segmentation for displaying an image is shown; Figure 5 A schematic diagram showing an example of rectangular strip segmentation for displaying an image is provided. Figure 6 A schematic diagram showing an example of an image divided into pieces, bricks, and rectangular strips is provided. Figure 7A An example diagram showing the CTB across the bottom image boundary is shown; Figure 7B An example image showing the CTB across the right edge of the image is shown; Figure 7CAn example image showing the CTB across the bottom right edge of the image is shown; Figure 8 A schematic diagram showing an example of an encoder block diagram is provided. Figure 9 The preprocessing unit and the postprocessing unit are shown; Figure 10 The architecture of the CNN in filter set 0 is shown; Figure 11 The implementation of the CNN in filter set 0 is shown; Figure 12 Encoder optimization 2 is shown; Figures 13A to 13C The architectures of the CNNs in filter set 1 are shown respectively; Figure 14 The time-domain loop filter is shown. Only the header is shown; the other parts are shown separately. Figure 13B and Figure 13C Keep them the same. {Col 0, Col 1} refers to the cosine samples from the first image in the two reference image lists; Figure 15A The parameter selection on the encoder side is shown; Figure 15B The parameter selection on the decoder side is shown; Figure 16 This demonstrates how intra-frame prediction patterns based on neural networks can be used to... Context of surrounding reference samples Predict the current situation piece A schematic diagram. Here, ; Figure 17 It shows that it will revolve around the current piece Context of reference sample Decomposed into usable reference samples and unavailable reference points .here, In the case shown, the number of reference samples cannot reach its maximum value; Figure 18 It shows the current Intra-predictive mode signaling for the luminance CB (highlighted by a dashed line in orange). The coordinates of the top-left pixel of this CB are... The binary value of nnFlag is shown in coarse gray. Here, , , ,and ; Figure 19 A flowchart of a method for video processing according to some embodiments of the present disclosure is shown; Figure 20 A flowchart is shown for another method for video processing according to some embodiments of the present disclosure; Figure 21 A flowchart of another method for video processing according to some embodiments of the present disclosure is shown; and Figure 22 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.

[0018] In all accompanying drawings, the same or similar reference numerals usually refer to the same or similar elements. Detailed Implementation

[0019] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.

[0020] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0021] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, it is claimed that, whether explicitly described or not, such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.

[0022] It should be understood that although the terms “first” and “second”, etc., may be used herein to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0023] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.

[0024] Example Environment Figure 1 This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0025] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.

[0026] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded images and associated data. An encoded image is an encoded representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to destination device 120 via network 130A through I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.

[0027] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.

[0028] The video encoder 114 and the video decoder 124 can operate according to video compression standards such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or future standards.

[0029] Figure 2 This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.

[0030] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0031] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.

[0032] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, in which at least one reference picture is the picture in which the current video block is located.

[0033] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2The examples are shown separately.

[0034] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0035] The mode selection unit 203 can, for example, select one of several coding modes (intra-frame coding or inter-frame coding) based on the error result, and provide the resulting intra-frame coded or inter-frame coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select a combined intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).

[0036] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.

[0037] Motion estimation unit 204 and motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks independent of macroblocks within the same image.

[0038] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0039] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search reference images in list 0 to find a reference video block for the current video block, and can also search reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference images containing multiple reference video blocks in lists 0 and 1, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. Motion estimation unit 204 can output the multiple reference indices and multiple motion vectors of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.

[0040] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0041] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0042] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0043] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.

[0044] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0045] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.

[0046] In other examples, such as in skip mode, residual data for the current video block may not exist, and residual generation unit 207 may not perform a subtraction operation.

[0047] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0048] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0049] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in the buffer 213.

[0050] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0051] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0052] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.

[0053] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0054] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.

[0055] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which involves deriving several most likely candidates based on data from neighboring PBs and reference pictures. Motion information typically includes horizontal motion vector displacement values ​​and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially or temporally neighboring blocks.

[0056] The motion compensation unit 302 can generate motion compensation blocks, possibly by performing interpolation based on an interpolation filter. Identifiers for interpolation filters used with sub-pixel precision can be included in the syntax elements.

[0057] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of a video block to calculate the interpolated values ​​of sub-integer pixels for the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate a prediction block.

[0058] Motion compensation unit 302 may use at least some of the syntax information to determine the size of the blocks used to encode the encoded video sequence (multiple frames) and / or (multiple stripes), segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some respects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be an entire image or a region of an image.

[0059] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially neighboring blocks. Dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.

[0060] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.

[0061] Some exemplary embodiments of this disclosure will be described in detail below. It should be noted that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section. Furthermore, although some embodiments are described with reference to multi-function video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Furthermore, although some embodiments describe video encoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. Additionally, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at different compression bitrates.

[0062] 1. Brief Overview This disclosure relates to video codec technology. Specifically, it relates to loop filters in image / video codecs. It can be applied to existing video codec standards, such as High Efficiency Video Codec (HEVC), Multi-Functional Video Codec (VVC), or standards yet to be completed (e.g., AVS3). It can also be applied to future video codec standards or video codecs, or used as a post-processing method outside the encoding / decoding process.

[0063] 2. Introduction Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Vision, and the two organizations jointly produced H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / HEVC standards. H.262, in particular, is based on a hybrid video codec architecture utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Group (JVET) was created between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) to work on a VVC standard with a 50% bitrate reduction compared to HEVC. VVC version 1 was completed in July 2020.

[0064] The Joint Video Exploration Team (JVET) of ITU-T VCEG and ISO / IEC MPEG is exploring potential neural network video codec technologies beyond the capabilities of VVC. This exploration is known as Neural Network-Based Video Coding (NNVC). Neural network-based (NN-based) codec tools are designed to enhance or replace legacy modules in existing VVC designs. The implementation of NN-based tools in NNVC 4 is based on a Small Ad hoc Deep Learning (SADL) library.

[0065] 2.1. NNVC-4.0 reference software is provided to demonstrate reference implementations of encoding and decoding processes, as well as training methods for neural network-based video encoding and decoding explored in JVET. Definition of a video unit. The image is divided into one or more slice rows and one or more slice columns. A slice is a CTU sequence that covers a rectangular area of ​​the image.

[0066] The sheet is divided into one or more blocks, each block consisting of a certain number of CTU rows within the sheet.

[0067] A slice that is not divided into multiple tiles is also called a tile. However, a tile that is a proper subset of a slice is not called a slice.

[0068] A strip contains a certain number of slices of a picture or a certain number of tiles of a slice.

[0069] Two strip modes are supported, namely the raster scan strip mode and the rectangular strip mode. In the raster scan strip mode, a strip contains a sequence of slices in the slice raster scan of the picture. In the rectangular strip mode, a strip contains a certain number of tiles of the picture, and these tiles together form a rectangular area of the picture. The tiles within a rectangular strip are in the order of the tile raster scan of the strip.

[0070] Figure 4 An example of the raster scan strip partitioning of a picture is shown, where the picture is divided into 12 slices and 3 raster scan strips. Figure 4 A picture with 18×12 luminance CTUs is shown, which is divided into 12 slices and 3 raster scan strips (informative).

[0071] Figure 5 An example of the rectangular strip partitioning of a picture is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips. Figure 5 A picture with 18×12 luminance CTUs is shown, which is divided into 24 slices and 9 rectangular strips (informative).

[0072] Figure 6 An example of a picture divided into slices, tiles, and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows), 11 tiles (the upper left slice contains 1 tile, the upper right slice contains 5 tiles, the lower left slice contains 2 tiles, and the lower right slice contains 3 tiles), and 4 rectangular strips. Figure 6 A picture divided into 4 slices, 11 tiles, and 4 rectangular strips is shown (informative).

[0073] 2.1.1. CTU / CTB Size In VVC, the CTU size signaled by the syntax element log2_ctu_size_minus2 in the SPS can be as small as 4×4.

[0074] 2.1.2. CTUs in a Picture Assume that the CTB / LCU size is indicated by M×N (usually M equals N, as defined in HEVC / VVC), and for a CTB located at the boundary of a picture (or a slice or a strip or other types of boundaries, taking the picture boundary as an example), K×L samples are within the picture boundary, where K < M or L < N. For Figures 7A to 7CThose CTBs depicted in [description] have CTB sizes still equal to M×N. However, the bottom / right boundaries of the CTBs are outside the picture.

[0075] Figure 7A An example diagram showing a CTB (where K = M, L < N) spanning the bottom picture boundary is shown.

[0076] Figure 7B An example diagram showing a CTB (where K < M, L = N) spanning the right picture boundary is shown.

[0077] Figure 7C An example diagram showing a CTB (where K < M, L < N) spanning the bottom - right picture boundary is shown.

[0078] 2.2. Encoding and decoding streams of typical video codecs Figure 8 An example of the encoder block diagram of VVC is shown, which includes three loop - filtering blocks: the de - block filter (DF), sample - adaptive offset (SAO), and ALF. Different from D which uses predefined filters, SAO and ALF utilize the original samples of the current picture, reduce the mean - square error between the original samples and the reconstructed samples by adding an offset and by applying a finite - impulse response (FIR) filter respectively, where the encoded - decoded side information signals the offset and filter coefficients. ALF is located in the last processing stage of each picture and can be regarded as a tool to attempt to capture and repair artifacts caused by previous stages. Figure 8 A schematic diagram 800 showing an example of the encoder block diagram is shown.

[0079] 2.3. Neural - network - based video encoding and decoding (NNVC) 2.3.1. Neural - network - based loop - filter set 0 2.3.1.1. Pre - processing and post - processing of chrominance In filter set 0, filters with a single model are designed to process three components. Due to the different resolutions of luminance and chrominance, pre - processing steps and post - processing steps are introduced to upsample and downsample the chrominance component respectively, as Figure 9 shown. In the resampling process, the nearest - neighbor interpolation method is used. Figure 9 An example of the pre - processing unit and the post - processing unit is shown.

[0080] 2.3.1.2. Neural network The network structure of the CNN filter is as Figure 10As shown. In addition to the reconstructed image (Rec_yuv), additional side information, such as the predicted image (Pred_yuv), strip QP, base QP, and strip type, is also fed into the network. In the residual block (Resblock), the number of channels is first increased before the activation layer and then decreased after the activation layer. Specifically, K and M are set to 64 and 160, respectively, and the number of residual blocks is set to 32. Figure 10 The architecture of the CNN in filter set 0 is shown.

[0081] 2.3.1.3. Combination with traditional filters Figure 11 This illustrates the implementation of a CNN in filter set 0. For example... Figure 11 As shown, the reconstructed samples prior to DBK are fed into a CNN-based filter (CNNLF), and then the final filtered samples are generated by mixing the results of CNNLF and SAO. This mixing process can be briefly described as follows: .

[0082] For the blending weight, there are four candidates: 1, 0.75, 0.5, and an adaptive weight. The adaptive weight is derived based on the least squares method. If the adaptive weight is selected, the blending weight is transmitted via signal transmission for each color component in the strip header.

[0083] 2.3.1.4. Mode Selection CNN filters can be turned on / off at both the CTU and strip levels. For each enable type, there are four hybrid modes. Therefore, nine modes need to be evaluated at the encoder via RDO. The final selected mode will be transmitted via signal transmission in the strip header.

[0084] Table 1. Parameter selection for filter set 0

[0085] 2.3.1.5. Basic QP Adjustment like Figure 10 As shown, the base QP is fed into the CNN filter. To improve adaptability, an offset can be added to the base QP at the strip level (the adjusted base QP is used as input to the NN filter). The offset candidates are {-5, 5}. For example, given an offset of -5, for the current strip, the actual input base QP to the filter becomes (base QP - 5).

[0086] Encoder method The proposed encoder filters only one of every four CTUs during the process of selecting the optimal base QP offset, thus saving encoding time. For example... Figure 12As shown, only the shaded CTU is considered to calculate the distortion using different BaseQP candidates {BaseQP, BaseQP-5, BaseQP+5}. After selecting the candidate with the minimum cost, the encoder applies the optimal offset to the remaining CTUs ( Figure 12 The non-shadowed CTU in the image is filtered. Figure 12 Encoder optimization 2 is shown.

[0087] 2.3.1.6. Encoder optimization only To more accurately estimate the rate-distortion (RD) cost of utilizing the integrated NN-based loop filter, the encoder-only NN filter is involved in the segmentation decision process. In the segmentation mode decision, the distortion between the NN-filtered samples and the original samples is calculated, and the optimal segmentation mode is then selected based on the calculated distortion to make the segmentation decision more accurate. To reduce complexity, only a small number of residual blocks are used in the network structure (see Section 2.3.1.2). The NN filter in the RDO process is implemented using SADL with integer (int) 16 precision. This encoder-only NN tool is disabled by default.

[0088] 2.3.1.7. Reasoning Details SADL (see Section 2.3.4) is used to perform inference for CNN filters. It supports both floating-point and fixed-point implementations. In the fixed-point implementation, static quantization is used to represent both weights and feature maps with integer precision of 16. Network information during the inference phase is shown in Table 2.

[0089] Table 2. Network information for filter set 0 during the inference phase

[0090] 2.3.2. Loop Filter Set Based on Neural Networks 1 2.3.2.1. Neural Network for Luminance Component Filter set 1 contains two conventional networks, one for luminance and one for chrominance.

[0091] The input to the luminance network includes reconstructed luminance samples (rec), predicted luminance samples (pred), boundary intensities (bs), QP, and block type (IPB). The number of feature maps and residual blocks are set to 96 and 8, respectively. The structure of the luminance network is as follows: Figures 13A to 13C As shown.

[0092] Figures 13A to 13C The architectures of the CNN in filter set 1 are shown respectively. Figure 13A The head of the luminance network is shown. The inputs are combined to form the input y to the next part of the network. Figure 13BThe k-th residual block (k = 0..7) is shown. The output y of the header is fed into the first residual block (input z0 = y). The output z1 is then fed into another such residual block. Figure 13C This shows that the output of the last residual block is fed into this final part of the network.

[0093] 2.3.2.2. Neural Networks for Chromaticity Components Luminance information is used as additional input to the loop filter for chroma. Considering that luminance has a higher resolution than chroma in the YUV 4:2:0 format, features are first extracted from luminance and chroma separately. The luminance features are then downsampled and concatenated with the chroma features. The input to the chroma network includes reconstructed luminance samples (recY), reconstructed chroma samples (recUV), predicted chroma samples (predUV), boundary intensity (bsUV), and QP. Regarding the network backbone, the chroma component uses the same backbone as the luminance component.

[0094] 2.3.2.3. Time-domain filter Filter set 1 contains additional loop filters, i.e., time-domain filters, which extract co-occurring blocks from the first image in the two reference image lists to improve performance. For example... Figure 14 As shown, two co-located blocks are directly concatenated and fed into the network. When temporal filtering is enabled, the temporal filter is applied to the luminance component of the image in the three highest temporal layers, while the regular luminance and chrominance filters are used for other purposes. By default, temporal filtering is disabled.

[0095] Figure 14 The time-domain loop filter is shown. Only the header is shown; the other parts are shown separately. Figures 13B to 13C Keep them the same. {Col 0, Col 1} refers to the cosine samples from the first image in the two reference image lists.

[0096] 2.3.2.4. Adaptive Inference Granularity The granularity of filter determination and parameter selection depends on the resolution and QP. Given a higher resolution and a larger QP, determination and selection will be performed over a larger area.

[0097] 2.3.2.5. Parameter Selection Each strip or block can determine whether to apply a CNN-based filter. Once it's determined that a CNN-based filter will be applied to a strip / block, it can be further decided which conditional parameter to choose from a candidate list including three candidates derived from the QP. Let the sequence-level QP be denoted as q, and the candidate list include the conditional parameters {Param_1, Param_2, Param_3}. For lower temporal layers, Param_1 = q, Param_2 = q 5. Param_3 = q 10. For higher time-domain layers, Param_1 = q, Param_2 = q 5. Param_3 = q 5. In other words, the third candidate is different across different time-domain layers.

[0098] The selection process is based on the rate-distortion cost on the encoder side. If necessary, the on / off control indication and condition parameter index are transmitted via signaling in the bitstream. Figure 15A and Figure 15B The parameter selection diagrams for the encoder and decoder sides are shown. All blocks in the current frame need to be processed first using three conditional parameters. Then, five costs, Cost_0, ..., Cost_5, are calculated and compared with each other to achieve optimal rate-distortion performance. In Cost_0, CNN-based filters are disabled for all blocks. In Cost_i, {i = 1, 2, 3}, the parameter Param_i is used for all blocks. In Cost_4, different blocks can be biased with different parameters, and for each block, information about whether to use CNN-based filters or which parameters to use is transmitted via signal transmission. On the decoder side, whether to use CNN-based filters or which parameters to use for a block is based on Param_Id parsed from the bitstream, such as... Figure 15B As shown.

[0099] Note that for the intra-frame configuration, parameter selection is disabled, while filter on / off control is still retained. Shared conditional parameters are used for both chroma components to alleviate the worst-case burden on the decoder side. Additionally, the maximum number of conditional parameter candidates can be specified on the encoder side.

[0100] 2.3.2.6. Residual Scaling When the NN filter is applied to the reconstructed image, a scaling factor is derived and transmitted via signal transmission for each color component in the strip header. The derivation is based on the least squares method. The difference between the input sample and the NN-filtered sample (residual) is scaled by the scaling factor before being added to the input sample.

[0101] 2.3.2.7. Combination with deblocking filters To enable the combination with deblocking, the input samples used in residual scaling are the output of the deblocking filter. The residual scaling process is shown below, where... These refer to the outputs of NN filtering and deblocking filtering, respectively.

[0102] .

[0103] 2.3.2.8. Encoder optimization only Unlike NNVC-2.0, EncDbOpt also enables AI configuration.

[0104] To better estimate the rate-distortion (RD) cost when using neural network (NN) filters, the proposed encoder introduces NN-based filtering into the rate-distortion optimization (RDO) process for segmentation mode selection. Specifically, the refinement distortion is calculated by comparing the NN-filtered samples with the original samples. The segmentation mode with the minimum rate-refinement distortion cost is selected as the optimal segmentation mode. To reduce complexity, several fast algorithms are applied. First, the NN model is simplified by using a smaller number of residual blocks. Second, parameter selection is not allowed for the NN filtering in the RDO process. Third, the proposed technique is only applied to encoder-decoder units with a height and width no greater than 64. The NN filters used in the RDO process are also implemented using SADL based on fixed-point computation. This NN-based encoder-only method is disabled by default.

[0105] 2.3.2.9. Reasoning Details SADL (see Section 2.3.4) is used to perform inference for CNN filters. It supports both floating-point and fixed-point implementations. In the fixed-point implementation, static quantization is used to represent both weights and feature maps with integer precision of 16. Network information during the inference phase is shown in Table 3.

[0106] Table 3. Network information of filter set 1 during the inference phase

[0107] 2.3.3. Intra-frame prediction based on neural networks 2.3.3.1. Neural Network Inference The neural network-based intra-frame prediction mode comprises seven neural networks, each predicting... Blocks of different sizes. Predicted size is... The block neural network is represented as ,in Collect its parameters. For a given piece Using the block located above it OK A reference sample point and its left side The context composed of reference samples Preprocessed version to provide Applying post-processing produce Prediction See Figure 16 .also, Return two indexes and . This represents the index that characterizes the LFNST kernel index, and when Whether the main transformation coefficients generated from the residuals of horizontal and vertical DCT-2 applications to neural network predictions are transposed, see [reference needed]. Figure 16 .also, Provide an index of the VVC intra-prediction modes (planar intra-prediction mode, DC intra-prediction mode, or directional intra-prediction mode). Among them, those from surrounding Reference sample The prediction best represents See Figure 16 . Figure 16 This demonstrates how intra-frame prediction patterns based on neural networks can be used to... Context of surrounding reference samples Predict the current situation piece A schematic diagram. Here, and .

[0108] if and :

[0109] otherwise: if :

[0110] otherwise:

[0111] if :

[0112] otherwise: .

[0113] if .otherwise, .

[0114] if .otherwise, .

[0115] 2.3.3.2. Preprocessing and Postprocessing 2.3.3.2.1. Preprocessing of the current block's context Figure 16 The “preprocessing” shown includes the following four steps.

[0116] • from Subtract Available reference samples mean μ See Figure 17 .

[0117] • If the neural network predicting the current block is in floating-point, then the context Multiply the reference sample points in , b This is the internal bit depth, which is 10 in VVC. Otherwise, the context... Multiply the reference sample points in , This indicates the input quantizer.

[0118] • Will All unavailable reference samples (See) Figure 17 Set it to 0.

[0119] • Flatten the context generated in the previous step to obtain a size of vector .

[0120] Figure 17 It shows that it will revolve around the current w×h piece Context of reference sample Decomposed into usable reference samples and unavailable reference points .here, and In the case shown, the number of unusable reference samples cannot reach its maximum value.

[0121] 2.3.3.2.2. Post-processing of neural network predictions Figure 16 The "post-processing" described in the text includes processing data of a size of [missing information]. Whh vector Remodeling to a height of h And the width is w The rectangle, the result of the reshaping is divided by ρ Add the mean of available reference samples in the context of the current block. μ And limited to Therefore, post-processing can be summarized as follows: .

[0122] 2.3.3.3. Adaptive Derivation of the MPM List When creating an MPM list for a given luminance CB, if the "left" luminance CB is predicted via a neural network-based intra-frame prediction mode, the neural network-based mode index can be replaced by the repIdx returned during the prediction of the "left" luminance CB, and becomes a candidate index to be added to the MPM list. Similarly, if the "upper" luminance CB is predicted via a neural network-based intra-frame prediction mode, the neural network-based mode index can be replaced by the repIdx returned during the prediction of the "upper" luminance CB, and becomes a candidate index to be inserted into the MPM list.

[0123] 2.3.3.4. Signaling for Intra-Frame Prediction Mode Based on Neural Networks 2.3.3.4.1. Signaling for Intra-Frame Prediction Mode Based on Neural Networks in Luminosity For the present w×h The luminance CB (whose top-left pixel is located at position (y,x) in the current luminance channel) is divided into two cases for intra-frame prediction mode signaling in luminance.

[0124] • if ,but nnFlag It appears in the intra-prediction mode signaling in the brightness. This means that the neural network-based intra-frame prediction mode is selected to predict the current brightness CB and END. This means that the neural network-based intra-frame prediction mode was not selected to predict the current luminance CB, and then the regular intra-frame prediction mode signaling in the luminance (represented as...) (Applicable, see also) Figure 18 .

[0125] Otherwise, regular intra-frame prediction mode signaling in brightness. Applicable.

[0126] Note that in " and nnFlag When =1”, if the context of the current luminance CB exceeds the boundary of the current luminance channel, i.e. Then, the intra-frame prediction based on the neural network is replaced by the plane.

[0127] .

[0128] Figure 18 It shows the current w×h Intra-prediction mode signaling for the luminance CB (highlighted by a dashed line in orange). The coordinates of the top-left pixel of this CB are ( y, x The binary value of nnFlag is shown in coarse gray. Here, , , ,and .

[0129] 2.3.3.4.2. Signaling for Intra-Frame Prediction Mode Based on Neural Networks in Chroma For the present w×h Chroma CB (its top left pixel is located in the current chroma channel) y, x Intra-predictive mode signaling in chroma is divided into two cases.

[0130] • If the luminance CB, which is in the same position as the chrominance CB, is predicted using an intra-frame prediction mode based on a neural network: o If Then DM becomes an intra-frame prediction mode based on neural networks.

[0131] Otherwise, DM is set to a plane.

[0132] • Otherwise: o If ,but nnFlagChroma It appears in the intra-predictive mode signaling in chroma. nnFlagChroma It is placed before the DM flag in the decision tree of the intra-predictive mode signaling in chroma. nnFlagChroma=1 This means that the neural network-based intra-frame prediction mode is selected to predict the current pair of chroma CB and END. nnFlagChroma=0 This means that the neural network-based intra-prediction mode was not selected to predict the current pair of chroma CBs, and then the regular intra-prediction mode signaling in the chroma was recovered from the DM flag.

[0133] Otherwise, the standard intra-frame prediction mode signaling in chroma applies.

[0134] Note that in " And in the case that DM becomes an intra-frame prediction mode based on neural networks, and in the case of "DM becoming an intra-frame prediction mode based on neural networks", and " and In the case of "", if the context of the current chroma CB exceeds the boundary of the current chroma channel, i.e. Then, the intra-frame prediction based on the neural network is replaced by the plane.

[0135] 2.3.3.5. Transformation of Context and Neural Network Prediction For a given w×h block, if Therefore, the intra-frame prediction mode based on neural networks must predict this block, but the intra-frame prediction mode based on neural networks does not include... In this case, the context of the current block can be... Figure 16 The vertical downsampling factor is referred to in the "preprocessing" step. δ and / or level downsampling factor γ And / or transpose. Then, the prediction of the current block can be... Figure 16 The step referred to as "post-processing" is followed by transposition and / or vertical upsampling factor. δ and / or level upsampling factor γ The context of the current block and the predicted transpose, δ and γ The selected neural network is used for prediction, as shown in Table 4 below.

[0136] Table 4: Current situation to be predicted w×h The context of the block and the decision of transposing the prediction of that block. γ value, δ The value and belonging to each Neural networks used for intra-frame prediction patterns based on neural networks

[0137] 2.3.4. Small Ad hoc deep learning (SADL) library SADL (Small Ad hoc Deep Learning Library) is a small, header-only library for neural network inference. SADL provides both floating-point and integer-based inference capabilities. Neural network inference in NNVC is based on SADL.

[0138] The following table summarizes the framework features.

[0139] Table 5. Characteristics of SADL

[0140] The NNVC repository uses SADL as a submodule, pointing to this repository: https: / / vcgit.hhi.fraunhofer.de / jvet-ahg-nnvc / sadl. Documentation is available in the repository's doc directory.

[0141] 3. Problem Current neural network-based loop filtering has the following problems: 1. During parameter selection, the parameter candidate list contains three candidates by default. However, using three candidates may lead to high coding complexity.

[0142] 2. During parameter selection, the candidate list is derived without considering the total number of allowed candidates. However, it may be beneficial to derive the candidate list by taking into account the total number of allowed candidates.

[0143] 3. The input to a neural network-based loop filter includes reconstructed samples and predicted samples. However, the difference between the reconstructed and predicted samples (also known as residual samples) can also be useful for the loop filtering process.

[0144] 4. To better estimate the rate-distortion (RD) cost in NNVC, a simplified version of the NN-based loop filter is introduced into the rate-distortion optimization (RDO) process for segmentation mode selection. However, the NN filter in RDO uses different models to handle different types of stripes. Applying a single unified model in RDO can help reduce model storage.

[0145] 5. In NNVC, the calculation of the RD cost is performed on video units (i.e., blocks or frames) to select the optimal parameters for the NN filter. However, to reduce the complexity of calculating the RD cost, sub-regions of video units can be used. Furthermore, the parameters of the NN filter for other video units can be derived.

[0146] 4. Detailed Solution The detailed embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.

[0147] One or more neural network (NN) filter models are trained as part of a loop filtering technique or used in post-processing to reduce distortion generated during compression. Samples with different characteristics are processed by different NN filter models or NN filter models with different parameters. This disclosure details how to adaptively derive a list of parameter candidates and how to utilize residual information in the design of NN filters.

[0148] It should be noted that the concepts of adaptive parameter candidate list derivation and residual information utilization can also be extended to other neural network-based encoding and decoding tools, such as neural network-based intra-frame prediction, neural network-based cross-component prediction, neural network-based inter-frame prediction, neural network-based super-resolution, neural network-based motion compensation, neural network-based reference frame generation, and neural network-based transform design. In the following example, we use neural network-based filtering techniques as an example.

[0149] It should also be noted that the concept of residual information utilization can be extended to non-NN-based encoding and decoding tools, such as non-NN-based loop filtering, non-NN-based intra-frame prediction, non-NN-based cross-component prediction, non-NN-based inter-frame prediction, non-NN-based super-resolution, non-NN-based motion compensation, non-NN-based reference frame generation, and non-NN-based transform design. For example, non-NN-based loop filters (deblocking filters, SAO, or ALF) can use residual information as an auxiliary input.

[0150] In this disclosure, the NN filter can be any kind of NN filter, such as a convolutional neural network (CNN) filter, a fully connected neural network filter, a transformer-based filter, or a recurrent neural network-based filter. In the following discussion, the NN filter may also be referred to as a CNN filter.

[0151] In the following discussion, a video unit can be a sequence, image, strip, slice, brick, sub-image, CTU / CTB, CTU / CTB line, one or more CU / CB, one or more CTU / CTB, one or more VPDU (Virtual Pipeline Data Unit), or a sub-region within an image / strip / slice / brick. A parent video unit represents a unit larger than the video unit. Typically, a parent unit will contain multiple video units. For example, when the video unit is a CTU, the parent unit can be a strip, a CTU line, multiple CTUs, etc.

[0152] 1. To address issue 1, the number of parameter candidates is set to 2 by default. In other words, the parameter candidate list contains two candidates by default.

[0153] a. In one example, the variable that depends on the QP (e.g., sequence-level QP) is denoted as q, and the candidate list includes parameters {Param_1, Param_2}, where Param_1 and Param_2 are derived based on q.

[0154] i. In one example, Param_1 = q + Param_2 = q + .

[0155] 1) In one example, q refers to sequence level QP.

[0156] 2) In one example, q refers to the block level QP.

[0157] 3) In one example, q refers to the strip level QP.

[0158] 4) In one example, It can be any negative number, positive number, or zero, with the constraint that... .

[0159] 5) In one example, q, and / or It can depend on any encoded or decoded information, such as stripe type, prediction type, CBF, etc.

[0160] 6) In one example, q, and / or The indication can be transmitted via signal, or merged from adjacent video units, non-adjacent video units, or temporal video units.

[0161] ii. Alternative location, Param_1 = q Param_2 = q , where q, It can have the same interpretation as the above items.

[0162] iii. In one example, Param_1 = Param_2 = ,in It is either a linear function or a nonlinear function.

[0163] iv. In one example, Param_1 = q, Param_2 = q 5.

[0164] v. In one example, Param_1 = q, Param_2 = q 10.

[0165] vi. In one example, Param_1 = q 5. Param_2 = q 10.

[0166] b. Alternatively, QP in the above items can be replaced by other terms / variables, such as energy, quantization step size, etc.

[0167] 2. To solve problem 2, we derive the parameter candidate list by considering the number of parameter candidates.

[0168] a. In one example, the number of parameter candidates is 2. Let q be the variable that depends on the QP (e.g., sequence-level QP), and let the candidate list be {Param_1, Param_2}. For lower temporal layers, Param_1 = Param_2 = For higher time-domain layers, Param_1 = Param_2 = ,in It is either a linear function or a nonlinear function.

[0169] i. In one example, Same, and They are different. In other words, the second candidate is different across different time-domain layers.

[0170] ii. In one example, for the lower time-domain layer, Param_1 = q, Param_2 = q 5. For higher time-domain layers, Param_1 = q, Param_2 = q 5.

[0171] iii. In one example, a low time-domain layer refers to a layer with tid equal to 0, 1, 2 or 3, while a high time-domain layer refers to a layer with tid equal to 4 or 5.

[0172] b. In one example, the number of parameter candidates is 3. Let q be the variable that depends on the QP (e.g., sequence-level QP), and let the candidate list be {Param_1, Param_2, Param_3}. For lower temporal layers, Param_1 = Param_2 = Param_3 = For higher time-domain layers, Param_1 = Param_2 = Param_3 = ,in It is either a linear function or a nonlinear function.

[0173] i. In one example, same, Same, and They are different. In other words, the third candidate is different across different time-domain layers.

[0174] ii. In one example, for the lower time-domain layer, Param_1 = q, Param_2 = q 5. Param_3 = q 10. For higher time-domain layers, Param_1 = q, Param_2 = q 5. Param_3 = q 5.

[0175] iii. In one example, a low time-domain layer refers to a layer with tid equal to 0, 1, 2 or 3, while a high time-domain layer refers to a layer with tid equal to 4 or 5.

[0176] 3. To address problem 3, the difference between the reconstructed samples and the predicted samples (also known as residual samples) is additionally fed into a NN-based loop filter.

[0177] a. In one example, the residual samples and the existing input are first concatenated and then fed into the loop filter.

[0178] b. In one example, the residual samples and the existing input are fed separately into the loop filter.

[0179] 4. To address problem 4, a unified NN filter model is required for RDO.

[0180] a. In one example, the unified NN filter in RDO can handle different types of stripes.

[0181] b. In one example, the uniform NN filter in RDO can handle different types of color components.

[0182] c. In one example, the unified NN filter in RDO can handle different types of stripes and different types of color components.

[0183] d. In one example, the unified NN filter in RDO can take at least one indicator that can be related to the strip type as input.

[0184] e. In one example, the unified NN filter in RDO can take at least one indicator that can be related to the encoding / decoding mode as input.

[0185] f. In one example, the unified NN filter in RDO can take encoded and / or reconstructed and / or derived information from encoded and decoded information as input.

[0186] i. In one example, the encoded or decoded information can be residual information or information derived from residual information, such as the difference between reconstructed samples and predicted samples.

[0187] 5. To address problem 5, sub-regions of video units (i.e., blocks or frames) are used to calculate the RD cost.

[0188] a. In one example, the location of the sub-region can be predefined.

[0189] i. In one example, the sub-region is located in the center of the video unit.

[0190] ii. In one example, the sub-region is located at the top left of the video unit.

[0191] b. In one example, the size of the sub-region can be predefined.

[0192] i. In one example, the size of the sub-region can be one-quarter of the video unit.

[0193] c. In one example, the optimal parameters chosen for the NN filter for one video unit can be used to derive the parameters for other video units.

[0194] i. In one example, the optimal parameters selected for a video cell are directly applied to adjacent video cells.

[0195] ii. In one example, the optimal parameters selected for a luma video unit can be directly applied to the corresponding chroma video unit.

[0196] As used herein, the term "machine learning model" may also be referred to as "machine learning model". A machine learning model can include any kind of model, such as a neural network (NN) model (also known as an "NN filter" or "NN filter model"), a convolutional neural network (CNN) model, etc. Alternatively, in some embodiments, the machine learning model may also include a non-NN-based model or a non-NN-based filter. The scope of this disclosure is not limited in this respect.

[0197] As used herein, the term "video unit" or "video block" can be a sequence, picture, strip, slice, brick, sub-picture, codec tree unit (CTU) / codec tree block (CTB), CTU / CTB line, one or more codec units (CU) / codec blocks (CB), one or more CTU / CTB, one or more Virtual Pipeline Data Units (VPDU), or a sub-region within a picture / strip / slice / brick. As used herein, the term "parent video unit" can refer to a unit larger than a video unit. A parent unit will contain multiple video units. For example, if the video unit is a CTU, the parent unit can be a strip, a CTU line, multiple CTUs, etc.

[0198] Figure 19 A flowchart of a method 1900 for video processing according to an embodiment of the present disclosure is shown. Method 1900 is implemented during the conversion between video units of a video and a bitstream of a video.

[0199] At box 1910, a conversion between the current video unit and the video bitstream is performed, wherein a neural network (NN)-based loop filter is applied to this conversion. During the parameter selection process for the NN-based loop filter, the number of parameter candidates in the parameter candidate list for the NN-based loop filter is less than a threshold number. In some embodiments, the threshold number is 3. For example, the parameter selection process may be to select parameters from two parameter candidates.

[0200] Method 1900 allows for the use of fewer parameter candidates during parameter selection. This reduces encoding and decoding complexity.

[0201] In some embodiments, the parameter candidate list includes a first parameter candidate and a second parameter candidate, which are based on at least one of the following: the quantization parameter (QP) of the current video cell, the energy of the current video cell, or the quantization step size of the current video cell. As used herein, the term "energy" may refer to the energy of a video block or video cell, which can be determined by adding the energies of the samples in the video block or video cell.

[0202] In some embodiments, QP includes sequence-level QP, block-level QP, or stripe-level QP.

[0203] In some embodiments, the first parameter candidate and the second parameter candidate are determined by Param_1 = q + M1 and Param_2 = q + M2, where Param_1 represents the first parameter candidate, Param_2 represents the second parameter candidate, q represents at least one of the following: the QP, the energy, or the quantization step size, M1 represents the first factor, and M2 represents the second factor.

[0204] In some embodiments, the first parameter candidate and the second parameter candidate are determined by Param_1 = q × M1 and Param_2 = q × M2, where Param_1 represents the first parameter candidate, Param_2 represents the second parameter candidate, q represents at least one of the following: the QP, the energy or the quantization step size, M1 represents the first factor, and M2 represents the second factor.

[0205] In some embodiments, the first factor is different from the second factor, and the first factor or the second factor is negative, positive, or zero.

[0206] In some embodiments, at least one of q, M1, or M2 is based on encoded information, which includes at least one of the following: stripe type, prediction type, or codec block flag (such as cbf).

[0207] In some embodiments, at least one indication of q, M1, or M2 is included in the bitstream, or merged from adjacent or non-adjacent video units or temporal video units of the current video unit.

[0208] In some embodiments, the first parameter candidate and the second parameter candidate are determined by Param_1 = f1(q) and Param_2 = f2(q), where Param_1 represents the first parameter candidate, Param_2 represents the second parameter candidate, q represents at least one of the following: the QP, the energy, or the quantization step size, and f1 and f2 are linear or nonlinear functions.

[0209] In some embodiments, f1 includes Param_1 = q and f2 includes Param_2 = q5, or f1 includes Param_1 = q and f2 includes Param_2 = q10, or f1 includes Param_1 = q5 and f2 includes Param_2 = q10.

[0210] In some embodiments, the number of parameter candidates in the parameter candidate list is 2, and wherein for at least one low time domain layer, the parameter candidate list includes a first parameter candidate Param_1 = f1(q) and a second parameter candidate Param_2 = f2(q), and for at least one high time domain layer, the parameter candidate list includes a first parameter candidate Param_1 = f3(q) and a second parameter candidate Param_2 = f4(q), wherein f1, f2, f3 and f4 are linear functions or nonlinear functions.

[0211] In some embodiments, f1 and f2 are the same, and f3 and f4 are different.

[0212] In some embodiments, for at least one low time-domain layer, Param_1 = q and Param_2 = q5, and for at least one high time-domain layer, Param_1 = q and Param_2 = q5.

[0213] In some embodiments, at least one time domain identifier (tid) of at least one low time domain layer is greater than or equal to 0 and less than or equal to a first value, and at least one tid of at least one high time domain layer is greater than the first value and less than or equal to a second value.

[0214] In some embodiments, the first value is 3 and the second value is 5.

[0215] In some embodiments, the number of parameter candidates in the other parameter candidate list of the video unit is 3, and wherein for at least one low temporal layer, the other parameter candidate list includes a first parameter candidate Param_1 = f1(q), a second parameter candidate Param_2 = f2(q), and a third parameter candidate Param_2 = f3(q), and for at least one high temporal layer, the other parameter candidate list includes a first parameter candidate Param_1 = f4(q), a second parameter candidate Param_2 = f5(q), and a third parameter candidate Param_3 = f6(q), wherein f1, f2, f3, f4, f5, and f6 are linear functions or nonlinear functions.

[0216] In some embodiments, f1 and f4 are the same, f2 and f5 are the same, and f3 and f6 are different.

[0217] In some embodiments, for at least one low time-domain layer, Param_1 = q, Param_2 = q-5 and Param_3 = q-10, and for at least one high time-domain layer, Param_1 = q, Param_2 = q-5 and Param_3 = q+5.

[0218] In some embodiments, at least one time domain identifier (tid) of at least one low time domain layer is greater than or equal to 0 and less than or equal to a first value, and at least one tid of at least one high time domain layer is greater than the first value and less than or equal to a second value.

[0219] In some embodiments, the first value is 3 and the second value is 5.

[0220] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: generating the bitstream of video, wherein a neural network (NN)-based loop filter is applied to the generation, and wherein during a parameter selection process for the NN-based loop filter, the number of parameter candidates in a parameter candidate list for the NN-based loop filter is less than a threshold number.

[0221] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: generating a bitstream of video, wherein a neural network (NN)-based loop filter is applied to the generation, and wherein during a parameter selection process for the NN-based loop filter, the number of parameter candidates in a parameter candidate list for the NN-based loop filter is less than a threshold number; and storing the bitstream in a non-transitory computer-readable recording medium.

[0222] Figure 20A flowchart of a method 2000 for video processing according to an embodiment of the present disclosure is shown. Method 2000 is implemented during the conversion between video units of a video and a bitstream of a video.

[0223] At box 2010, a conversion between the current video unit and the video bitstream is performed, wherein a neural network (NN) based loop filter is applied to the conversion. At least one residual sample between at least one reconstructed sample and at least one predicted sample of the current video unit is fed into the NN-based loop filter.

[0224] Method 2000 enables the use of residual samples in the loop filtering process. Therefore, encoding / decoding efficiency and effectiveness can be improved.

[0225] In some embodiments, at least one residual sample and at least one additional input are concatenated and then fed into an NN-based loop filter.

[0226] In some embodiments, at least one residual sample and at least one additional input are fed separately into an NN-based loop filter.

[0227] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: generating the bitstream of video, wherein a neural network (NN)-based loop filter is applied to the generation, and wherein at least one residual sample between at least one reconstructed sample and at least one predicted sample of a current video unit is fed into the NN-based loop filter.

[0228] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: generating a bitstream of video, wherein a neural network (NN)-based loop filter is applied to the generation, and wherein at least one residual sample between at least one reconstructed sample and at least one predicted sample of a current video unit is fed into the NN-based loop filter; and storing the bitstream in a non-transitory computer-readable recording medium.

[0229] Figure 21 A flowchart of a method 2100 for video processing according to an embodiment of the present disclosure is shown. Method 2100 is implemented during the conversion between video units of a video and a bitstream of a video.

[0230] At box 2110, a conversion between the current video unit and the video bitstream is performed, where a unified neural network (NN)-based filter is applied to the rate-distortion optimization (RDO) process. The NN-based filter is applied to different types of stripes and / or different color components.

[0231] Method 2100 enables the application of unified NN-based filters to video encoding and decoding. Model storage can therefore be reduced.

[0232] In some embodiments, the unified NN-based filter in the RDO takes at least one indicator related to at least one stripe type as input.

[0233] In some embodiments, the unified NN-based filter in the RDO takes at least one indicator related to at least one encoding / decoding mode as input.

[0234] In some embodiments, the unified NN filter in the RDO takes at least one of the following as input: encoded or decoded information, reconstructed information, or information derived from encoded or decoded information.

[0235] In some embodiments, the encoded and decoded information includes at least one of the following: residual information, information derived from the residual information, or at least one difference between at least one reconstructed sample and at least one predicted sample.

[0236] In some embodiments, at least one sub-region of the current video unit is used for determination rate distortion (RD) cost, and the current video unit includes a video block or frame.

[0237] In some embodiments, at least one location of at least one sub-region is predefined.

[0238] In some embodiments, at least one sub-region includes at least one of the following: a sub-region at the center of the current video unit, or a sub-region at the upper left position of the current video unit.

[0239] In some embodiments, the size of at least one sub-region is predefined.

[0240] In some embodiments, the size of at least one sub-region is one-quarter of the current video unit.

[0241] In some embodiments, at least one parameter is selected from a list of candidate parameters for a NN-based filter, and the selected at least one parameter is used to determine at least one parameter for another video unit.

[0242] In some embodiments, at least one parameter selected for a video unit is applied to the adjacent video units of the video unit.

[0243] In some embodiments, at least one parameter selected for the luminance video unit is applied to the corresponding chrominance video unit.

[0244] In some embodiments, the conversion includes encoding the current video unit into a bitstream.

[0245] In some embodiments, the conversion includes decoding the current video unit from the bitstream.

[0246] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: generating the bitstream of video, wherein a unified neural network (NN)-based filter is applied to a rate-distortion optimization (RDO) process, and wherein the unified NN-based filter is applied to different types of stripes and / or different color components.

[0247] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: generating a bitstream of video, wherein a unified neural network (NN)-based filter is applied to a rate-distortion optimization (RDO) process, and wherein the unified NN-based filter is applied to different types of stripes and / or different color components; and storing the bitstream in a non-transitory computer-readable recording medium.

[0248] It should be understood that Method 1900, Method 2000 and / or Method 2100 can be applied individually or in any combination.

[0249] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.

[0250] Item 1. A method for video processing, comprising: performing a conversion between a current video unit of a video and a bitstream of the video, wherein a neural network (NN)-based loop filter is applied to the conversion, and wherein during a parameter selection process for the NN-based loop filter, the number of parameter candidates in a parameter candidate list for the NN-based loop filter is less than a threshold number.

[0251] Item 2. The method according to Item 1, wherein the number of thresholds is 3.

[0252] Item 3. The method according to Item 1 or 2, wherein the parameter candidate list includes a first parameter candidate and a second parameter candidate, the first parameter candidate and the second parameter candidate being based on at least one of the following: the quantization parameter (QP) of the current video unit, the energy of the current video unit, or the quantization step size of the current video unit.

[0253] Item 4. The method according to Item 3, wherein the QP includes sequence-level QP, block-level QP, or strip-level QP.

[0254] Item 5. The method according to Item 3 or 4, wherein the first parameter candidate and the second parameter candidate are determined by: Param_1 = q + M1 and Param_2 = q + M2, wherein Param_1 represents the first parameter candidate, Param_2 represents the second parameter candidate, q represents at least one of the following: the QP, the energy or the quantization step size, M1 represents the first factor, and M2 represents the second factor.

[0255] Item 6. The method according to Item 3 or 4, wherein the first parameter candidate and the second parameter candidate are determined by: Param_1 = q × M1 and Param_2 = q × M2, wherein Param_1 represents the first parameter candidate, Param_2 represents the second parameter candidate, q represents at least one of the following: the QP, the energy or the quantization step size, M1 represents the first factor, and M2 represents the second factor.

[0256] Item 7. The method according to Item 5 or 6, wherein the first factor is different from the second factor, and the first factor or the second factor is negative, positive or zero.

[0257] Item 8. The method according to any one of items 5 to 7, wherein at least one of q, M1 or M2 is based on encoded information, said encoded information including at least one of: stripe type, prediction type, or codec block flag.

[0258] Item 9. The method according to any one of Items 5 to 8, wherein at least one indication of at least one of q, M1 or M2 is included in the bitstream or merged from adjacent video units or non-adjacent video units or temporal video units of the current video unit.

[0259] Item 10. The method according to Item 3 or 4, wherein the first parameter candidate and the second parameter candidate are determined by Param_1 = f1(q) and Param_2 = f2(q), wherein Param_1 represents the first parameter candidate, Param_2 represents the second parameter candidate, q represents at least one of the following: the QP, the energy, or the quantization step size, and f1 and f2 are linear or nonlinear functions.

[0260] Item 11. The method according to Item 10, wherein f1 includes Param_1 = q, and f2 includes Param_2 = q. 5, or where f1 includes Param_1 = q, and f2 includes Param_2 = q. 10, or where f1 includes Param_1 = q 5, and f2 includes Param_2 = q 10.

[0261] Item 12. The method according to any one of items 1 to 11, wherein the number of parameter candidates in the parameter candidate list is 2, and wherein for at least one low time-domain layer, the parameter candidate list includes a first parameter candidate Param_1 = f1(q) and a second parameter candidate Param_2 = f2(q), and for at least one high time-domain layer, the parameter candidate list includes a first parameter candidate Param_1 = f3(q) and a second parameter candidate Param_2 = f4(q), wherein f1, f2, f3 and f4 are linear functions or nonlinear functions.

[0262] Item 13. The method according to Item 12, wherein f1 and f2 are the same, and f3 and f4 are different.

[0263] Item 14. The method according to Item 12 or 13, wherein for the at least one low-time-domain layer, Param_1 = q and Param_2 = q 5, and for the at least one high-temporal layer, Param_1 = q and Param_2 = q 5.

[0264] Item 15. The method according to any one of items 12 to 14, wherein at least one time domain identifier (tid) of the at least one low time domain layer is greater than or equal to 0 and less than or equal to a first value, and at least one tid of the at least one high time domain layer is greater than the first value and less than or equal to a second value.

[0265] Item 16. The method according to Item 15, wherein the first value is 3 and the second value is 5.

[0266] Item 17. The method according to any one of items 1 to 11, wherein the number of parameter candidates in another parameter candidate list of the video unit is three, and wherein for at least one low temporal layer, the other parameter candidate list includes a first parameter candidate Param_1 = f1(q), a second parameter candidate Param_2 = f2(q) and a third parameter candidate Param_2 = f3(q), and for at least one high temporal layer, the other parameter candidate list includes a first parameter candidate Param_1 = f4(q), a second parameter candidate Param_2 = f5(q) and a third parameter candidate Param_3 = f6(q), wherein f1, f2, f3, f4, f5 and f6 are linear functions or nonlinear functions.

[0267] Item 18. The method according to Item 17, wherein f1 and f4 are the same, f2 and f5 are the same, and f3 and f6 are different.

[0268] Item 19. The method according to Item 17 or 18, wherein for the at least one low time-domain layer, Param_1 = q, Param_2 = q-5, and Param_3 = q-10, and for the at least one high time-domain layer, Param_1 = q, Param_2 = q-5, and Param_3 = q+5.

[0269] Item 20. The method according to any one of items 17 to 19, wherein at least one time domain identifier (tid) of the at least one low time domain layer is greater than or equal to 0 and less than or equal to a first value, and at least one tid of the at least one high time domain layer is greater than the first value and less than or equal to a second value.

[0270] Item 21. The method according to Item 20, wherein the first value is 3 and the second value is 5.

[0271] Item 22. A method for video processing, comprising: performing a conversion between a current video unit of a video and a bitstream of the video, wherein a neural network (NN)-based loop filter is applied to the conversion, and wherein at least one residual sample between at least one reconstructed sample and at least one predicted sample of the current video unit is fed into the NN-based loop filter.

[0272] Item 23. The method according to Item 22, wherein the at least one residual sample and at least one additional input are concatenated and then fed into the NN-based loop filter.

[0273] Item 24. The method according to Item 22, wherein the at least one residual sample and at least one additional input are fed separately into the NN-based loop filter.

[0274] Item 25. A method for video processing, comprising: performing a conversion between a current video unit of a video and a bitstream of the video, wherein a unified neural network (NN)-based filter is applied to a rate-distortion optimization (RDO) process, and wherein the unified NN-based filter is applied to different types of stripes and / or different color components.

[0275] Item 26. The method according to Item 25, wherein the unified NN-based filter in the RDO takes at least one indicator related to at least one stripe type as input.

[0276] Item 27. The method according to Item 25, wherein the filter based on the unified NN in the RDO takes at least one indicator related to at least one encoding / decoding mode as input.

[0277] Item 28. The method according to any one of items 25 to 27, wherein the unified NN filter in the RDO takes at least one of the following as input: encoded or decoded information, reconstructed information, or information derived from the encoded or decoded information.

[0278] Item 29. The method according to Item 28, wherein the encoded / decoded information includes at least one of the following: residual information, information derived based on the residual information, or at least one difference between at least one reconstructed sample and at least one predicted sample.

[0279] Item 30. The method according to any one of Items 25 to 29, wherein at least one sub-region of the current video unit is used for determination rate distortion (RD) cost, the current video unit comprising a video block or frame.

[0280] Item 31. According to the method of Item 30, at least one location of the at least one sub-region is predefined.

[0281] Item 32. The method according to Item 30 or 31, wherein the at least one sub-region includes at least one of the following: a sub-region at the center of the current video unit or a sub-region at the upper left position of the current video unit.

[0282] Item 33. The method according to any one of items 30 to 32, wherein the size of the at least one sub-region is predefined.

[0283] Item 34. The method according to Item 33, wherein the size of the at least one sub-region is one-quarter of the current video unit.

[0284] Item 35. The method according to any one of items 30 to 34, wherein at least one parameter is selected from a list of candidate parameters for the NN-based filter, and the selected at least one parameter is used to determine at least one parameter for another video unit.

[0285] Item 36. The method according to Item 35, wherein the at least one parameter selected for the video unit is applied to the adjacent video units of the video unit.

[0286] Item 37. The method according to Item 35, wherein the at least one parameter selected for the luminance video unit is applied to the corresponding chroma video unit.

[0287] Item 38. The method according to any one of items 1 to 37, wherein the conversion includes encoding the current video unit into the bitstream.

[0288] Item 39. The method according to any one of items 1 to 37, wherein the conversion includes decoding the current video unit from the bitstream.

[0289] Item 40. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1 to 39.

[0290] Item 41. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of items 1 to 39.

[0291] Item 42. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by an apparatus for video processing, wherein the method includes: generating the bitstream of the video, wherein a neural network (NN) based loop filter is applied to the generation, and wherein during a parameter selection process for the NN-based loop filter, the number of parameter candidates in a parameter candidate list for the NN-based loop filter is less than a threshold number.

[0292] Item 43. A method for storing a bitstream of video, comprising: generating the bitstream of the video, wherein a neural network (NN) based loop filter is applied to the generation, and wherein during a parameter selection process for the NN-based loop filter, the number of parameter candidates in a parameter candidate list for the NN-based loop filter is less than a threshold number; and storing the bitstream in a non-transitory computer-readable recording medium.

[0293] Item 44. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: generating the bitstream of the video, wherein a neural network (NN) based loop filter is applied to the generation, and wherein at least one residual sample between at least one reconstructed sample and at least one predicted sample of a current video unit is fed into the NN-based loop filter.

[0294] Item 45. A method for storing a bitstream of video, comprising: generating the bitstream of the video, wherein a neural network (NN) based loop filter is applied to the generation, and wherein at least one residual sample between at least one reconstructed sample and at least one predicted sample of a current video unit is fed into the NN-based loop filter; and storing the bitstream in a non-transitory computer-readable recording medium.

[0295] Item 46. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by an apparatus for video processing, wherein the method includes: generating the bitstream of the video, wherein a filter based on a unified neural network (NN) is applied to a rate-distortion optimization (RDO) process, and wherein the filter based on the unified NN is applied to different types of stripes and / or different color components.

[0296] Item 47. A method for storing a bitstream of video, comprising: generating the bitstream of the video, wherein a filter based on a unified neural network (NN) is applied to a rate-distortion optimization (RDO) process, and wherein the filter based on the unified NN is applied to different types of stripes and / or different color components; and storing the bitstream in a non-transitory computer-readable recording medium.

[0297] Example device Figure 22A block diagram of a computing device 2200 in which various embodiments of the present disclosure may be implemented is shown. The computing device 2200 may be implemented as a source device 110 (or video encoder 114 or 200) or a target device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a target device 120 (or video decoder 124 or 300).

[0298] It should be understood that, Figure 22 The computing device 2200 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.

[0299] like Figure 22 As shown, computing device 2200 includes general-purpose computing device 2200. Computing device 2200 may include at least one or more processors or processing units 2210, memory 2220, storage unit 2230, one or more communication units 2240, one or more input devices 2250, and one or more output devices 2260.

[0300] In some embodiments, computing device 2200 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server provided by a service provider, a large computing device, etc. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that computing device 2200 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).

[0301] Processing unit 2210 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 2220. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 2200. Processing unit 2210 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.

[0302] Computing device 2200 typically includes various computer storage media. Such media can be any media accessible by computing device 2200, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 2220 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 2230 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 2200.

[0303] The computing device 2200 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 22 Not shown, but may provide disk drives for reading from and / or writing to removable non-volatile disks, and optical disc drives for reading from and / or writing to removable non-volatile optical discs. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.

[0304] Communication unit 2240 communicates with another computing device via a communication medium. Furthermore, the functionality of components in computing device 2200 can be implemented by a single computing cluster or by multiple computing machines communicating via communication connections. Therefore, computing device 2200 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0305] Input device 2250 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 2260 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 2240, computing device 2200 can also communicate with one or more external devices (not shown), such as storage devices and display devices, and / or with one or more devices that enable a user to interact with computing device 2200, or any device that enables computing device 2200 to communicate with one or more other computing devices (e.g., network card, modem, etc.), if needed. Such communication can be performed via an input / output (I / O) interface (not shown).

[0306] In some embodiments, some or all components of computing device 2200 may be arranged in a cloud computing architecture, rather than integrated into a single device. In a cloud computing architecture, components may be provided remotely and work together to achieve the functions described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider offers applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed at locations in remote data centers. Cloud computing infrastructure may provide services through shared data centers, although they may appear as a single access point for users. Therefore, cloud computing architectures can be used to provide the components and functions described herein from service providers at remote locations. Alternatively, they may be provided from conventional servers or directly installed or otherwise installed on client devices.

[0307] In embodiments of this disclosure, computing device 2200 can be used to implement video encoding / decoding. Memory 2220 may include one or more video codec modules 2225 having one or more program instructions. These modules can be accessed and executed by processing unit 2210 to perform the functions of the various embodiments described herein.

[0308] In an example embodiment of performing video encoding, input device 2250 may receive video data as input 2270 to be encoded. The video data may be processed, for example, by video codec module 2225 to generate an encoded bitstream. The encoded bitstream may be provided as output 2280 via output device 2260.

[0309] In an example embodiment of performing video decoding, input device 2250 may receive an encoded bitstream as input 2270. The encoded bitstream may be processed, for example, by a video codec module 2225 to generate decoded video data. The decoded video data may be provided as output 2280 via output device 2260.

[0310] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These changes are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.

Claims

1. A method for video processing, comprising: A conversion is performed between the current video unit of the video and the bitstream of the video, wherein a neural network (NN) based loop filter is applied to the conversion, and wherein during the parameter selection process for the NN-based loop filter, the number of parameter candidates in the parameter candidate list for the NN-based loop filter is less than a threshold number.

2. The method according to claim 1, wherein the number of thresholds is 3.

3. The method according to claim 1 or 2, wherein the parameter candidate list includes a first parameter candidate and a second parameter candidate, the first parameter candidate and the second parameter candidate being based on at least one of the following: the quantization parameter (QP) of the current video unit, the energy of the current video unit, or the quantization step size of the current video unit.

4. The method according to claim 3, wherein the QP includes sequence-level QP, block-level QP, or strip-level QP.

5. The method according to claim 3 or 4, wherein the first parameter candidate and the second parameter candidate are determined by: Param_1 = q + M1 and Param_2 = q + M2. Where Param_1 represents the first parameter candidate, Param_2 represents the second parameter candidate, q represents at least one of the following: the QP, the energy, or the quantization step size, M1 represents the first factor, and M2 represents the second factor.

6. The method according to claim 3 or 4, wherein the first parameter candidate and the second parameter candidate are determined by: Param_1 = q × M1 and Param_2 = q × M2, Where Param_1 represents the first parameter candidate, Param_2 represents the second parameter candidate, q represents at least one of the following: the QP, the energy, or the quantization step size, M1 represents the first factor, and M2 represents the second factor.

7. The method according to claim 5 or 6, wherein the first factor is different from the second factor, and the first factor or the second factor is negative, positive or zero.

8. The method according to any one of claims 5 to 7, wherein at least one of q, M1 or M2 is based on encoded information, said encoded information including at least one of: stripe type, prediction type, or codec block flag.

9. The method according to any one of claims 5 to 8, wherein at least one indication of at least one of q, M1 or M2 is included in the bitstream or merged from adjacent video units or non-adjacent video units or temporal video units of the current video unit.

10. The method of claim 3 or 4, wherein the first parameter candidate and the second parameter candidate are determined by: Param_1 = f 1(q) and Param_2 = f 2(q), Where Param_1 represents the first parameter candidate, Param_2 represents the second parameter candidate, and q represents at least one of the following: the QP, the energy, or the quantization step size, and f 1 and f 2 is a linear function or a nonlinear function.

11. The method of claim 10, wherein f 1 includes Param_1 = q, and f 2 includes Param_2 = q 5, or in f 1 includes Param_1 = q, and f 2 includes Param_2 = q 10, or in f 1 includes Param_1 = q 5, and f 2 includes Param_2 = q 10.

12. The method according to any one of claims 1 to 11, wherein the number of parameter candidates in the parameter candidate list is 2, and For at least one low-time-domain layer, the parameter candidate list includes a first parameter candidate Param_1 = f 1(q) and the second parameter candidate Param_2 = f 2(q), and for at least one high-temporal layer, the parameter candidate list includes a first parameter candidate Param_1 = f 3(q) and the second parameter candidate Param_2 = f 4(q), where f 1. f 2. f 3 and f 4 is a linear function or a nonlinear function.

13. The method of claim 12, wherein f 1 and f 2 are the same, and f 3 and f 4. Different.

14. The method of claim 12 or 13, wherein for the at least one low-time-domain layer, Param_1 = q and Param_2 = q 5, and for the at least one high-temporal layer, Param_1 = q and Param_2 = q 5.

15. The method according to any one of claims 12 to 14, wherein at least one time domain identifier (tid) of the at least one low time domain layer is greater than or equal to 0 and less than or equal to a first value, and at least one tid of the at least one high time domain layer is greater than the first value and less than or equal to a second value.

16. The method of claim 15, wherein the first value is 3 and the second value is 5.

17. The method according to any one of claims 1 to 11, wherein the number of parameter candidates in another parameter candidate list of the video unit is three, and For at least one low-time-domain layer, the other parameter candidate list includes a first parameter candidate Param_1 = f 1(q), Second parameter candidate Param_2 = f 2(q) and the third parameter candidate Param_2 = f 3(q), and for at least one high-temporal layer, the other parameter candidate list includes a first parameter candidate Param_1 = f 4(q), Second parameter candidate Param_2 = f 5(q) and the third parameter candidate Param_3 = f 6(q), where f 1. f 2. f 3. f 4. f 5 and f 6 is a linear function or a nonlinear function.

18. The method of claim 17, wherein f 1 and f 4 are the same. f 2 and f 5 are the same, and f 3 and f 6. Different.

19. The method of claim 17 or 18, wherein for the at least one low time-domain layer, Param_1 = q, Param_2 = q-5, and Param_3 = q-10, and for the at least one high time-domain layer, Param_1 = q, Param_2 = q-5, and Param_3 = q+5.

20. The method according to any one of claims 17 to 19, wherein at least one time domain identifier (tid) of the at least one low time domain layer is greater than or equal to 0 and less than or equal to a first value, and at least one tid of the at least one high time domain layer is greater than the first value and less than or equal to a second value.

21. The method of claim 20, wherein the first value is 3 and the second value is 5.

22. A method for video processing, comprising: A conversion is performed between the current video unit of the video and the bitstream of the video, wherein a neural network (NN) based loop filter is applied to the conversion, and wherein at least one residual sample between at least one reconstructed sample and at least one predicted sample of the current video unit is fed into the NN-based loop filter.

23. The method of claim 22, wherein the at least one residual sample and at least one additional input are concatenated and then fed into the NN-based loop filter.

24. The method of claim 22, wherein the at least one residual sample and at least one additional input are fed separately into the NN-based loop filter.

25. A method for video processing, comprising: Perform the conversion between the current video unit of the video and the bitstream of the video, wherein a unified neural network (NN) based filter is applied to the rate-distortion optimization (RDO) process, and wherein the unified NN based filter is applied to different types of stripes and / or different color components.

26. The method of claim 25, wherein the unified NN-based filter in the RDO takes at least one indicator related to at least one stripe type as input.

27. The method of claim 25, wherein the unified NN-based filter in the RDO takes at least one indicator related to at least one encoding / decoding mode as input.

28. The method according to any one of claims 25 to 27, wherein the unified NN filter in the RDO takes at least one of the following as input: encoded or decoded information, reconstructed information, or information derived from the encoded or decoded information.

29. The method of claim 28, wherein the encoded / decoded information comprises at least one of the following: residual information, information derived based on the residual information, or at least one difference between at least one reconstructed sample and at least one predicted sample.

30. The method according to any one of claims 25 to 29, wherein at least one sub-region of the current video unit is used for determination rate distortion (RD) cost, the current video unit comprising a video block or frame.

31. The method of claim 30, wherein at least one location of the at least one sub-region is predefined.

32. The method according to claim 30 or 31, wherein the at least one sub-region comprises at least one of the following: a sub-region at the center of the current video unit or a sub-region at the upper left position of the current video unit.

33. The method according to any one of claims 30 to 32, wherein the size of the at least one sub-region is predefined.

34. The method of claim 33, wherein the size of the at least one sub-region is one-quarter of the current video unit.

35. The method of any one of claims 30 to 34, wherein at least one parameter is selected from a list of candidate parameters for the NN-based filter, and the selected at least one parameter is used to determine at least one parameter for another video unit.

36. The method of claim 35, wherein the at least one parameter selected for the video unit is applied to the adjacent video units of the video unit.

37. The method of claim 35, wherein the at least one parameter selected for the luminance video unit is applied to the corresponding chroma video unit.

38. The method according to any one of claims 1 to 37, wherein the conversion comprises encoding the current video unit into the bitstream.

39. The method according to any one of claims 1 to 37, wherein the conversion comprises decoding the current video unit from the bitstream.

40. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 39.

41. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 39.

42. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: The bitstream of the video is generated, wherein a neural network (NN) based loop filter is applied to the generation, and wherein during the parameter selection process for the NN-based loop filter, the number of parameter candidates in the parameter candidate list for the NN-based loop filter is less than a threshold number.

43. A method for storing a bitstream of video, comprising: The bitstream of the video is generated, wherein a neural network (NN) based loop filter is applied to the generation, and wherein during the parameter selection process for the NN-based loop filter, the number of parameter candidates in the parameter candidate list for the NN-based loop filter is less than a threshold number. as well as The bitstream is stored in a non-transitory computer-readable recording medium.

44. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: The bitstream of the video is generated, wherein a neural network (NN) based loop filter is applied to the generation, and wherein at least one residual sample between at least one reconstructed sample and at least one predicted sample of the current video unit is fed into the NN-based loop filter.

45. A method for storing a bitstream of video, comprising: The bitstream of the video is generated, wherein a neural network (NN) based loop filter is applied to the generation, and wherein at least one residual sample between at least one reconstructed sample and at least one predicted sample of the current video unit is fed into the NN-based loop filter. as well as The bitstream is stored in a non-transitory computer-readable recording medium.

46. ​​A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: The bitstream of the video is generated, wherein a unified neural network (NN)-based filter is applied to a rate-distortion optimization (RDO) process, and wherein the unified NN-based filter is applied to different types of stripes and / or different color components.

47. A method for storing a bitstream of video, comprising: The bitstream of the video is generated, wherein a unified neural network (NN)-based filter is applied to a rate-distortion optimization (RDO) process, and wherein the unified NN-based filter is applied to different types of stripes and / or different color components; and The bitstream is stored in a non-transitory computer-readable recording medium.