Neural network based post filter for video coding
By introducing a neural network filter model into the video bitstream, the problem of low bandwidth utilization efficiency in existing video encoding and decoding technologies is solved, achieving more efficient video quality and compression effects, and improving the flexibility and adaptability of video processing.
Patent Information
- Application Number
- CN202210358053.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-04-06
- Filing Date
- 2022-04-06
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2042-04-06
AI Technical Summary
Existing video encoding and decoding technologies suffer from inefficient bandwidth utilization, especially in the Internet and digital communication networks. As the number of connected user devices increases, the bandwidth demand for video data continues to grow, and existing loop filters are unable to effectively improve video quality and compression efficiency.
A neural network-based filter model is introduced, which instructs the selection and enabling/disabling of filter models for video units or samples by including supplementary enhancement information (SEI) messages in the video bitstream. Combined with arithmetic encoding/decoding or exponential Columbus encoding/decoding, intelligent filtering processing of video units is achieved.
It improves the quality and compression efficiency of video encoding and decoding, enhances bandwidth utilization, adapts to the characteristics of different video units, and strengthens the flexibility and effectiveness of video processing.
Smart Images

Figure CN115209143B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This patent application claims the benefit of U.S. Provisional Patent Application No. 63 / 171,415, filed April 6, 2021, entitled “Neural Network-Based Post Filter For Video Coding,” which is incorporated herein by reference. Technical Field
[0003] This disclosure generally relates to video encoding and decoding, and more specifically, to loop filters in image / video encoding and decoding. Background Technology
[0004] Digital video consumes the largest share of bandwidth on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention
[0005] The disclosed aspects / embodiments describe one or more bitstreams where their Supplemental Enhancement Information (SEI) messages include indicators specifying one or more neural network (NN) filter model candidates or selections for a video cell or a sample within a video cell. Furthermore, one or more bitstreams are described where their SEI messages include indicators specifying one or more neural network (NN) filter model indices. Additionally, one or more bitstreams are described where their SEI messages include syntax elements indicating whether one or more neural network (NN) filter models are enabled or disabled for a video cell.
[0006] The first aspect relates to a method for processing video. The method includes determining a supplementary enhancement information (SEI) message for a bitstream, including indicators specifying one or more neural network (NN) filter model candidates or selections for video units or samples within video units, and performing a conversion between a video media file and a bitstream comprising the video units based on the indicators.
[0007] Alternatively, in any of the foregoing aspects, another implementation of this aspect provides whether the SEI message includes an indicator based on whether the NN filter is enabled for the video unit.
[0008] Alternatively, in any of the foregoing aspects, another implementation of this aspect provides that when the NN filter is disabled for the video unit, the indicator is excluded from the SEI message.
[0009] Optionally, in any of the preceding aspects, another embodiment of the aspect provides coding the indicator into the bitstream using arithmetic coding or exponential Golomb (EG) coding.
[0010] Optionally, in any of the preceding aspects, another embodiment of the aspect provides the indicator in the SEI message further specifies one or more NN filter model indices.
[0011] Optionally, in any of the preceding aspects, another embodiment of the aspect provides determining the SEI message of the bitstream includes the indicator specifying the one or more NN filter model indices.
[0012] Optionally, in any of the preceding aspects, another embodiment of the aspect provides coding the indicator specifying the one or more NN filter model indices into the bitstream using fixed length coding.
[0013] Optionally, in any of the preceding aspects, another embodiment of the aspect provides deriving the one or more NN filter model indices or deriving on / off control of the one or more NN filter models based on the SEI message.
[0014] Optionally, in any of the preceding aspects, another embodiment of the aspect provides the deriving is dependent on a parameter of a current video unit or a neighboring video unit of the current video unit.
[0015] Optionally, in any of the preceding aspects, another embodiment of the aspect provides the parameter includes a quantization parameter (QP) associated with the video unit or the neighboring video unit.
[0016] Optionally, in any of the preceding aspects, another embodiment of the aspect provides the SEI message includes a first indicator in a parent unit of the bitstream, and wherein the first indicator indicates how the one or more NN filter model indices are included in the SEI message for each video unit included in the parent unit or how the NN filter is used for each video unit included in the parent unit.
[0017] Optionally, in any of the preceding aspects, another embodiment of the aspect provides determining the SEI message of the bitstream includes a syntax element indicating whether one or more neural network (NN) filter models are enabled or disabled for a video unit.
[0018] Optionally, in any of the preceding aspects, another embodiment of the aspect provides context coding or bypass coding the syntax element in the bitstream.
[0019] Optionally, in any of the preceding aspects, another embodiment of the aspect provides arithmetic coding the syntax element using a single context or multiple contexts.
[0020] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that the indicator is included in a plurality of levels of the SEI message, and wherein the plurality of levels comprises a picture level, a slice level, a coding tree block (CTB) level, and a coding tree unit (CTU) level.
[0021] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that the plurality of levels comprises a first level and a second level, and wherein whether the second level is included depends on a value of a syntax element of the first level.
[0022] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that the converting comprises encoding the video media file into the bitstream.
[0023] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that the converting comprises decoding the bitstream to obtain the video media file.
[0024] A second aspect relates to an apparatus for coding video data, comprising a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to determine that a supplemental enhancement information (SEI) message of a bitstream includes an indicator of one or more neural network (NN) filter model candidates or selections that specify a video unit or samples within the video unit, and to convert between a video media file comprising the video unit and the bitstream based on the indicator.
[0025] A third aspect relates to a non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing apparatus. In an embodiment, the method comprises determining that a supplemental enhancement information (SEI) message of the bitstream includes an indicator of one or more neural network (NN) filter model candidates or selections that specify a video unit or samples within the video unit, and generating the bitstream based on the SEI message.
[0026] For clarity, any of the above-mentioned embodiments can be combined with any one or more of the other above-mentioned embodiments to create a new embodiment within the scope of the present disclosure.
[0027] These and other features will become more apparent from the following detailed description, taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF DRAWINGS
[0028] For a more complete understanding of the present disclosure, reference is now made to the following brief description of the drawings taken in conjunction with the detailed description below, in which like reference numerals represent like elements, and in which:
[0029] Figure 1 is an example of raster-scan slice partitioning of a picture.
[0030] Figure 2 is an example of a rectangular slice partition of a picture.
[0031] Figure 3 is an example of a picture partition into slices, tiles, and rectangular slices.
[0032] Figure 4A is an example of a coding tree block (CTB) that crosses a bottom picture boundary.
[0033] Figure 4B is an example of a CTB that crosses a right picture boundary.
[0034] Figure 4C is an example of a CTB that crosses a bottom-right picture boundary.
[0035] Figure 5 is an example of an encoder.
[0036] Figure 6 is an illustration of samples within an 8x8 sample block.
[0037] Figure 7 is an example of pixels involved in filter on / off decision and strong / weak filter selection.
[0038] Figure 8 shows four one-dimensional (1-D) directional schemes for edge offset (EO) sample classification.
[0039] Figure 9 shows an example of a geometric-alternative loop filter (GALF) filter shape.
[0040] Figure 10 shows an example of relative coordinates for 5x5 diamond filter support.
[0041] Figure 11 shows another example of relative coordinates for 5x5 diamond filter support.
[0042] Figure 12A is an example architecture of the proposed CNN filter.
[0043] Figure 12B is an example of a residual block (ResBlock) construction.
[0044] Figure 13 shows an embodiment of a video bitstream.
[0045] Figure 14 is a block diagram showing an example video processing system.
[0046] Figure 15 is a block diagram of a video processing device.
[0047] Figure 16 is a block diagram illustrating an example video coding system.
[0048] Figure 17 is a block diagram illustrating an example of a video encoder.
[0049] Figure 18 is a block diagram illustrating an example of a video decoder.
[0050] Figure 19 is a method for coding video data according to an embodiment of the disclosure.
[0051] Figure 20 is another method for coding video data according to an embodiment of the disclosure.
[0052] Figure 21 is another method for coding video data according to an embodiment of the disclosure. DETAILED DESCRIPTION
[0053] It should be understood at the outset that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or in existence. The disclosure should in no way be limited to the illustrative implementations, drawings, and techniques illustrated below, including the exemplary design and implementation illustrated and described herein, but can be modified in any manner within the scope of the appended claims.
[0054] The use of H.266 terminology in certain descriptions is merely for ease of understanding and is not intended to limit the scope of the disclosed technology. Thus, the technology described herein is applicable to other video codec protocols and designs as well.
[0055] Video coding standards have evolved primarily through the development of the well-known International Telecommunication Union - Telecommunication (ITU-T) and International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) standards. The ITU-T developed H.261 and H.263, ISO / IEC developed Motion Pictures Expert Group (MPEG)-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 Video and H.263 / MPEG-4 Advanced Video Coding (AVC) and H.265 / High Efficiency Video Coding (HEVC) standards.
[0056] Since H.262, video coding standards are based on a hybrid video coding structure, in which temporal prediction plus transform coding is employed. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. Since then, many new methods have been adopted by JVET and applied to the reference software named Joint Exploration Model (JEM).
[0057] In April 2018, the Joint Video Expert Team (JVET) was created between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG), which is committed to researching the multi-functional video coding (VVC) standard, aiming to reduce 50% of bit rate compared with HEVC. VVC version 1 was completed in July 2020.
[0058] Color spaces and chroma subsampling are discussed. A color space, also called a color model (or color system), is an abstract mathematical model that simply describes a range of colors as a tuple of numbers, usually 3 or 4 values or color components (e.g., red green blue (RGB)). Fundamentally, a color space is an exposition of a coordinate system and subspace.
[0059] For video compression, the most commonly used color spaces are YCbCr and RGB. Y’CbCr or Y Pb / Cb Pr / Cr, also written as YCBCR or Y’CBCR, is a family of color spaces used as part of the color image pipeline in video and digital photography systems. Y’ is the luminance component, CB and CR are the blue-difference and red-difference chrominance components. Y’ (with the prime) is different from Y, which is luminance, meaning that the light intensity is based on gamma-corrected RGB primary non-linear encoding.
[0060] Chroma subsampling is the practice of encoding an image with lower precision for chrominance information than for luminance, exploiting the fact that the human visual system is less sensitive to color differences than to luminance differences.
[0061] For 4:4:4 chroma subsampling, each of the three Y’CbCr components has the same sampling rate, so there is no chroma subsampling. This scheme is sometimes used for high-end film scanners and movie post-production.
[0062] For 4:2:2 chroma subsampling, the two chrominance components are sampled at half the luminance sampling rate: horizontal chroma precision is halved. This reduces the bandwidth of the uncompressed video signal by one third, but there is little visual difference.
[0063] For 4:2:0 chroma subsampling, the horizontal sampling is doubled compared to 4:1 :1, but the vertical precision is halved since in this scheme the Cb and Cr channels are only sampled on every other line. Thus, the data rate is the same. Cb and Cr are subsampled by a factor of 2 in the horizontal and vertical direction, respectively. There are three variants of the 4:2:0 scheme with different horizontal and vertical addressing.
[0064] In MPEG-2, Cb and Cr are co-located in the horizontal direction. Cb and Cr are addressed between pixels in the vertical direction (gap addressing). In Joint Photographic Experts Group (JPEG) / JPEG File Interchange Format (JFIF), H.261 and MPEG-1, Cb and Cr are gap addressed, located in the middle of the alternate luma samples. In 4:2:0 DV, Cb and Cr are co-located in the horizontal direction. In the vertical direction, they are co-located on alternate lines.
[0065] Definitions of video units are provided. A picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of coding tree units (CTUs) covering a rectangular region of a picture. A tile is divided into one or more bricks, each consisting of multiple CTU rows within the tile. A tile that is not divided into multiple bricks is also referred to as a brick. However, a brick that is a proper subset of a tile is not referred to as a tile. A slice includes multiple tiles of a picture or multiple bricks of a tile.
[0066] Two slice modes are supported, namely, a raster-scan slice mode and a rectangular slice mode. In the raster-scan slice mode, a slice includes a sequence of tiles in a raster-scan of tiles of a picture. In the rectangular slice mode, a slice includes multiple bricks of a picture that collectively form a rectangular region of the picture. The bricks within a rectangular slice are arranged in the order of brick raster-scan of the slice.
[0067] Figure 1 An example of raster-scan slice partitioning for picture 100, in which the picture is divided into twelve tiles 102 and three raster-scan slices 104. As shown, each tile 102 and slice 104 includes multiple CTUs 106.
[0068] Figure 2 An example of rectangular slice partitioning for picture 200 according to the VVC specification, in which the picture is divided into twenty-four tiles 202 (six tile columns 203 and four tile rows 205) and nine rectangular slices 204. As shown, each tile 202 and slice 204 includes multiple CTUs 206.
[0069] Figure 3An example of partitioning picture 300 into slices, tiles, and rectangular slices according to the VVC specification, where the picture is divided into four slices 302 (two slice columns 303 and two slice rows 305), eleven tiles 304 (one tile in the top-left slice, five tiles in the top-right slice, two tiles in the bottom-left slice, and three tiles in the bottom-right slice), and four rectangular slices 306.
[0070] CTU and coding tree block (CTB) sizes are discussed. In VVC, the coding tree unit (CTU) size, signaled in the sequence parameter set (SPS) by the syntax element log2_ctu_size_minus2, can be as small as 4x4. The sequence parameter set (SPS) syntax is as follows.
[0071]
[0072]
[0073] log2_ctu_size_minus2 plus 2 specifies the luma coding tree block size of each CTU.
[0074] log2_min_luma_coding_block_size_minus2 plus 2 specifies the minimum luma coding block size.
[0075] The variables CtbLog2SizeY, CtbSizeY, MinCbLog2SizeY, MinCbSizeY, MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, MaxTbSizeY, PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC are derived as follows:
[0076] CtbLog2SizeY = log2_ctu_size_minus2 + 2 (7-9)
[0077] CtbSizeY = 1 « CtbLog2SizeY (7-10)
[0078] MinCbLog2SizeY = log2_min_luma_coding_block_size_minus2 + 2 (7-11)
[0079] MinCbSizeY = 1 « MinCbLog2SizeY (7-12)
[0080] MinTbLog2SizeY = 2 (7-13)
[0081] MaxTbLog2SizeY = 6 (7-14)
[0082] MinTbSizeY = 1 « MinTbLog2SizeY (7-15)
[0083] MaxTbSizeY = 1 « MaxTbLog2SizeY (7-16)
[0084] PicWidthInCtbsY = Ceil( pic_width_in_luma_samples ÷ CtbSizeY ) (7-17)
[0085] PicHeightInCtbsY = Ceil( pic_height_in_luma_samples ÷ CtbSizeY ) (7-18)
[0086] PicSizeInCtbsY = PicWidthInCtbsY * PicHeightInCtbsY (7-19)
[0087] PicWidthInMinCbsY = pic_width_in_luma_samples ÷ MinCbSizeY (7-20)
[0088] PicHeightInMinCbsY = pic_height_in_luma_samples ÷ MinCbSizeY (7-21)
[0089] PicSizeInMinCbsY = PicWidthInMinCbsY * PicHeightInMinCbsY (7-22)
[0090] PicSizeInSamplesY = pic_width_in_luma_samples * pic_height_in_luma_samples (7-23)
[0091] PicWidthInSamplesC = pic_width_in_luma_samples ÷ SubWidthC (7-24)
[0092] PicHeightInSamplesC = pic_height_in_luma_samples / SubHeightC (7-25)
[0093] Figure 4A is an example of a CTB that crosses the bottom picture boundary. Figure 4B is an example of a CTB that crosses the right picture boundary. Figure 4C is an example of a CTB that crosses the bottom-right picture boundary. In Figures 4A-4C , there are K=M, L<N; K
[0094] Reference is made to Figures 4A-4C CTUs in picture 400 are discussed. Assume that the CTB / maximum coding unit (LCU) size is denoted by MxN (typically M is equal to N as defined in HEVC / VVC), and for a CTB located at the picture (or slice or tile or other type, taking picture boundaries as an example) boundary, KxL samples are within the picture boundary, where K Figures 4A-4C For those CTBs 402 depicted in
[0095] The coding flow of a typical video encoder / decoder (a.k.a. codec) is discussed. Figure 5 is an example of an encoder block diagram of VVC, which contains three in-loop filtering blocks: the de-blocking filter (DF), the sample adaptive offset (SAO), and the adaptive loop filter (ALF). Unlike DF, which uses a pre-defined filter, SAO and ALF exploit the original samples of the current picture by means of coded side information signaling the offsets and filter coefficients to reduce the mean square error between the original and reconstructed samples, respectively, by adding offsets and applying a finite impulse response (FIR) filter. ALF is located at the last processing stage of each picture and can be seen as a tool trying to capture and fix artifacts established by previous stages.
[0096] Figure 5An encoder 500 is shown in a schematic diagram. The encoder 500 is suitable for implementing the VVC technique. The encoder 500 comprises three in-loop filters, namely a deblocking filter (DF) 502, a sample adaptive offset (SAO) filter 504 and an ALF 506. Unlike the DF 502 which uses a pre-defined filter, the SAO filter 504 and the ALF 506 exploit the original samples of the current picture by means of signaling the coding side information of the offsets and filter coefficients to reduce the mean square error between the original and reconstructed samples through adding offsets and applying a FIR filter respectively. The ALF 506 is located at the last processing stage of each picture and can be regarded as a tool trying to capture and fix the artifacts established by previous stages.
[0097] The encoder 500 further comprises an intra prediction component 508 and a motion estimation / compensation (ME / MC) component 510 configured to receive an input video. The intra prediction component 508 is configured to perform intra prediction while the ME / MC component 510 is configured to perform inter prediction with reference pictures obtained from a reference picture buffer 512. The residual blocks from either inter or intra prediction are fed to a transform component 514 and a quantization component 516 to generate quantized residual transform coefficients which are fed to an entropy coding component 518. The entropy coding component 518 entropy codes the prediction results and the quantized transform coefficients and sends them to a video decoder (not shown). The quantized components output from the quantization component 516 can be fed to a dequantization component 520, an inverse transform component 522 and a reconstruction (REC) component 524. The REC component 524 is capable of outputting pictures to the DF 502, the SAO filter 504 and the ALF 506 for filtering before the pictures are stored in the reference picture buffer 512.
[0098] The input to the DF 502 is the reconstructed samples before the in-loop filters. First, vertical edges in the picture are filtered. Then, horizontal edges in the picture are filtered with the samples modified by the vertical edge filtering process as input. The vertical and horizontal edges in the CTBs of each CTU are processed individually on a coding unit basis. The vertical edges of the coding blocks in a coding unit are filtered starting from the left edge of the coding blocks and proceeding through the edges to the right side of the coding blocks in their geometric order. The horizontal edges of the coding blocks in a coding unit are filtered starting from the top edge of the coding blocks and proceeding through the edges to the bottom side of the coding blocks in their geometric order.
[0099] Figure 6 A diagram 600 of a sample 602 within an 8x8 sample block 604. As shown, the diagram 600 includes horizontal block boundaries 606 and vertical block boundaries 608 on an 8x8 grid, respectively. In addition, the diagram 600 depicts an 8x8 sample non-overlapping block 610, which can be deblocked in parallel.
[0100] The boundary decision is discussed. The filter is applied to the 8x8 block boundary. In addition, it must be a transform block boundary or a coded sub-block boundary (e.g., due to the use of affine motion prediction, optional temporal motion vector prediction (ATMVP)). For those boundaries that are not such boundaries, the filter is disabled.
[0101] The boundary strength calculation is discussed. For a transform block boundary / coded sub-block boundary, if it is located in an 8x8 block boundary, the transform block boundary / coded sub-block boundary can be filtered, and the bS[xD i ][yD j ](where [xD i ][yD j ] denotes the coordinates) of this edge is defined in Table 1 and Table 2, respectively.
[0102] Table 1. Boundary strength (when SPS Intra Block Copy (IBC) is disabled)
[0103]
[0104] Table 2. Boundary strength (when SPS IBC is enabled)
[0105]
[0106]
[0107] The deblocking decision for the luma component is discussed.
[0108] Figure 7 Example 700 of pixels involved in filter on / off decision and strong / weak filter selection. The wider stronger luma filter is used only when condition 1, condition 2 and condition 3 are all true. Condition 1 is the “large block condition”. This condition detects whether the samples on the P side and the Q side belong to a large block, denoted by variables bSidePisLargeBlk and bSideQisLargeBlk, respectively. The definitions of bSidePisLargeBlk and bSideQisLargeBlk are as follows.
[0109] bSidePisLargeBlk = ((edge type is vertical, and p0 belongs to a CU with width >= 32) || (edge type is horizontal, and p0 belongs to a CU with height >= 32))? TRUE : FALSE
[0110] bSideQisLargeBlk = ((edge type is vertical, and q0 belongs to a CU with width >= 32) || (edge type is horizontal, and q0 belongs to a CU with height >= 32))? TRUE : FALSE
[0111] Based on bSidePisLargeBlk and bSideQisLargeBlk, condition 1 is defined as follows.
[0112] Condition 1 = (bSidePisLargeBlk || bSideQisLargeBlk)? TRUE : FALSE
[0113] Next, if condition 1 is true, condition 2 will be further checked. First, the following variables are derived.
[0114] In HEVC, dp0, dp3, dq0, dq3 are first derived.
[0115] If (p-side is greater than or equal to 32)
[0116] dp0 = (dp0 + Abs(p50 - 2*p40 + p30) + 1) » 1
[0117] dp3 = (dp3 + Abs(p53 - 2*p43 + p33) + 1) » 1
[0118] If (q-side is greater than or equal to 32)
[0119] dq0 = (dq0 + Abs(q50 - 2*q40 + q30) + 1) » 1
[0120] dq3 = (dq3 + Abs(q53 - 2*q43 + q33) + 1) » 1
[0121] Condition 2 = (d < β)? TRUE : FALSE
[0122] Where d = dp0 + dq0 + dp3 + dq3.
[0123] If condition 1 and condition 2 are valid, it further checks whether any block uses sub-block.
[0124]
[0125]
[0126] Finally, if both condition 1 and condition 2 are valid, the proposed deblocking method will check condition 3 (large block strong filter condition), which is defined as follows.
[0127] In condition 3 StrongFilterCondition, the following variables are derived.
[0128]
[0129]
[0130] As in HEVC, StrongFilterCondition = (dpq < (β » 2), sp3 + sq3 < (3 * β » 5), and Abs(p0 - q0) < (5 * t C + 1) » 1)? TRUE: FALSE.
[0131] The strong de-blocking filter for luma (designed for larger blocks) is discussed.
[0132] When the samples on either side of the boundary belong to large blocks, bilinear filtering is used. The samples belonging to large blocks are defined when the width of the vertical edge is ≥ 32 and the height of the horizontal edge is ≥ 32.
[0133] The bilinear filter is listed as follows.
[0134] Then, in the above HEVC de-blocking, for block boundary samples p i and q i are replaced by linear interpolation, p i and q i are for filtering the i-th sample in a row of the vertical edge, or the i-th sample in a column of the horizontal edge, as follows.
[0135] p i ′ = (f i * Middle s,t + (64 - f i ) * P s + 32) » 6), clipped to p i ± tcPD i
[0136] q j ′ = (g j * Middle s,t + (64 - g j ) * Q s + 32) » 6), clipped to q j ± tcPD j
[0137] where tcPD i and tcPD j The terms are the position dependent clippings described below and g j , f i , Middle s,t , P s and Q s are given below.
[0138] The control of chroma deblocking is discussed.
[0139] A strong chroma filter is used on both sides of the block boundary. Here, the chroma filter is selected when both sides of the chroma edge are greater than or equal to 8 (chroma position) and the decision of the following three conditions is satisfied. The first decision considers the boundary strength and the decision of the large block. The proposed filter can be applied when the block width or height in the chroma sample domain orthogonal to the block edge is equal to or greater than 8. The second and third decisions are basically the same as the HEVC luma deblocking decisions, which are the on / off decision and the strong filter decision, respectively.
[0140] In the first decision, the boundary strength (bS) is modified for chroma filtering and the conditions are checked sequentially. If one condition is satisfied, the rest of the lower priority conditions are skipped.
[0141] The chroma deblocking is performed when bS is equal to 2 or bS is equal to 1 when a large block boundary is detected.
[0142] The second and third conditions are basically the same as the HEVC luma strong filter decisions, which are shown as follows.
[0143] In the second condition, d is derived as in the HEVC luma deblocking. The second condition is TRUE when d is less than β.
[0144] In the third condition, StrongFilterCondition is derived as follows.
[0145] sp3 = Abs(p3 - p0), derived as in HEVC
[0146] sq3 = Abs(q0 - q3), derived as in HEVC
[0147] As in the HEVC design, StrongFilterCondition = (dpq is less than (β » 2), sp3 + sq3 is less than (β » 3), Abs(p0 - q0) is less than (5 * t C + 1) » 1).
[0148] The strong deblocking filter for chroma is discussed. The following strong deblocking filter for chroma is defined.
[0149] p2' = (3 * p3 + 2 * p2 + p1 + p0 + q0 + 4) » 3
[0150] p1' = (2 * p3 + p2 + 2 * p1 + p0 + q0 + q1 + 4) » 3
[0151] p0' = (p3 + p2 + p1 + 2 * p0 + q0 + q1 + q2 + 4) » 3
[0152] The proposed chroma filter performs deblocking on a 4x4 chroma sample grid.
[0153] Position dependent clipping (tcPD) is discussed. Position dependent clipping tcPD is applied to the output samples of the luma filter process which involves modifying the strong and long filters for 7, 5 and 3 samples at the boundaries. Assuming a quantization error distribution, it is proposed to increase the clipping values for samples which are expected to have higher quantization noise, hence a larger deviation of the reconstructed sample value from the true sample value.
[0154] For each P or Q boundary filtered with asymmetric filters, a position dependent threshold table is selected from two tables (i.e. Tc7 and Tc3 in the table below) according to the result of the decision process in the boundary strength calculation and provided to the decoder as side information.
[0155] Tc7 = {6, 5, 4, 3, 2, 1, 1}; Tc3 = {6, 4, 2};
[0156] tcPD = (Sp == 3)? Tc3 : Tc7;
[0157] tcQD = (Sq == 3)? Tc3 : Tc7;
[0158] For P or Q boundaries filtered with short symmetric filters, lower amplitude position dependent thresholds are applied.
[0159] Tc3 = {3, 2, 1};
[0160] After defining the thresholds, the filtered p' i q' i sample values are clipped according to the tcP and tcQ clipping values.
[0161] p” i = Clip3(p' i + tcP i , p' i - tcP i , p' i );
[0162] q” j = Clip3(q' j + tcQ j , q' j - tcQ j , q' j );
[0163] where p' i and q' i are the filtered sample values p" i and q" j are the clipped output sample values, tcPi tcQ i is the clipping threshold derived from the VVC tc parameter and tcPDand tcQD. The function Clip3 is a clipping function specified in VVC.
[0164] Subblock deblocking adjustment is discussed. To enable parallel friendly deblocking using long filter and subblock deblocking, the long filter is restricted to modify at most 5 samples on the side using subblock deblocking (AFFINE or ATMVP or decoder side motion vector refinement (DMVR)) as shown in the luma control of long filter. In addition, subblock deblocking is adjusted such that subblock boundaries on 8x8 grid close to coding unit (CU) or implicit TU boundaries are restricted to modify at most two samples on each side.
[0165] The following applies to subblock boundaries that are not aligned with CU boundaries.
[0166]
[0167] where edge equal to 0 corresponds to a CU boundary, edge equal to 2 or equal to orthogonalLength - 2 corresponds to a subblock boundary 8 samples away from a CU boundary, and so on, where implicit TU is true if implicit partitioning of TUs is used.
[0168] Sample adaptive offset (SAO) is discussed. The input to SAO is the deblocked reconstructed samples (DB). The concept of SAO is to reduce the average sample distortion of a region by first classifying the region samples into multiple categories with a selected classifier, obtaining an offset for each category, and then adding the offset to each sample of the category, where the classifier index and the offset of a region are coded in the bitstream. In HEVC and VVC, a region (a unit signaled by SAO parameters) is defined as a CTU.
[0169] HEVC employs two types of SAO that can satisfy low complexity requirement. The two types are edge offset (EO) and band offset (BO), which will be discussed in detail below. The index of SAO type is coded (in the range of [0, 2]). For EO, the sample classification is based on the comparison between the current sample and the neighboring samples, according to a one-dimensional (1-D) directional scheme, such as horizontal, vertical, 135° diagonal, and 45° diagonal.
[0170] Figure 8 Four 1-D directional schemes 800 for EO sample classification are shown: horizontal (EO classification = 0), vertical (EO classification = 1), 135° diagonal (EO classification = 2), and 45° diagonal (EO classification = 3).
[0171] For a given EO class, each sample within the CTB is classified into one of five categories. The current sample value, denoted as "c", is compared with its two neighboring sample values along the selected 1-D pattern. The classification rule for each sample is summarized in Table 3. Categories 1 and 4 are associated with local valleys and local peaks along the selected 1-D pattern, respectively. Categories 2 and 3 are associated with concave and convex corners along the selected 1-D pattern, respectively. If the current sample does not belong to EO classes 1-4, it belongs to class 0 and SAO is not applied.
[0172] Table 3: Sample classification rule for edge offset
[0173]
[0174] A geometric transform based adaptive loop filter in the Joint Exploration Model (JEM) is discussed. The input to the DB is the reconstructed sample after DB and SAO. The sample classification and filtering process is based on the reconstructed sample after DB and SAO.
[0175] In JEM, a geometric transform based adaptive loop filter (GALF) with block based filter adaptation is applied. For the luma component, one of 25 filters is selected for each 2x2 block according to the direction and significance of the local gradient.
[0176] The shape of the filter is discussed. Figure 9 An example of the GALF filter shape 900 is shown, including a 5x5 diamond on the left, a 7x7 diamond on the right, and a 9x9 diamond in the middle. In JEM, up to three diamond filter shapes can be selected for the luma component (as shown in Figure 9 The index is signaled at the picture level to indicate the filter shape used for the luma component. Each square represents a sample, and Ci (i is 0-6 (left), 0-12 (right), 0-20 (middle)) represents the coefficient applied to that sample. For the chroma components in the picture, a 5x5 diamond is always used.
[0177] Block classification is discussed. Each 2x2 block is classified into one of 25 classes. The classification index c is derived based on its directionality D and significance of the quantized values of the gradients in the horizontal, vertical and two diagonal directions, as follows.
[0178]
[0179] To compute D and The gradients in the horizontal, vertical and two diagonal directions are first computed using 1-D Laplacian operators.
[0180]
[0181]
[0182] H k,l = |2R(k,l) - R(k-1,l) - R(k+1,l)|,
[0183]
[0184]
[0185] The indices i and j refer to the coordinates of the top-left sample in the 2x2 block, and R(i,j) denotes the reconstructed sample at coordinate (i,j).
[0186] The maximum and minimum of the horizontal and vertical gradients are then set to:
[0187]
[0188] The maximum and minimum of the two diagonal gradients are set to:
[0189]
[0190] To derive the value of the directionality D, these values are compared to each other and to two thresholds t1 and t2:
[0191] Step 1. If and are true, then D is set to 0.
[0192] Step 2. If continue with step 3; otherwise continue with step 4.
[0193] Step 3. If D is set to 2; otherwise D is set to 1.
[0194] Step 4. If D is set to 4; otherwise D is set to 3.
[0195] The significance value A is computed as follows:
[0196]
[0197] A is further quantized to the range 0 to 4, and the quantized value is denoted as
[0198] For two chroma components in a picture, the classification method is not applied, i.e. a single set of ALF coefficients is applied for each chroma component.
[0199] A geometric transformation of the filter coefficients is discussed.
[0200] Figure 10An example of relative coordinates 1000 supporting 5x5 diamond filters such as (from left to right) diagonal, vertical flip and rotation respectively is shown.
[0201] Before filtering each 2x2 block, a geometric transform such as rotation or diagonal anti-flip and vertical flip is applied to the filter coefficients f(k,l) associated with coordinates (k,l) depending on the gradient value computed for this block. This is equivalent to applying these transforms to the samples in the filter support region. The idea is to make the different blocks for which ALF is applied more similar by aligning their directionality.
[0202] Three geometric transforms are introduced, including diagonal, vertical flip and rotation:
[0203] Diagonal: f D (k,l) = f(l,k),
[0204] Vertical flip: f V (k,l) = f(k,K-l-1), (9)
[0205] Rotation: f R (k,l) = f(K-l-1,k).
[0206] where K is the size of the filter and 0≤k,l≤K-1 are the coefficient coordinates such that position (0,0) is at the top-left corner and position (K-1,K-1) is at the bottom-right corner. The transform is applied to the filter coefficients f(k,l) depending on the gradient value computed for this block. Table 4 summarizes the relationship between the transform and the four gradients for the four directions.
[0207] Table 1: Mapping between the gradient computed for a block and the transform
[0208]
[0209] Signaling of filter parameters is discussed. In JEM, GALF filter parameters are signaled for the first CTU, i.e. after the slice header of the first CTU and before the SAO parameters. Up to 25 sets of luma filter coefficients can be sent. To reduce the bit overhead, filter coefficients of different categories can be merged. In addition, the GALF coefficients of a reference picture are stored and allowed to be reused as the GALF coefficients of the current picture. The current picture can choose to use the GALF coefficients stored for a reference picture and bypass the GALF coefficient signaling. In this case, only the index of the reference picture is signaled and the stored GALF coefficients of the indicated reference picture are inherited by the current picture.
[0210] To support GALF temporal prediction, a candidate list of GALF filter sets is maintained. At the beginning of the decoding of a new sequence, the candidate list is empty. After the decoding of a picture, the corresponding filter set can be added to the candidate list. Once the size of the candidate list reaches the maximum allowed value (i.e. 6 in the current JEM), the new filter set overwrites the oldest set in decoding order, that is, the first-in first-out (FIFO) rule is applied to update the candidate list. To avoid duplication, a set is only added to the list if the corresponding picture does not use GALF temporal prediction. To support temporal scalability, there are multiple candidate lists of filter sets and each candidate list is associated with a temporal layer. More specifically, each array assigned by a temporal layer index (TempIdx) can consist of filter sets with previously decoded pictures equal to lower TempIdx. For example, the k-th array is assigned to be associated with TempIdx equal to k, and the k-th array only contains filter sets from pictures with TempIdx less than or equal to k. After a certain picture is coded, the filter set associated with the picture will be used to update those arrays associated with TempIdx equal to or higher.
[0211] Temporal prediction of GALF coefficients is used for inter-coded frames to minimize the signaling overhead. For intra frames, temporal prediction is not available and a set of 16 fixed filters is assigned to each category. To indicate the use of a fixed filter, a flag for each category is signaled and, if needed, the index of the selected fixed filter is also signaled. Even when a fixed filter is selected for a given category, the coefficients of the adaptive filter f(k,l) can still be transmitted for that category, in which case the coefficients of the filter to be applied to the reconstructed image are the sum of the two sets of coefficients.
[0212] The filtering process for the luma component can be controlled at the CU level. A flag is signaled to indicate whether GALF is applied to the luma component of a CU. For the chroma components, whether to apply GALF is only indicated at the picture level.
[0213] The filtering process is discussed. At the decoder side, when GALF is enabled for a block, each sample R(i,j) within the block is filtered, resulting in a sample value R'(i,j) as shown below, where L denotes the filter length and f(k,l) denotes the decoded filter coefficients.
[0214]
[0215] Figure 11 Another example of relative coordinates for a 5x5 diamond filter support is shown assuming the coordinates (i,j) of the current sample are (0,0). The samples in different coordinates that are filled with the same color are multiplied by the same filter coefficient.
[0216] The geometry-based adaptive loop filter (GALF) in VVC is discussed. In VVC test model 4.0 (VTM4.0), the filtering process of the adaptive loop filter is performed as follows:
[0217] O(x, y) = ∑ (i,j) w(i, j).I(x + i, y + j), (11)
[0218] where the sample I(x + i, y + j) is the input sample, O(x, y) is the filtered output sample (i.e., the filtering result), and w(i, j) represents the filter coefficient. In fact, VTM4.0 is using integer operations to implement fixed-point precision calculation
[0219]
[0220] where L represents the filter length, and w(i, j) is the filter coefficient of fixed-point precision.
[0221] Compared with JEM, the current design of GALF in VVC has the following main changes:
[0222] 1) The adaptive filter shape is removed. The luma component only allows 7x7 filter shape, and the chroma component only allows 5x5 filter shape.
[0223] 2) The signaling of ALF parameters is moved from the slice / picture level to the CTU level.
[0224] 3) The calculation of the class index is performed at the 4x4 level instead of the 2x2 level. In addition, as proposed in JVET-L0147, the sub-sampling Laplacian calculation method for ALF classification is utilized. More specifically, there is no need to calculate the horizontal / vertical / 45-degree diagonal / 135-degree gradient for each sample within a block. Instead, 1:2 sub-sampling is used.
[0225] Regarding the filtering reconstruction, the non-linear ALF in the current VVC is discussed.
[0226] Equation (11) can be re-expressed as the following expression without affecting the coding efficiency:
[0227] O(x, y) = I(x, y) + ∑ (i,j)≠(0,0) w(i, j).(I(x + i, y + j) - I(x, y)), (2)
[0228] where w(i, j) is the same filter coefficient as in equation (11) [except that w(0, 0) is equal to 1 in equation (13), while it is equal to 1 - ∑ (i,j)≠(0,0) w(i, j) in equation (11)].
[0229] Using the above filter formula of equation (13), VVC introduced non-linearity to make ALF more effective by using a simple clipping function to reduce the impact when the neighboring sample values (I(x+i, y+j)) differ too much from the filtered current sample value (I(x, y)).
[0230] More specifically, the ALF filter is modified as follows:
[0231] O'(x, y) = I(x, y) +∑ (i,)≠(0,0) w(i, j).K(I(x+i, y+j)-I(x, y), k(i, j)), (14)
[0232] where K(d, b) = min(b, max(-b, d)) is the clipping function and k(i, j) is the clipping parameter depending on the (i, j) filter coefficient. The encoder performs an optimization to find the best k(i, j).
[0233] In the JVET-N0242 implementation, a clipping parameter k(i, j) is specified for each ALF filter and one clipping value is signaled for each filter coefficient. This means that up to 12 clipping values can be signaled in the bitstream for each luma filter and up to 6 clipping values for each chroma filter.
[0234] To limit the signaling cost and encoder complexity, only 4 fixed values are used, which are the same for INTER and INTRA slices.
[0235] Because the difference of local differences of luma filters is usually higher than the difference of local differences of chroma filters, two different sets of luma and chroma filters are applied. A maximum sample value (here 1024 for 10 bit depth) in each set is also introduced so that clipping can be disabled when not necessary.
[0236] Table 5 provides the set of clipping values used in the JVET-N0242 tests. These 4 values are chosen by roughly equally dividing the full range of sample values (coded in 10 bits) for luma filters and the range from 4 to 1024 for chroma filters in the log domain.
[0237] More precisely, the luma filter table of clipping values is obtained by the following formula:
[0238]
[0239] Similarly, the chroma filter table of clipping values is obtained by the following formula:
[0240]
[0241] Table 5: authorized clipping values
[0242]
[0243] The selected clipping values are coded in the "alf_data" syntax element by using a Golomb coding scheme corresponding to the clipping value index in the above table 5. This coding scheme is the same as the one used for the filter index.
[0244] Convolutional neural network based loop filters for video coding are discussed.
[0245] In deep learning, a convolutional neural network (CNN or ConvNet) is a class of deep neural networks, most commonly applied to analyzing visual imagery. They have very successful applications in image and video recognition / processing, recommender systems, image classification, medical image analysis, and natural language processing.
[0246] CNNs are regularized versions of multilayer perceptrons. Multilayer perceptrons generally mean fully connected networks, i.e., each neuron in one layer is connected to all neurons in the next layer. The "full connectivity" of these networks makes them prone to overfitting the data. Typical regularization methods include adding some form of weight magnitude measure to the loss function. CNNs take a different approach to regularization. CNNs exploit the hierarchical scheme in the data and use smaller and simpler schemes to assemble more complex methods. Thus, CNNs are at the lower end of the spectrum in terms of connectivity and complexity.
[0247] Compared to other image classification / processing algorithms, CNNs use relatively little preprocessing. This means that the network learns the filters that are hand-designed in traditional algorithms. This independence from existing knowledge and manual effort in feature design is a major advantage.
[0248] Image / video compression based on deep learning generally has two meanings and types: 1) end-to-end compression based on neural networks, and 2) traditional framework enhanced by neural networks. End-to-end compression based on neural networks is discussed in Johannes Balle, Valero Laparra, and Eero P. Simoncelli, “End-to-end optimization of nonlinear transformations for perceptual quality-based rate-distortion code,” in Picture Coding Symposium (PCS), 2016, pp. 1-5, Institute of Electrical and Electronics Engineers (IEEE), and Lucas Theis, Wenzhe Shi, Andrew Cunningham, and Ferenc Huszar, “Lossy image compression with compressive autoencoders,” arXiv preprint arXiv: 1703.00395 (2017). Traditional framework enhanced by neural networks is discussed in Jiahao Li, Bin Li, Jizheng Xu, Ruiqin Xiong, and Wen Gao, “Image coding intra prediction based on fully connected network,” IEEE Transactions on Image Processing, 27, 7 (2018), 3236-3247, Yuanying Dai, Dong Liu, and Feng Wu, “Convolutional neural network method of post-processing in HEVC intra coding,” Multimedia Tools and Applications, Springer, 28-39, Rui Song, Dong Liu, Houqiang Li, and Feng Wu, “Neural network based arithmetic coding for intra prediction mode in HEVC,” VCIP. IEEE, 1-4, and J. Pfaff, P. Helle, D. Maniry, S. Kaltenstadler, W. Samek, H. Schwarz, D. Marpe, and T. Wiegand, “Neural network-based intra prediction for video coding,” Digital Image Processing Applications XLI, Vol. 10752, International Society for Optics and Photonics, 1075213.
[0249] The first type usually adopts an autoencoder-like structure, implemented by either convolutional neural networks or recurrent neural networks. Although relying solely on neural networks for image / video compression can avoid any manual optimization or hand-crafted design, the compression efficiency can not be satisfactory. Therefore, researches devoted to the second type aim to enhance the traditional compression framework by replacing or augmenting certain modules with the aid of neural networks. In this way, they can inherit the advantages of the highly optimized traditional framework. For example, the fully connected network for intra prediction proposed in HEVC discussed in Jiahao Li, Bin Li, Jizheng Xu, Ruiqin Xiong, and Wen Gao, “Fully connected network based intra prediction for image coding,” IEEE Transactions on Image Processing 27, 7 (2018), pp. 3236-3247.
[0250] In addition to intra prediction, deep learning based image / video compression is also used to enhance other modules. For example, the in-loop filter of HEVC is replaced by a convolutional neural network, and satisfactory results are achieved in Yuanying Dai, Dong Liu, and Feng Wu, “Convolutional neural network approach to post-processing in HEVC intra coding,” MMM. Springer, 28-39. Research in Rui Song, Dong Liu, Houqiang Li, and Feng Wu, “Neural network based arithmetic coding for HEVC intra prediction mode,” VCIP. IEEE, 1-4 applies neural networks to improve the arithmetic coding engine.
[0251] In-loop filtering based on convolutional neural networks is discussed. In lossy image / video compression, the reconstructed frame is an approximation of the original frame because the quantization process is irreversible, resulting in distortion of the reconstructed frame. To mitigate this distortion, a convolutional neural network can be trained to learn the mapping from the distorted frame to the original frame. In practice, training must be performed before deploying the CNN-based in-loop filter.
[0252] Training is discussed. The goal of the training process is to find the best values of the parameters, including weights and biases.
[0253] First, the codec (e.g., HM, JEM, VTM, etc.) is used to compress the training dataset to generate distorted reconstructed frames. Then, the reconstructed frames are input to the CNN, and the cost is calculated using the output of the CNN and the true frame (original frame). Commonly used cost functions include the sum of absolute differences (SAD) and the mean squared error (MSE). Next, the gradient of the cost with respect to each parameter is derived by the backpropagation algorithm. With the gradient, the parameter values can be updated. The above process is repeated until the convergence criterion is met. After completing the training, the best parameters derived are saved for the inference phase.
[0254] The convolution process is discussed. In the convolution process, the filter moves over the image from left to right and from top to bottom, changing one column of pixels when moving horizontally and one row of pixels when moving vertically. The amount of movement between the application of the filter to the input image is referred to as the stride, and it is almost always symmetric in the height and width dimensions. The default stride(s) in two dimensions is (1, 1) for both height and width movement.
[0255] Figure 12A is an example architecture 1200 of the proposed CNN filter, and Figure 12B is an example of the construction 1250 of a Residual Block (ResBlock). In most deep convolutional neural networks, Residual Blocks are used as the basic module and are stacked several times to build the final network, where in one example, the Residual Block is obtained by combining a convolutional layer, a ReLU / PReLU activation function, and a convolutional layer, as shown in Figure 12B .
[0256] Figure 13 An embodiment of a video bitstream 1300 is shown. As used herein, the video bitstream 1300 can also be referred to as a coded video bitstream, a bitstream, or variations thereof. As shown in Figure 13 , the bitstream 1300 includes one or more of the following: a decoding capability information (DCI) 1302, a video parameter set (VPS) 1304, a sequence parameter set (SPS) 1306, a picture parameter set (PPS) 1308, a picture header (PH) 1310, a picture 1314, and an SEI message 1322. Each of the DCI 1302, the VPS 1304, the SPS 1306, and the PPS 1308 can be collectively referred to as a parameter set. In an embodiment, other parameter sets not shown in Figure 13 , such as, for example, an adaptation parameter set (APS), which is a syntax structure containing syntax elements that apply to zero or more slices determined by zero or more syntax elements found in slice headers, can also be included in the bitstream 1300.
[0257] DCI 1302, which can also be referred to as a decoding parameter set (DPS) or decoder parameter set, is a syntax structure containing syntax elements that apply to the entire bitstream. DCI 1302 includes parameters that remain constant for the lifetime of the video bitstream (e.g., bitstream 1300), which can transition to the lifetime of a session. DCI 1302 can include profile, level, and sub-level information to determine a maximum complexity interoperability point that is guaranteed never to be exceeded, even if splicing of video sequences occurs within a session. It also optionally includes constraint flags that indicate the video bitstream will be constrained from using certain features indicated by the values of those flags. In this way, a bitstream can be marked as not using certain tools, which allows, among other things, for resource allocation in decoder implementations. Like all parameter sets, DCI 1302 is present at the first occurrence of a reference and is referenced by the first picture in a video sequence, which means it must be conveyed between the first network abstraction layer (NAL) units in the bitstream. While multiple DCI 1302 can be in bitstream 1300, the values of the syntax elements therein must be consistent at the time of reference.
[0258] VPS 1304 includes decoding dependencies or information for the construction of the reference picture set of enhancement layers. VPS 1304 provides an overall view of a scalable sequence or view, including providing what types of operation points, the profile, tier, and level of the operation points, and some other high-level properties of the bitstream that can be used as a basis for session negotiation and content selection, etc.
[0259] In embodiments, when some layers are indicated to use ILP, VPS 1304 indicates that the total number of OLSs specified by the VPS is equal to the number of layers, indicates that the i-th OLS includes layers with layer indices from 0 to i, inclusive, and indicates that for each OLS, only the highest layer in the OLS is output.
[0260] SPS 1306 contains data common to all pictures in a sequence of pictures (SOP). SPS 1306 is a syntax structure containing syntax elements that apply to zero or more complete coded layer video sequences (CLVSs) as determined by the content found in the syntax elements found in PPS 1308, which are referenced by syntax elements found in each picture header 1310. In contrast, PPS 1308 contains data common to an entire picture 1314. PPS 1308 is a syntax structure containing syntax elements that apply to zero or more complete coded pictures as determined by syntax elements found in each picture header (e.g., PH 1310).
[0261] DCI 1302, VPS 1304, SPS 1306, and PPS 1308 are contained in different types of network abstraction layer (NAL) units. NAL units are syntax structures that contain an indication of the type of data to follow (e.g., coded video data). NAL units are classified as video coding layer (VCL) and non-VCL NAL units. VCL NAL units contain data representing sample values in a video picture, and non-VCL NAL units contain any associated additional information, such as parameter sets (important data that can be applied to multiple VCL NAL units) and supplemental enhancement information (timing information and other supplemental data that can enhance the usability of the decoded video signal, but is not essential to decode the sample values in the video picture).
[0262] In embodiments, DCI 1302 is contained in a non-VCL NAL unit designated as a DCI NAL unit or a DPS NAL unit. That is, a DCI NAL unit has a DCI NAL unit type (NUT), and a DPS NAL unit has a DPS NUT. In embodiments, VPS 1304 is contained in a non-VCL NAL unit designated as a VPS NAL unit. Thus, a VPS NAL unit has a VPS NUT. In embodiments, SPS 1306 is a non-VCL NAL unit designated as an SPS NAL unit. Thus, an SPS NAL unit has an SPS NUT. In embodiments, PPS 1308 is contained in a non-VCL NAL unit designated as a PPS NAL unit. Thus, a PPS NAL unit has a PPS NUT.
[0263] PH 1310 is a syntax structure that contains syntax elements that apply to all slices (e.g., slices 1318) of a coded picture (e.g., picture 1314). In embodiments, PH 1310 is a non-VCL NAL unit designated as a PH NAL unit. Thus, a PH NAL unit has a PH NUT (e.g., PH NUT).
[0264] In embodiments, a PH NAL unit associated with PH 1310 has a temporal identifier (ID) and a layer ID. The temporal ID indicates the location in time of the PH NAL unit relative to other PH NAL units in the bitstream (e.g., bitstream 1300). The layer ID indicates the layer in which the PH NAL unit is contained. In embodiments, the temporal ID is similar to, but different from, a picture order count (POC). A POC uniquely identifies each picture in order. In a single-layer bitstream, the temporal ID and the POC will be the same. In a multi-layer bitstream, pictures in the same AU will have different POCs, but the same temporal ID.
[0265] In an embodiment, the PH NAL unit precedes the VCL NAL units containing the first slice 1318 of the associated picture 1314. This establishes the association between the PH 1310 and the slices 1318 of the picture 1314 associated with the PH 1310 without the need to have a picture header ID signaled in the PH 1310 and referenced from the slice header 1320. Thus, it can be inferred that all VCL NAL units between two PHs 1310 belong to the same picture 1314 and that the picture 1314 is associated with the first PH 1310 between the two PHs 1310. In an embodiment, the first VCL NAL unit following the PH 1310 contains the first slice 1318 of the picture 1314 associated with the PH 1310.
[0266] In an embodiment, the PH NAL unit follows a picture level parameter set (e.g., PPS) or a higher level parameter set, such as DCI 1302 (also referred to as DPS), VPS 1304, SPS 1306, PPS 1308, and the like, with a temporal ID and a layer ID that are less than the temporal ID and the layer ID of the PH NAL unit, respectively. Thus, these parameter sets are not repeated within a picture or access unit. As a result of this ordering, the PH 1310 can be resolved immediately. That is, the parameter set containing parameters related to the entire picture precedes the PH NAL unit in the bitstream. Anything containing parameters of a part of the picture follows the PH NAL unit.
[0267] In an alternative, the PH NAL unit follows a picture level parameter set and a prefix supplemental enhancement information (SEI) message, or a higher level parameter set, such as DCI 1302 (also referred to as DPS), VPS 1304, SPS 1306, PPS 1308, APS, SEI message 1322, and the like.
[0268] The picture 1314 is an array of luma samples in monochrome format, or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0269] The picture 1314 can be a frame or a field. However, in a CVS 1316, all pictures 1314 are frames or all pictures 1314 are fields. The CVS 1316 is a coded video sequence of each coded layer video sequence (CLVS) in the video bitstream 1300. Notably, the CVS 1316 and the CLVS are the same when the video bitstream 1300 includes a single layer. The CVS 1316 and the CLVS are different only when the video bitstream 1300 includes multiple layers.
[0270] Each picture 1314 includes one or more slices 1318. A slice 1318 is an integer number of complete tiles or an integer number of consecutive complete coding tree units (CTU) rows within a tile of a picture (e.g., picture 1314). Each slice 1318 is exclusively contained in a single NAL unit (e.g., VCL NAL unit). A tile (not shown) is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture (e.g., picture 1314). A CTU (not shown) is a coding tree block (CTB) of luma samples, two corresponding CTBs of chroma samples of a picture having three sample arrays, or a CTB of samples of a monochrome picture or a picture coded using three separate color planes and syntax structures for coding samples. A CTB (not shown) is an NxN block of samples for some value N, such that partitioning components into CTBs is a kind of partitioning. A block (not shown) is an MxN (M columns by N rows) array of samples (e.g., pixels), or an MxN array of transform coefficients.
[0271] In embodiments, each slice 1318 includes a slice header 1320. The slice header 1320 is a part of the coded slice 1318 that includes data elements pertaining to all tiles or CTU rows within the slice represented in the slice 1318. That is, the slice header 1320 includes information about the slice 1318, such as, for example, the slice type, which reference picture will be used, and so on.
[0272] The pictures 1314 and their slices 1318 include data associated with an image or video being encoded or decoded. Thus, the pictures 1314 and their slices 1318 can be simply referred to as the payload or data carried in the bitstream 1300.
[0273] The bitstream 1300 also includes one or more SEI messages, such as SEI messages 1322, that include supplemental enhancement information. SEI messages can include various types of data that indicate timing of video pictures, or describe various properties of the coded video or how to use or enhance the coded video. SEI messages are also defined to include arbitrary user- defined data. SEI messages do not affect the core decoding process, but can indicate how the video is suggested to be post-processed or displayed. Some other high-level properties of the video content are conveyed in the video usability information (VUI), such as an indication of the color space used to interpret the video content. As new color spaces have been developed, such as high dynamic range and wide color gamut video, additional VUI identifiers have been added to indicate them.
[0274] Those skilled in the art will understand that the bitstream 1300 can include other parameters and information in practical applications.
[0275] In an embodiment, the SEI message 1322 can be a serial digital interface (SDI) SEI message. The SDI SEI message can be used to indicate which primary layers are associated with an auxiliary layer when auxiliary information is present in the bitstream. For example, the SDI SEI message can include one or more syntax elements 1324 to indicate which primary layers are associated with an auxiliary layer when auxiliary information is present in the bitstream. A discussion of various SEI messages and syntax elements included in those SEI messages is provided below.
[0276] To address the above problems, methods summarized as follows are disclosed. The techniques should be considered examples of explaining general concepts and should not be construed as being narrow. Moreover, the techniques can be applied individually or in any combination.
[0277] Inference is discussed. In the inference phase, the reconstructed frame of distortion is fed into a CNN and processed by a CNN model whose parameters have been determined in the training phase. The input samples of the CNN can be the reconstructed samples before DB or after DB, or the reconstructed samples before SAO or after SAO, or the reconstructed samples before ALF or after ALF.
[0278] Current CNN-based in-loop filtering has the following problems. CNN-based filtering is applied in the codec loop, ignoring the processing capability of the decoder.
[0279] Techniques are disclosed herein to address one or more of the foregoing problems. For example, the disclosure provides one or more neural network (NN) filter models trained as part of an in-loop filtering technique or a filtering technique used in the post-processing phase for reducing distortion caused during compression. Samples with different characteristics are processed by different NN filter models. The disclosure sets forth how to design multiple NN filter models, how to select from the multiple NN filter models, and how to signal the selected NN filter index.
[0280] Video coding is a lossy process. A convolutional neural network (CNN) can be trained to recover details lost in the compression process. That is, an artificial intelligence (AI) process can create a CNN filter based on training data.
[0281] Different CNN filter models are best suited for different situations. The encoder and the decoder have access to multiple NN filter models, such as CNN filter models, which have been trained in advance (also referred to as pre-trained). This disclosure describes methods and techniques that allow the encoder to signal to the decoder which NN filter model to use for each video unit. The video unit can be a sequence of pictures, a picture, a slice, a tile, a brick, a sub-picture, a coding tree unit (CTU), a CTU row, a coding unit (CU), and so on. As an example, different NN filters can be used for different layers, different components (e.g., luma, chroma, Cb, Cr, and so on), different specific video units, and so on. A flag and / or an index can be signaled via one or more rules to indicate which NN filter model should be used for each video item. The NN filter model can be signaled based on a reconstruction quality level of the video unit. The reconstruction quality level of a video unit is a coding unit that indicates the quality of the video unit. For example, the reconstruction quality level of a video unit can refer to one or more of QP, bitrate, constant rate factor value, and other metrics.
[0282] Also provided is the inheritance of NN filters between parent and child nodes when a tree is used to partition a video unit.
[0283] The embodiments listed below should be considered as examples to explain the general concepts. The embodiments should not be interpreted in a narrow way. Furthermore, the embodiments can be combined in any way.
[0284] In this disclosure, the NN filter can be any kind of NN filter, such as a convolutional neural network (CNN) filter. In the following discussion, the NN filter can also be referred to as a CNN filter.
[0285] In the following discussion, the video unit can be a sequence, a picture, a slice, a tile, a brick, a sub-picture, a CTU / CTB, a CTU / CTB row, one or more CUs / coding blocks (CBs), one or more CTUs / CTBs, one or more virtual pipeline data units (VPDUs), a sub-region within a picture / slice / tile / brick. A parent video unit represents a unit that is larger than the video unit. Typically, the parent unit will contain several video units, for example, when the video unit is a CTU, the parent unit can be a slice, a CTU row, multiple CTUs, and so on. In some embodiments, the video unit can be a sample / pixel.
[0286] In this disclosure, SEI (Supplemental Enhancement Information) messages refer to any messages that help the process related to decoding, display, or other purposes. However, the decoding process does not need to construct luma or chroma samples. A conforming decoder does not need to process this information for output. Examples of SEI messages are SEI and VUI in HEVC and VVC. SEI messages can include multiple individual levels, including, for example, picture level, slice level, CTB level, and CTU level.
[0287] A discussion of model selection is provided
[0288] Example 1
[0289] 1. In a first embodiment, an indication of the selected NN filter model(s) and / or NN filter model candidates allowed for a video unit or for samples within a video unit can be presented in at least one SEI message.
[0290] a. In one example, the NN filter model can be selected from multiple NN filter model candidates for a video unit.
[0291] i. In one example, the video unit is a picture / slice. And the selected NN filter model is applied to all samples within the video unit.
[0292] ii. In one example, the video unit is a CTB / CTU. And the selected NN filter model is applied to all samples within the video unit.
[0293] b. In one example, the model selection is performed at the video unit level (e.g., at the CTU level).
[0294] i. In addition, alternatively, for a given video unit (e.g., CTU), the NN filter model index is first derived / signaled from the SEI message. Given the model index, the given video unit (e.g., CTU) will be processed based on the derived / signaled model index (e.g., by the corresponding selected NN filter model associated with the model index).
[0295] c. Alternatively, for a video unit, multiple candidates can be selected and / or signaled in the SEI message.
[0296] d. In one example, luma and chroma components in a video unit can use different sets of NN filter model candidates.
[0297] i. In one example, the same number of NN filter model candidates are allowed for luma and chroma components, but the filter model candidates are different for luma and chroma components.
[0298] ii. For example, there can be L NN filter model candidates for the luma component, and there can be C NN filter model candidates for the chroma component, where L is not equal to C, e.g., L > C.
[0299] iii. In one example, chroma components (such as Cb and Cr, or U and V) in a video unit can share the same NN filter model candidate.
[0300] e. In one example, the number of NN filter model candidates can be signaled to the decoder in an SEI message.
[0301] f. In one example, the NN filter model candidate can be different for different video units (e.g., sequence / picture / slice / tile / subpicture / CTU / CTU row / CU).
[0302] g. In one example, the NN filter model candidate can depend on the tree partition structure (e.g., dual tree or single tree).
[0303] 2. An indicator of one or several NN filter model indices can be signaled in at least one SEI message for a video unit.
[0304] a. In one example, the indicator is the NN filter model index.
[0305] i. Alternatively, more than one syntax element can be signaled in or derived from an SEI message to represent one NN filter model index.
[0306] 1) In one example, a first syntax element can be signaled in an SEI message to indicate whether the index is not greater than K0 (e.g., K0 = 0).
[0307] a. Alternatively, a second syntax element can be signaled in an SEI message to indicate whether the index is not greater than K0 (e.g., K1 = 1).
[0308] b. Alternatively, a second syntax element can be signaled in an SEI message to indicate whether the index is subtracted by K0.
[0309] b. In one example, the index corresponds to a NN filter model candidate.
[0310] c. In one example, different color components (including luma and chroma) in a video unit can share the same signaled NN filter model index(es).
[0311] i. Alternatively, a filter index is signaled in an SEI message for each color component in a video unit.
[0312] ii. Alternatively, a first filter index is signaled in the SEI message for a first color component, such as luma, and a second filter index is signaled in the SEI message for a second and third color component, such as Cb and Cr, or U and V.
[0313] iii. Alternatively, an indicator (e.g., a flag) is signaled in the SEI message to indicate whether all color components will share the same NN filter model index.
[0314] 1) In one example, when the flag is true, one CNN filter model index is signaled in the SEI message for the video unit. Otherwise, the filter index is signaled in the SEI message according to bullet i or bullet ii above.
[0315] iv. Alternatively, an indicator (e.g., a flag) is signaled in the SEI message to indicate whether two components (e.g., the second and third color components, or Cb and Cr, or U and V) will share the same NN filter model index.
[0316] 1) In one example, when the flag is true, one CNN filter model index is signaled in the SEI message for the two components. Otherwise, a separate filter index will be signaled in the SEI message for each of the two components.
[0317] v. Alternatively, an indicator (e.g., a flag) is signaled in the SEI message to indicate whether a NN filter will be used for the current video unit.
[0318] 1) In one example, if the flag is false, the current video unit will not be processed by a NN filter, meaning no transmission of any NN filter model index. Otherwise, the NN filter model index is signaled in the SEI message according to bullet i, bullet ii, bullet iii, or bullet iv above.
[0319] vi. Alternatively, an indicator (e.g., a flag) is signaled in the SEI message to indicate whether a NN filter will be used for the two components (e.g., the second and third color components, or Cb and Cr, or U and V) in the current video unit.
[0320] 1) In one example, if the flag is false, the two components will not be processed by a NN filter, meaning no transmission of any NN filter model index for the two components. Otherwise, the NN filter model index is signaled in the SEI message according to bullet i, bullet ii, bullet iii, or bullet iv above.
[0321] vii.The indicator can be coded with one or more contexts in arithmetic coding. A context in arithmetic coding can refer to a symbol preceding the current symbol of the code and thus defining the current context.
[0322] 1) In one example, the indicator of the index can be binarized into a string of bins, and at least one bin can be coded with one or more contexts.
[0323] 2) Alternatively, the indicator of the index can be first binarized into a string of bins, and at least one bin can be coded in bypass mode.
[0324] 3) The context can be derived from the parameters or coding information of the current unit and / or neighboring units.
[0325] viii.In the above example, a particular value of the indicator (e.g., equal to 0) can be further used to indicate whether to apply the NN filter or not.
[0326] ix.The indicator can be binarized with a fixed-length code, or a unary code, or a truncated unary code, or an exponential Golomb (EG) code (e.g., the Kth EG code, where K = 0), or a truncated exponential Golomb code. An EG code is a general-purpose code in which any non-negative integer x is coded using an EG code, where a) x + 1 is written in binary, and b) the number of leading zero bits is counted and written before the previous string of bits.
[0327] x.The indicator can be binarized based on the coding information (such as QP) of the current and / or neighboring blocks.
[0328] 1) In one example, let q represent the QP of the current video unit, then K CNN filter models are trained to correspond to q1, q2,..., qK, respectively. k where q1, q2,..., qK are K different QPs. k
[0329] a. Alternatively, shorter code lengths are assigned to the indices corresponding to the models based on q i = q.
[0330] b. Alternatively, shorter code lengths are assigned to the indices corresponding to the models based on q i , where abs(q j -q) is the smallest one compared to other q i .
[0331] 2) In one example, let the QP of the current video unit be denoted as q, then three CNN filter models are trained to correspond to q1, q, q2 respectively, where q1<q<q2. A first flag is then signaled to indicate whether the model corresponding to QP=q is used. If the flag is false, a second flag is signaled to indicate whether the model corresponding to QP=q1 is used. If the second flag is still false, a third flag is further signaled to indicate whether the model corresponding to QP=q2 is used.
[0332] d. In one example, the NN filter model index can be coded in a predictive manner.
[0333] i. For example, the previously coded / decoded NN filter model index can be used as a prediction for the current NN filter model index.
[0334] 1) A flag can be signaled to indicate whether the current NN filter model index is equal to the previously coded / decoded NN filter model index.
[0335] e. In one example, an indicator of one or several NN filter model indices in the current video unit can be inherited from the previously coded / adjacent video unit.
[0336] i. In one example, the NN filter model index of the current video unit can be inherited from the previously coded / adjacent video unit.
[0337] 1) In one example, let the number of previously coded / adjacent video unit candidates be denoted as C. An inheritance index (ranging from 0 to C-1) is then signaled for the current video unit to indicate the candidate to be inherited.
[0338] ii. In one example, the NN filter on / off control of the current video unit can be inherited from the previously coded / adjacent video unit. The on / off control refers to the parameters that can enable or disable the NN filter for the current video unit.
[0339] f. In the above example, the signaled indicator can further indicate whether the NN filtering method is applied.
[0340] i. In one example, assuming M filter models are allowed for a video unit, the indicator can be signaled in the following manner:
[0341]
[0342] 1) Alternatively, the "0" and "1" values of the kth binary bit in the above table can be swapped.
[0343] ii. Alternatively, assuming M filter models are allowed for a video unit, the indicator can be signaled in the following way:
[0344]
[0345] g. In the above example, the same value of the indicator for two video units can represent different filter models applied to the two video units.
[0346] i. In one example, how to select a filter model for a video unit can depend on the decoded value of the indicator and the characteristics of the video unit (e.g., QP / prediction mode / other coding information).
[0347] 3. For encoder and decoder, one or several NN filter model index and / or on / off control can be derived from SEI message for a unit in the same way.
[0348] a. The derivation can depend on the coding information of the current unit and / or neighboring units.
[0349] i. The coding information includes QP.
[0350] 4. An indicator of multiple NN filter model indexes can be signaled in SEI message for one video unit, and the selection of the indexes for sub-regions within one video unit can be further determined on the fly.
[0351] a. In one example, an indicator of a first set of NN filter models can be signaled in SEI message at a first level (e.g., picture / slice), and an indicator of the indexes within the first set can be further signaled in SEI message or derived from SEI message at a second level (where the number of samples at the second level is smaller than that at the first level).
[0352] i. In one example, the second level is CTU / CTB level.
[0353] ii. In one example, the second level is a fixed MxN region.
[0354] 5. A syntax element (e.g., a flag) can be signaled / derived from SEI message to indicate whether to enable or disable NN filter for a video unit.
[0355] a. In one example, the syntax element can be context coded or bypass coded.
[0356] i. In one example, one context is used to code the syntax element.
[0357] ii.In one example, multiple contexts can be utilized, e.g., the selection of the context can depend on the value of a syntax element associated with a neighboring block.
[0358] b.In one example, the syntax element can be signaled in the SEI message at multiple levels (e.g., at picture and slice level; at slice and CTB / CTU level).
[0359] i.Alternatively, whether the syntax element is signaled at the second level (e.g., CTB / CTU) can be based on the value of the syntax element at the first level (e.g., slice).
[0360] 6.An indicator of multiple NN filter model indices can be conditionally signaled in the SEI message.
[0361] a.In one example, whether the indicator is signaled can depend on whether the NN filter is enabled for the video unit.
[0362] i.Alternatively, the signaling of the indicator is skipped when the NN filter is disabled for the video unit.
[0363] 7.A first indicator can be signaled in the SEI message in a parent unit to indicate, for each video unit contained in the parent unit, how the NN filter model indices are signaled in the SEI message or how the NN filter is to be used for each video unit contained in the parent unit.
[0364] a.In one example, the first indicator can be used to indicate whether all samples within the parent unit share the same on / off control.
[0365] i.Alternatively, a second indicator of video units within the parent unit indicating the use of the NN filter can be conditionally signaled based on the first indicator.
[0366] b.In one example, the first indicator can be used to indicate which model index is used for all samples within the parent unit.
[0367] i.Alternatively, if the first indicator indicates the use of a given model index, a second indicator of video units within the parent unit indicating the use of the NN filter can be further signaled.
[0368] ii.Alternatively, if the first indicator indicates the use of a given model index, the second indicator of video units within the parent unit indicating the use of the NN filter can be skipped.
[0369] c.In one example, the first indicator can be used to indicate whether an index of video units within the parent unit is further signaled in the SEI message.
[0370] d. In one example, the indicator can have K+2 options, where K is the number of NN filter model candidates.
[0371] i. In one example, when the indicator is 0, the NN filter is disabled for all video units contained in the parent unit.
[0372] ii. In one example, when the indicator is i (1≤i≤K), the i-th NN filter model is to be used for all video units contained in the parent unit. Obviously, for the K+1 options just mentioned, there is no need to signal any NN filter model index for any video unit contained in the parent unit.
[0373] 1) Alternatively, the information of on / off control of the video unit is still signaled, and when it is on, the i-th NN filter model is used.
[0374] iii. In one example, when the indicator is K+1, the NN filter model index (e.g., filter index, and / or on / off control) is to be signaled for each video unit contained in the parent unit.
[0375] e. In one example, the parent unit can be a slice or a picture.
[0376] i. Alternatively, the video unit within the parent unit can be a CTB / CTU / region within the parent unit.
[0377] 8. For one video unit, multiple (i.e., more than 1) model indices of the NN filter can be signaled in the SEI message or the video bitstream.
[0378] a. In one example, for a sample within the video unit, it can select one NN filter according to one of the multiple model indices.
[0379] i. Alternatively, the selection can depend on the decoded information (e.g., prediction mode).
[0380] b. In one example, for a sample within the video unit, the NN filter according to the multiple model indices can be applied.
[0381] 9. For the above bullets, the indicator and / or syntax elements presented in the SEI message can be coded with a fixed length code.
[0382] a. Alternatively, at least one of them can be coded with an arithmetic coder.
[0383] b. Alternatively, at least one of them can be coded with a k-th order EG code / truncated unary / truncated binary.
[0384] 10. For example, at least one NN parameter can be signaled in at least one SEI message.
[0385] a. For example, the NN model structure can be signaled in an SEI message.
[0386] i. For example, an indication of the NN model structure selected from some candidates can be signaled in an SEI message.
[0387] b. For example, a set of NN model parameters (such as network layers, connections, number of feature maps, kernel size, convolution stride, padding method, etc.) that can represent the model can be signaled in an SEI message.
[0388] i. The parameters can be coded in a predictive manner.
[0389] Example 2
[0390] 11. In a second embodiment, an indication of whether a NN filter is enabled for a video unit or for samples within a video unit can be presented in at least one SEI message.
[0391] a. In one example, for the luma component, a first indicator is signaled.
[0392] i. Alternatively, the first indicator is signaled to indicate that the NN filter is used for all three color components.
[0393] ii. Alternatively, the first indicator is signaled to indicate that the NN filter is used for at least one of the three color components.
[0394] b. In one example, for the two chroma components, a second indicator is signaled.
[0395] i. Alternatively, for each of the two chroma components, a second indicator and a third indicator are signaled.
[0396] Example 3
[0397] Embodiments of applying a NN-based filtering method are discussed. An example SEI message syntax is provided below.
[0398]
[0399]
[0400] luma_parameters_present_flag equal to 1 indicates that luma cnn parameters will be present. luma_parameters_present_flag equal to 0 indicates that luma cnn parameters will not be present.
[0401] cb_parameters_present_flag equal to 1 indicates that there will be cb cnn parameters. cb_parameters_present_flag equal to 0 indicates that there will be no cb cnn parameters.
[0402] cr_parameters_present_flag equal to 1 indicates that there will be cr cnn parameters. cr_parameters_present_flag equal to 0 indicates that there will be no cr cnn parameters.
[0403] luma_cnn_filter_slice_indication equal to 0 indicates that the luma cnn filter will not be used for the luma component of any block in the slice. luma_cnn_filter_slice_indication equal to i (i = 1, 2, or 3) indicates that the i-th luma cnn filter will be used for the luma component of all blocks in the slice. luma_cnn_filter_slice_indication equal to 4 indicates that luma_cnn_block_idc will be parsed for each block in the slice.
[0404] cb_cnn_filter_slice_indication equal to 0 indicates that the cb cnn filter will not be used for the cb component of any block in the slice. cb_cnn_filter_slice_indication equal to i (i = 1, 2, or 3) indicates that the i-th cb cnn filter will be used for the cb component of all blocks in the slice. cb_cnn_filter_slice_indication equal to 4 indicates that cb_cnn_block_idc will be parsed for each block in the slice.
[0405] cr_cnn_filter_slice_indication equal to 0 indicates that the cr cnn filter will not be used for the cr component of any block in the slice. cr_cnn_filter_slice_indication equal to i (i = 1, 2, or 3) indicates that the i-th cr cnn filter will be used for the cr component of all blocks in the slice. cr_cnn_filter_slice_indication equal to 4 indicates that cr_cnn_block_idc will be parsed for each block in the slice.
[0406] luma_cnn_block_idc[ i ][ j ] equal to 0 indicates that no luma cnn filter will be used for the luma component of this block. luma_cnn_block_idc[ i ][ j ] equal to i (i = 1, 2 or 3) indicates that the i-th luma cnn filter will be used for the luma component of this block.
[0407] cb_cnn_block_idc[ i ][ j ] equal to 0 indicates that no cb cnn filter will be used for the cb component of this block. cr_cnn_block_idc[ i ][ j ] equal to i (i = 1, 2 or 3) indicates that the i-th cb cnn filter will be used for the cb component of this block.
[0408] cr_cnn_block_idc[ i ][ j ] equal to 0 indicates that no cr cnn filter will be used for the cr component of this block. cr_cnn_block_idc[ i ][ j ] equal to i (i = 1, 2 or 3) indicates that the i-th cr cnn filter will be used for the cr component of this block.
[0409] Alternatively, the u(N) coding method of one or more syntax elements in the SEI message can be replaced by ae(v) or ue(v) or f(n) or se(v).
[0410] Figure 14 is a block diagram illustrating an example video processing system 1400 in which various techniques disclosed herein can be implemented. Various implementations can include some or all of the components of the video processing system 1400. The video processing system 1400 can include an input 1402 for receiving video content. The video content can be received in a raw or uncompressed format, e.g., 8-bit or 10-bit multi-component pixel values, or can be received in a compressed or encoded format. The input 1402 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical networks (PONs), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0411] The video processing system 1400 can include a codec component 1404 that can implement various coding or encoding methods described in this document. The codec component 1404 can reduce the average bitrate of a video from an input 1402 to an output of the codec component 1404 to produce a coded representation of the video. Thus, the coding techniques are sometimes called video compression or video transcoding techniques. The output of the codec component 1404 can be stored or transmitted over a communication (as shown by component 1406) that is connected. The stored or transmitted bitstream (or coded) representation of the video received at the input 1402 can be used by a component 1408 to generate pixel values or a displayable video to a display interface 1410. The process of generating a user-viewable video from a bitstream representation is sometimes called video decompression. Also, although certain video processing operations are referred to as “coding” operations or tools, it should be understood that the coding tools or operations are used at an encoder, and corresponding decoding tools or operations that reverse the results of the coding will be performed by a decoder.
[0412] Examples of peripheral bus interfaces or display interfaces can include a Universal Serial Bus (USB) or a High Definition Multimedia Interface (HDMI) or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described herein can be implemented in various electronic devices such as mobile phones, laptops, smart phones, or other devices capable of performing digital data processing and / or video display.
[0413] Figure 15 is a block diagram of a video processing apparatus 1500. The apparatus 1500 can be used to implement one or more methods described herein. The apparatus 1500 can be implemented in a smart phone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 1500 can include one or more processors 1502, one or more memories 1504, and video processing hardware 1506. The processor(s) 1502 can be configured to implement one or more methods described in the present document. The memory (memories) 1504 can be used for storing data and code used for implementing the methods and techniques described herein. The video processing hardware 1506 can be used to implement, using hardware circuitry, some of the techniques described in the present document. In some embodiments, the hardware 1506 can be implemented completely or partially in the processor 1502, e.g., as a graphics co-processor.
[0414] Figure 16 is a block diagram that describes an example video coding system 1600 that can utilize the techniques of this disclosure. As Figure 16As shown, video coding system 1600 can include a source device 1610 and a destination device 1620. Source device 1610 generates encoded video data, which can be referred to as a video encoding device. Destination device 1620 can decode the encoded video data generated by source device 1610, which can be referred to as a video decoding device.
[0415] Source device 1610 can include a video source 1612, a video encoder 1614, and an input / output (VO) interface 1616.
[0416] Video source 1612 can include a source such as a video capture device, an interface to receive video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. Video data can comprise one or more pictures. Video encoder 1614 encodes video data from video source 1612 to generate a bitstream. The bitstream can include a sequence of bits that form a coded representation of the video data. The bitstream can include coded pictures and associated data. A coded picture is a coded representation of a picture. Associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. VO interface 1616 can include a modulator / demodulator (modem) and / or a transmitter. Encoded video data can be transmitted directly to destination device 1620 via VO interface 1616 and network 1630. The encoded video data can also be stored onto a storage medium / server 1640 for access by destination device 1620.
[0417] Destination device 1620 can include VO interface 1626, video decoder 1624, and display device 1622.
[0418] VO interface 1626 can include a receiver and / or a modem.
[0419] VO interface 1626 can obtain encoded video data from source device 1610 or storage medium / server 1640. Video decoder 1624 can decode encoded video data. Display device 1622 can display the decoded video data to a user. Display device 1622 can be integrated with destination device 1620, or can be external to destination device 1620, which can be configured to interface with an external display device.
[0420] Video encoder 1614 and video decoder 1624 can operate according to a video compression standard, such as High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard, and other current and / or further standards.
[0421] Figure 17 is a block diagram illustrating an example of a video encoder 1700 that can be Figure 16Video encoder 1614 in video coding system 1600 as illustrated in FIG. 16.
[0422] Video encoder 1700 can be configured to perform any or all of the techniques of this disclosure. In Figure 17 In an example, video encoder 1700 includes a plurality of functional components. The techniques described in this disclosure can be shared among the various components of video encoder 1700. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0423] The functional components of video encoder 1700 can include partition unit 1701, prediction unit 1702, which can include mode select unit 1703, motion estimation unit 1704, motion compensation unit 1705, and intra-prediction unit 1706, residual generation unit 1707, transform unit 1708, quantization unit 1709, inverse quantization unit 1710, inverse transform unit 1711, reconstruction unit 1712, buffer 1713, and entropy encoding unit 1714.
[0424] In other examples, video encoder 1700 can include more, less, or different functional components. In one example, prediction unit 1702 can include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0425] Furthermore, some components, such as motion estimation unit 1704 and motion compensation unit 1705, can be highly integrated, but are represented separately for illustrative purposes. Figure 17
[0426] Partition unit 1701 can partition a picture into one or more video blocks. Figure 16 Video encoder 1614 and video decoder 1624 in FIG. 16 can support various video
[0427] Mode select unit 1703 can select one of the coding modes (intra or inter, for example, based on error results), and provide the resulting intra or inter coded block to residual generation unit 1707 to generate residual block data, and to reconstruction unit 1712 to reconstruct the coded block for use as a reference picture. In some examples, mode select unit 1703 can select a combination of intra and inter prediction (CIIP) mode, where the prediction is based on both an inter prediction signal and an intra prediction signal. In the case of inter prediction, mode select unit 1703 can also select a resolution for the motion vectors (e.g., sub-pixel or integer pixel precision).
[0428] To perform inter prediction on a current video block, motion estimation unit 1704 can generate motion information for the current video block by comparing one or more reference frames from buffer 1713 to the current video block. Motion compensation unit 1705 can determine a predicted video block for the current video block based on the motion information and decoded samples for pictures from buffer 1713 other than the picture associated with the current video block.
[0429] Motion estimation unit 1704 and motion compensation unit 1705 can perform different operations on a current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0430] In some examples, motion estimation unit 1704 can perform uni-prediction for a current video block, and motion estimation unit 1704 can search reference pictures of list 0 or list 1 for a reference video block for the current video block. Motion estimation unit 1704 can then generate a reference index that indicates a reference picture of list 0 or list 1 that contains the reference video block and a motion vector that indicates a spatial displacement between the current video block and the reference video block. Motion estimation unit 1704 can output the reference index, the prediction direction indicator, and the motion vector as motion information for the current video block. Motion compensation unit 1705 can generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0431] In other examples, motion estimation unit 1704 can perform bi-prediction for a current video block, and motion estimation unit 1704 can search reference pictures of list 0 for a reference video block for the current video block and also search reference pictures of list 1 for another reference video block for the current video block. Motion estimation unit 1704 can then generate reference indices that indicate reference pictures of list 0 and list 1 that contain the reference video block and the other reference video block and motion vectors that indicate spatial displacements between the reference video blocks and the current video block. Motion estimation unit 1704 can output the reference indices and the motion vectors as motion information for the current video block. Motion compensation unit 1705 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information for the current video block.
[0432] In some examples, motion estimation unit 1704 can output a full set of motion information for a current video block for decoding processing at a decoder.
[0433] In some examples, motion estimation unit 1704 can not output a full set of motion information for a current video block. Instead, motion estimation unit 1704 can signal motion information for the current video block with reference to motion information for another video block. For example, motion estimation unit 1704 can determine that the motion information for the current video block is sufficiently similar to the motion information for a neighboring video block.
[0434] In one example, the motion estimation unit 1704 can indicate a value in a syntax structure associated with the current video block that indicates to the video decoder 1624 that the current video block has the same motion information as another video block.
[0435] In another example, the motion estimation unit 1704 can identify, in a syntax structure associated with the current video block, another video block and a motion vector difference (MVD). The motion vector difference represents a difference between a motion vector of the current video block and a motion vector of the indicated video block. The video decoder 1624 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0436] As discussed above, the video encoder 1614 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 1614 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0437] The intra prediction unit 1706 can perform intra prediction on the current video block. When the intra prediction unit 1706 performs intra prediction on the current video block, the intra prediction unit 1706 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0438] The residual generation unit 1707 can generate residual data for the current video block by subtracting (e.g., indicated by a negative sign) the predicted video block for the current video block from the current video block. The residual data for the current video block can include a residual video block that corresponds to different sample components of samples in the current video block.
[0439] In other examples, the residual data for the current video block can not exist for the current video block, such as in a skip mode, and the residual generation unit 1707 can not perform the subtraction operation.
[0440] The transform processing unit 1708 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0441] After the transform processing unit 1708 generates the transform coefficient video block associated with the current video block, the quantization unit 1709 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0442] The inverse quantization unit 1710 and the inverse transform unit 1711 can apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 1712 can add the reconstructed residual video block to corresponding samples of one or more prediction video blocks generated from the prediction unit 1702 to produce a reconstructed video block associated with the current block for storage in the buffer 1713.
[0443] After the reconstruction unit 1712 reconstructs the video block, loop filtering operations can be performed to reduce video block artifacts in the video block.
[0444] The entropy encoding unit 1714 can receive data from other functional components of the video encoder 1700. When the entropy encoding unit 1714 receives data, the entropy encoding unit 1714 can perform one or more entropy encoding operations to produce entropy encoded data and output a bitstream that includes the entropy encoded data.
[0445] Figure 18 is an example block diagram illustrating a video decoder 1800 that can be Figure 16 the video decoder 1624 in the video coding system 1600 illustrated in FIG.
[0446] The video decoder 1800 can be configured to perform any or all of the techniques of this disclosure. In Figure 18 In examples, the video decoder 1800 includes a number of functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 1800. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0447] In Figure 18 In examples, the video decoder 1800 includes an entropy decoding unit 1801, a motion compensation unit 1802, an intra prediction unit 1803, a dequantization unit 1804, an inverse transform unit 1805, and a reconstruction unit 1806 and a buffer 1807. In some examples, the video decoder 1800 can perform a decoding pass generally reciprocal to the encoding pass described with respect to the video encoder 1614 (e.g., Figure 16 ).
[0448] The entropy decoding unit 1801 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded video data blocks). The entropy decoding unit 1801 can decode the entropy encoded video data, and from the entropy decoded video data, the motion compensation unit 1802 can determine motion information including motion vectors, motion vector precisions, reference picture list indices, and other motion information. For example, the motion compensation unit 1802 can determine this information by performing AMVP and merge modes.
[0449] Motion compensation unit 1802 can generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of the interpolation filter used at sub-pixel precision can be included in the syntax elements.
[0450] Motion compensation unit 1802 can calculate interpolated values for sub-integer pixels of the reference block using an interpolation filter used by video encoder 1614 during encoding of the video block. Motion compensation unit 1802 can determine the interpolation filter used by video encoder 1614 from the received syntax information and use the interpolation filter to produce the prediction block.
[0451] Motion compensation unit 1802 can use some of the syntax information to determine the size of blocks used to encode frames and / or slices of the encoded video sequence, partition information describing how each macroblock of a picture of the encoded video sequence is partitioned, modes indicating how each partition is encoded, one or more reference frames (and lists of reference frames) for each inter-coded block, and other information to decode the encoded video sequence.
[0452] Intra prediction unit 1803 can use, for example, intra prediction modes received in the bitstream to form the prediction block from spatially neighboring blocks. Dequantization unit 1804 dequantizes, i.e., inverse quantizes, quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 1801. Inverse transform unit 1805 applies an inverse transform.
[0453] Reconstruction unit 1806 can add the residual block to the corresponding prediction block generated by motion compensation unit 1802 or intra prediction unit 1803 to form a decoded block. If desired, a deblocking filter can also be applied to filter the decoded block to remove blocking artifacts. The decoded video blocks are then stored in buffer 1807, which provides reference blocks for subsequent motion compensation / intra prediction, and also generates decoded video for presentation on a display device.
[0454] Figure 19 Method 1900 of processing video data according to embodiments of the disclosure. Method 1900 can be performed by a codec device (e.g., an encoder) having a processor and a memory. In block 1902, the codec device determines that a supplemental enhancement information (SEI) message of a bitstream includes an indicator of one or more neural network (NN) filter model candidates or selections that specify a sample within a video unit or a video unit.
[0455] In block 1904, the codec converts between a video media file comprising video units and a bitstream based on the indicator. When implemented in an encoder, the conversion includes receiving a media file (e.g., video units), and encoding the indicator into a bitstream. When implemented in a decoder, the conversion includes receiving a bitstream comprising the indicator, and decoding the bitstream to obtain a video media file. In embodiments, the method 1900 can utilize or incorporate one or more of the features or processes of other methods disclosed herein.
[0456] In embodiments, an apparatus for coding video data comprises a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to determine that a supplemental enhancement information (SEI) message of a bitstream includes an indicator that specifies one or more neural network (NN) filter model candidates or selections for a video unit or samples within the video unit, and convert between a video media file comprising the video unit and the bitstream based on the indicator.
[0457] In embodiments, a non-transitory computer-readable medium comprising a computer program product for use by a codec comprises computer executable instructions stored thereon which, when executed by one or more processors, cause the codec to determine that a supplemental enhancement information (SEI) message of a bitstream includes an indicator that specifies one or more neural network (NN) filter model candidates or selections for a video unit or samples within the video unit, and convert between a video media file comprising the video unit and the bitstream based on the indicator.
[0458] Figure 20 is another method 1920 of processing video data according to embodiments of the disclosure. The method 1920 can be performed by a codec (e.g., an encoder) having a processor and a memory. In block 1922, the codec determines that a supplemental enhancement information (SEI) message of a bitstream includes an indicator that specifies one or more neural network (NN) filter model indices.
[0459] In block 1924, the codec converts between a video media file and the bitstream based on the indicator. When implemented in an encoder, the conversion includes receiving a media file (e.g., video units), and encoding the indicator into a bitstream. When implemented in a decoder, the conversion includes receiving a bitstream comprising the indicator, and decoding the bitstream to obtain a video media file. In embodiments, the method 1920 can utilize or incorporate one or more of the features or processes of other methods disclosed herein.
[0460] In an embodiment, an apparatus for coding video data includes a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to determine that a supplemental enhancement information (SEI) message of a bitstream includes an indicator that specifies one or more neural network (NN) filter model indexes, and convert between a video media file and the bitstream based on the indicator.
[0461] In an embodiment, a non-transitory computer-readable medium including a computer program product for use by a coding apparatus includes computer executable instructions stored thereon which, when executed by one or more processors, cause the coding apparatus to determine that a supplemental enhancement information (SEI) message of a bitstream includes an indicator that specifies one or more neural network (NN) filter model indexes, and convert between a video media file and the bitstream based on the indicator.
[0462] Figure 21 Another method 1940 of processing video data according to embodiments of the disclosure is described. The method 1940 can be performed by a coding apparatus (e.g., an encoder) having a processor and a memory. In block 1942, the coding apparatus determines that a supplemental enhancement information (SEI) message of a bitstream includes a syntax element that indicates whether one or more neural network (NN) filter models are enabled or disabled for a video unit.
[0463] In block 1944, the coding apparatus converts between a video media file and the bitstream based on the syntax element. When implemented in an encoder, the conversion includes receiving a media file (including a plurality of video layers), and encoding the indicator to the bitstream. When implemented in a decoder, the conversion includes receiving the bitstream including the indicator, and decoding the bitstream to obtain the video media file. In an embodiment, the method 1940 can utilize or incorporate one or more of the features or processes of other methods disclosed herein.
[0464] In an embodiment, an apparatus for coding video data includes a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to determine that a supplemental enhancement information (SEI) message of a bitstream includes an indicator that specifies one or more neural network (NN) filter model indexes, and convert between a video media file and the bitstream based on the indicator.
[0465] In an embodiment, a non-transitory computer-readable medium comprising a computer program product for use by a coding device, the computer program product comprising computer executable instructions stored on the non-transitory computer readable medium such that when executed by one or more processors cause the coding device to: determine that a supplemental enhancement information (SEI) message of a bitstream includes a syntax element indicating whether one or more neural network (NN) filter models are enabled or disabled for a video unit, and convert between a video media file comprising the video unit and the bitstream based on the syntax element.
[0466] Next, a list of some embodiments preferred solutions is provided.
[0467] The following solutions show example embodiments of the techniques discussed in the previous section.
[0468] 1. A method of video processing, comprising performing a conversion between a video comprising a video unit and a bitstream of the video according to a rule, wherein the rule specifies that a supplemental enhancement information (SEI) message indicates one or more neural network (NN) filter model candidates or a NN filter model that is applicable to in-loop filtering of the video unit or samples within the video unit.
[0469] 2. The method of claim 1, wherein the conversion comprises selecting a NN filter model from the one or more NN filter model candidates.
[0470] 3. The method of claim 2, wherein the selecting is performed at a video unit level.
[0471] 4. The method of claim 1, wherein the rule specifies that a luma component and a chroma component of the video unit are converted using different NN filter models.
[0472] 5. The method of claim 1 or 4, wherein the SEI message indicates L NN filter model candidates for luma and C filter model candidates for chroma, wherein L is not equal to C
[0473] 6. The method of any of claims 1-5, wherein the video unit belongs to a video region, and wherein different SEI messages are associated with different video regions of the video.
[0474] 7. The method of claim 1, wherein the rule depends on a tree partitioning structure of the video unit.
[0475] 8. A method of video processing, comprising performing a conversion between a video comprising a video unit and a bitstream of the video according to a rule, wherein the rule specifies that a supplemental enhancement information (SEI) message indicates an index of one or more neural network filter models of loop filtering that are applicable to the video unit or samples within the video unit.
[0476] 9. The method of claim 8, wherein the SEI message comprises an indicator of the index.
[0477] 10. The method of claim 9, wherein the index is shared between luma and chroma components of the video unit.
[0478] 11. The method of claim 9, wherein the index comprises an index of a luma component separate from an index of one or more chroma components.
[0479] 12. The method of any of claims 8-11, wherein the index is context coded in the bitstream.
[0480] 13. The method of any of claims 8-12, wherein a particular value of the indicator indicates whether a NN filter is enabled for the conversion.
[0481] 14. The method of any of claims 8-13, wherein the index is binarized based on coding information of the video unit or a neighboring video unit.
[0482] 15. The method of any of claims 8-14, wherein the index is prediction coded or inherited.
[0483] 16. The method of any of claims 1-15, wherein the rule specifies that the index is indicated in response to whether a NN filter is enabled.
[0484] 17. The method of any of claims 1-16, wherein the predetermined rule allows derivation of a NN filter model index or determination of whether loop filtering is enabled for the video unit via an on / off control.
[0485] 18. The method of claim 17, wherein the derivation depends on coding information of the video unit.
[0486] 19. The method of any of claims 1-18, wherein the conversion further comprises determining, for a sub-region within the video unit, an index indicated from a SEI message of loop filtering for the sub-region.
[0487] 20. The method of claim 19, wherein the video unit is a video picture, and wherein the sub-region is a coding tree unit or a coding tree block or an MxN region of the video picture, wherein M and N are integers.
[0488] 21. The method of any of claims 1-20, wherein the information about whether a particular NN filter is enabled is indicated by a SEI message.
[0489] 22. The method of claim 21, wherein the information is explicitly indicated.
[0490] 23. The method of claim 21, wherein the information is implicitly indicated.
[0491] 24. The method of claim 23, wherein the information is indicated by a syntax element that is context coded or bypass coded.
[0492] 25. A method of video processing, comprising performing a conversion between a video comprising a video region including a video unit and a bitstream of the video according to a rule, wherein the rule specifies that a supplemental enhancement information (SEI) message of the video region indicates one or more neural network filter model indices of a loop filter that is applicable to the video unit and other video units in the video region.
[0493] 26. The method of claim 25, wherein the SEI message includes a first indicator indicating whether all samples in the video region share a same on / off filter control.
[0494] 27. The method of claim 26, wherein the rule specifies to conditionally include a second indicator in the SEI message to indicate that a NN filter is used in the video unit.
[0495] 28. The method of any of claims 26-27, wherein the first indicator indicates a model index that is applicable to all samples in the video region.
[0496] 29. The method of claim 28, wherein the first indicator takes one of K+2 possible values, wherein K is a number of NN filter models for the video region.
[0497] 30. The method of any of the above claims, wherein the video region comprises a video slice.
[0498] 31. A method of video processing, comprising performing a conversion between a video comprising a video unit and a bitstream of the video according to a rule, wherein the rule specifies whether or how a plurality of syntax element indications or model indices of a neural network (NN) filter model of a loop for the video unit are indicated in a supplemental enhancement information (SEI) message of the video unit or in the bitstream.
[0499] 32. The method of claim 31, wherein the predetermined rule is to determine the NN filter model from a NN filter model applicable to video samples of the video unit.
[0500] 33. The method of any of claims 1-32, wherein the SEI message is coded with a fixed length coding.
[0501] 34. The method of any of claims 1-32, wherein the SEI message is coded using arithmetic coding.
[0502] 35. The method of any of claims 1-34, wherein the rule further specifies that the SEI message comprises one or more parameters of the neural network model.
[0503] 36. The method of claim 35, wherein the one or more parameters are predicted based on previous usage to code.
[0504] 37. The method of any of the preceding claims, wherein the video unit comprises a coding unit, a transform unit, a prediction unit, a slice, or a subpicture.
[0505] 38. The method of any of the preceding claims, wherein the video unit is a coding block or a video slice or a video picture or a video tile or a video subpicture.
[0506] 39. The method of any of claims 1-21, wherein the conversion comprises generating the video from the bitstream or generating the bitstream from the video.
[0507] 40. A video decoding apparatus comprising a processor configured to implement a method recited in one or more of claims 1-22.
[0508] 41. A video encoding apparatus comprising a processor configured to implement a method recited in one or more of claims 1-22.
[0509] 42. A computer program product having computer code stored therein, the code, which when executed by a processor, causes the processor to implement a method recited in any of claims 1-22.
[0510] 43. A method, apparatus or system described in this document.
[0511] The disclosed and other solutions, examples, embodiments, modules and functional operations set forth in this document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structural equivalents of such as disclosed in this document, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine- readable propagated signal, or a combination of one or more of them. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine- readable propagated signal, or a combination of one or more of them. The term“data processing apparatus” encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. The propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus.
[0512] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and are interconnected by a communication network.
[0513] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).
[0514] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0515] Although the patent document includes many details, these should not be construed as limiting the scope of any invention or of the claimed scope but as a description of features that can be specific to particular embodiments of the invention. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any suitable subcombination. Furthermore, although features can be described above as acting in certain combinations and initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can then be directed to a subcombination or variation of a subcombination.
[0516] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such order nor that all illustrated operations be performed, to accomplish desirable results. Additionally, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0517] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.
[0518] Although this patent document contains many details, these should not be construed as limiting the scope of any invention or of the scope of patentable subject matter, but as merely describing a particular implementation of specific embodiments. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any suitable subcombination. Moreover, although features can be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can be directed to a subcombination or variation of a subcombination.
[0519] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such order nor that all illustrated operations be performed, to accomplish desirable results. Additionally, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0520] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A method for processing video data, comprising: The Supplemental Enhancement Information (SEI) message for determining the bitstream includes indicators that specify one or more neural network (NN) filter model candidates or selections for a video unit or samples within a video unit. as well as Converting between video media files and bitstreams, including video units, based on indicators. The indicator is included in at least one of multiple levels of the SEI message, and the multiple levels include picture level, stripe level, codec tree block (CTB) level, and codec tree unit (CTU) level. This also includes deriving the index of one or more NN filter models or deriving the on / off control of one or more NN filter models based on SEI messages. The derivation depends on the parameters of the current video unit or its neighboring video units. Among these parameters are the quantization parameters QP associated with the video unit or neighboring video units.
2. The method according to claim 1, wherein, Whether an indicator is included in the SEI message depends on whether the NN filter is enabled for the video unit.
3. The method according to claim 2, wherein, When the NN filter is disabled for a video unit, the indicator is excluded from the SEI message.
4. The method according to claim 1, wherein, Use arithmetic encoding / decoding or exponential Columbus EG encoding / decoding to encode the indicator into a bit stream.
5. The method according to claim 1, wherein, The indicator in the SEI message further specifies one or more NN filter model indices.
6. The method of claim 1, further comprising determining that the SEI message of the bitstream includes an indicator specifying one or more NN filter model indices.
7. The method according to claim 6, wherein, Use fixed-length encoding / decoding to encode / decode an indicator that specifies one or more NN filter model indices into a bitstream.
8. The method according to claim 6, wherein, The SEI message includes a first indicator in the parent unit of the bitstream, and the first indicator indicates how one or more NN filter model indices are included in the SEI message for each video unit included in the parent unit, or how NN filters are used for each video unit included in the parent unit.
9. The method according to claim 1, further comprising: The SEI message that determines the bitstream includes syntax elements indicating whether one or more neural network (NN) filter models are enabled or disabled for the video unit.
10. The method according to claim 9, wherein, Contextual encoding or bypass encoding / decoding of syntax elements in the bitstream.
11. The method according to claim 10, wherein, Arithmetic encoding and decoding of syntax elements can be performed using a single context or multiple contexts.
12. The method according to claim 1, wherein, Multiple levels include a first level and a second level, and whether or not a syntactic element of the second level is included depends on the value of the syntactic element of the first level.
13. The method according to claim 1, wherein, The conversion includes encoding the video media file into a bitstream.
14. The method according to claim 1, wherein, The conversion includes decoding the bitstream to obtain a video media file.
15. The method of claim 13, further comprising: The bit stream is stored in a non-transitory computer-readable recording medium.
16. An apparatus for encoding and decoding video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: The Supplemental Enhancement Information (SEI) message, which determines the bitstream, includes indicators specifying one or more neural network (NN) filter model candidates or selections for a video unit or samples within a video unit; and Converting between video media files and bitstreams, including video units, based on indicators. in, The indicator is included in at least one of multiple levels of the SEI message, and the multiple levels include the picture level, stripe level, codec tree block (CTB) level, and codec tree unit (CTU) level. This also includes deriving the index of one or more NN filter models or deriving the on / off control of one or more NN filter models based on SEI messages. The derivation depends on the parameters of the current video unit or its neighboring video units. Among these parameters are the quantization parameters QP associated with the video unit or neighboring video units.
Citation Information
Patent Citations
Supplemental enhancement information messages for neural network based video post processing
US20200304836A1