Systems and methods for reducing reconstruction error in video coding
By using cross-component filtering technology to reduce reconstruction errors in video coding, the problem of reconstruction errors affecting video quality and coding efficiency in existing technologies is solved, thus achieving higher quality and more efficient video coding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHARP KK
- Filing Date
- 2020-06-23
- Publication Date
- 2026-07-07
AI Technical Summary
Existing video coding technologies suffer from reconstruction errors during the reconstruction process, which affect video quality and coding efficiency.
Cross-component filtering technology is adopted. Before the adaptive loop filtering process, the cross-component filter coefficients and the reconstructed luminance component sample values are used to derive the filtered sample values, and the refined chrominance sample values are derived by summing the chrominance component sample values and the refined values, so as to reduce the reconstruction error.
It effectively reduces reconstruction errors in video encoding, improving video quality and encoding efficiency.
Smart Images

Figure CN122349010A_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese patent application "System and method for reducing reconstruction error in video coding based on cross-component correlation" (application number: 202080045948.2), filed on June 23, 2020. Technical Field
[0002] This disclosure relates to video coding, and more specifically to techniques for reducing reconstruction errors. Background Technology
[0003] Digital video functionality can be integrated into a wide variety of devices, including digital televisions, laptops or desktops, tablets, digital recording devices, digital media players, video game consoles, cellular phones (including so-called smartphones), medical imaging equipment, and more. Digital video can be encoded according to video coding standards. Video coding standards define the format for encapsulating encoded video data into compliant bitstreams. A compliant bitstream is a data structure that can be received and decoded by video decoding devices to generate reconstructed video data. Video coding standards can be combined with video compression techniques. Examples of video coding standards include ISO / IEC MPEG-4 Visual and ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) and High Efficiency Video Coding (HEVC). HEVC is described in the ITU-T H.265 Recommendation of December 2016, which is incorporated herein by reference and referred to herein as ITU-T H.265. Extensions and improvements to ITU-T H.265 are currently under consideration for developing next-generation video coding standards. For example, the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) (collectively referred to as the Joint Video Study Group (JVET)) are working to standardize video coding technologies with compression capabilities significantly exceeding the current HEVC standard. The Joint Exploratory Model 7 (JEM7), the algorithm description of Joint Exploratory Test Model 7 (JEM 7), and the ISO / IEC JTC1 / SC29 / WG11 document: JVET-G1001 (July 2017, Turin, Italy), which are incorporated herein by reference, describe the coding features of the JVET under the Joint Test Model Study, a technology that represents a potential enhancement to video coding beyond the capabilities of ITU-T H.265. It should be noted that the coding features of JEM 7 are implemented in the JEM reference software. As used herein, the term JEM can refer collectively to the algorithms included in JEM 7 and the specific implementations in the JEM reference software. In addition, in response to the “Joint Call for Proposals on Video Compression with Capabilities beyond HEVC” jointly released by VCEG and MPEG, various groups presented multiple descriptions of video coding tools at the 10th meeting of ISO / IEC JTC1 / SC29 / WG11 held in San Diego, CA, from April 16 to 20, 2018.Based on various descriptions of video coding tools, the final initial draft text of the video coding specification was described in "Versatile Video Coding (Draft 1)," also known as document JVET-J1001-v2, presented at the 10th meeting of ISO / IEC JTC1 / SC29 / WG11 held in San Diego, California, April 16-20, 2018. This document is incorporated herein by reference and referred to as JVET-J1001. The current development of the next-generation video coding standard for JVET and MPEG is known as the Universal Video Coding (VVC) project. "Versatile Video Coding (Draft 5)" (document JVET-N1001-v8, which is incorporated herein by reference and referred to as JVET-N1001), presented at the 14th meeting of ISO / IEC JTC1 / SC29 / WG11 held in Geneva, Switzerland, March 19-27, 2019, represents a new version of the draft text of the video coding specification corresponding to the VVC project. The “Versatile Video Coding (Draft 6)” (document JVET-O2001-vE, which is incorporated herein by reference and referred to as JVET-O2001) from the 15th meeting of ISO / IEC JTC1 / SC29 / WG11 held in Gothenburg, Sweden from July 3 to 12, 2019, refers to the current version of the draft text of the video coding specification corresponding to the VVC project.
[0004] Video compression techniques reduce the data requirements for storing and transmitting video data. Video compression can reduce data requirements by utilizing the inherent redundancy in video sequences. It can further divide a video sequence into smaller, consecutive parts (i.e., a set of images within a video sequence, images within a set of images, regions within images, sub-regions within regions, etc.). Intra-frame predictive coding techniques (e.g., spatial prediction within images) and inter-frame prediction techniques (i.e., temporal techniques between images) can be used to generate the difference between the unit of video data to be encoded and a reference unit of the video data. This difference can be called residual data. The residual data can be encoded as quantized transform coefficients. Syntax elements can relate to the residual data and the reference coding unit (e.g., intra-frame predictive mode index and motion information). Entropy coding can be applied to the residual data and syntax elements. The entropy-coded residual data and syntax elements can be included in a data structure that forms a compatible bitstream. Summary of the Invention
[0005] In one example, a method for filtering reconstructed video data is provided, the method comprising: inputting reconstructed luma component sample values; deriving filtered sample values by using cross-component filter coefficients and reconstructed luma component sample values prior to an adaptive loop filtering process; deriving refined values for chroma components by using the filtered sample values; and deriving refined chroma sample values by using the sample values of the chroma components and the sum of the refined values for the chroma components.
[0006] In one example, a decoder for decoding encoded data is provided, the decoder including: a processor and memory associated with the processor; wherein the processor is configured to perform the following steps: inputting reconstructed luminance component sample values; deriving filtered sample values by using cross-component filter coefficients and the reconstructed luminance component sample values prior to an adaptive loop filtering process; deriving refined values for chrominance components using the filtered sample values; and deriving refined chrominance sample values by using the sample values of the chrominance components and the sum of the refined values for the chrominance components.
[0007] In one example, an encoder for encoding video data is provided, the encoder including: a processor and memory associated with the processor; wherein the processor is configured to perform the following steps: inputting reconstructed luma component sample values; deriving filtered sample values by using cross-component filter coefficients and the reconstructed luma component sample values prior to an adaptive loop filtering process; deriving refined values for the chroma component using the filtered sample values; and deriving refined chroma sample values by using the sample values of the chroma component and the sum of the refined values for the chroma component. Attached Figure Description
[0008] [ Figure 1 ] Figure 1 This is a conceptual diagram illustrating an example of a set of pictures encoded according to quadtree / multitree partitioning based on one or more techniques of this disclosure.
[0009] [ Figure 2A ] Figure 2A This is a conceptual diagram illustrating an example of encoding video data blocks according to one or more techniques disclosed herein.
[0010] [ Figure 2B ] Figure 2B This is a conceptual diagram illustrating an example of encoding video data blocks according to one or more techniques disclosed herein.
[0011] [ Figure 3 ] Figure 3 This is a conceptual diagram illustrating examples of video component sampling formats usable according to one or more techniques of this disclosure.
[0012] [ Figure 4A ] Figure 4A This is a conceptual diagram illustrating examples of location types of video component sampling formats that can be used according to one or more techniques of this disclosure.
[0013] [ Figure 4B ] Figure 4B This is a conceptual diagram illustrating examples of location types of video component sampling formats that can be used according to one or more techniques of this disclosure.
[0014] [ Figure 4C ] Figure 4C This is a conceptual diagram illustrating examples of location types of video component sampling formats that can be used according to one or more techniques of this disclosure.
[0015] [ Figure 4D ] Figure 4D This is a conceptual diagram illustrating examples of location types of video component sampling formats that can be used according to one or more techniques of this disclosure.
[0016] [ Figure 4E ] Figure 4E This is a conceptual diagram illustrating examples of location types of video component sampling formats that can be used according to one or more techniques of this disclosure.
[0017] [ Figure 4F ] Figure 4F This is a conceptual diagram illustrating examples of location types of video component sampling formats that can be used according to one or more techniques of this disclosure.
[0018] [ Figure 5 ] Figure 5 This is a block diagram illustrating an example of a system that can be configured to encode and decode video data according to one or more techniques of this disclosure.
[0019] [ Figure 6 ] Figure 6 This is a block diagram illustrating an example of a video encoder that can be configured to encode video data according to one or more techniques of this disclosure.
[0020] [ Figure 7 ] Figure 7 This is a block diagram illustrating an example of a cross-component filter unit that can be configured to encode video data according to one or more techniques of this disclosure.
[0021] [ Figure 8 ] Figure 8 This is a conceptual diagram illustrating an example of reconstruction error for multiple components of video data according to one or more techniques of this disclosure.
[0022] [ Figure 9A ] Figure 9AThis is a conceptual diagram illustrating a sample example of a support for cross-component filtering based on one or more techniques according to this disclosure.
[0023] [ Figure 9B ] Figure 9B This is a conceptual diagram illustrating a sample example of a support for cross-component filtering based on one or more techniques according to this disclosure.
[0024] [ Figure 9C ] Figure 9C This is a conceptual diagram illustrating a sample example of a support for cross-component filtering based on one or more techniques according to this disclosure.
[0025] [ Figure 9D ] Figure 9D This is a conceptual diagram illustrating a sample example of a support for cross-component filtering based on one or more techniques according to this disclosure.
[0026] [ Figure 9E ] Figure 9E This is a conceptual diagram illustrating a sample example of a support for cross-component filtering based on one or more techniques according to this disclosure.
[0027] [ Figure 9F ] Figure 9F This is a conceptual diagram illustrating a sample example of a support for cross-component filtering based on one or more techniques according to this disclosure.
[0028] [ Figure 10 ] Figure 10 This is a conceptual diagram illustrating an example of using cross-component filtering to reduce reconstruction error according to one or more techniques of this disclosure.
[0029] [ Figure 11A ] Figure 11A This is a block diagram illustrating an example of a cross-component filter unit that can be configured to reduce reconstruction error according to one or more techniques of this disclosure.
[0030] [ Figure 11B ] Figure 11B This is a block diagram illustrating an example of a cross-component filter unit that can be configured to reduce reconstruction error according to one or more techniques of this disclosure.
[0031] [ Figure 11C ] Figure 11C This is a block diagram illustrating an example of a cross-component filter unit that can be configured to reduce reconstruction error according to one or more techniques of this disclosure.
[0032] [ Figure 11D ] Figure 11D This is a block diagram illustrating an example of a cross-component filter unit that can be configured to reduce reconstruction error according to one or more techniques of this disclosure.
[0033] [ Figure 12A ] Figure 12A This is a block diagram illustrating an example of a cross-component filter unit that can be configured to reduce reconstruction error according to one or more techniques of this disclosure.
[0034] [ Figure 12B ] Figure 12B This is a block diagram illustrating an example of a cross-component filter unit that can be configured to reduce reconstruction error according to one or more techniques of this disclosure.
[0035] [ Figure 13A ] Figure 13A This is a block diagram illustrating an example of a cross-component filter unit that can be configured to reduce reconstruction error according to one or more techniques of this disclosure.
[0036] [ Figure 13B ] Figure 13B This is a block diagram illustrating an example of a cross-component filter unit that can be configured to reduce reconstruction error according to one or more techniques of this disclosure.
[0037] [ Figure 13C ] Figure 13C This is a block diagram illustrating an example of a cross-component filter unit that can be configured to reduce reconstruction error according to one or more techniques of this disclosure.
[0038] [ Figure 14A ] Figure 14A This is a conceptual diagram illustrating an example of the location of filter coefficients that can be used for cross-component filtering according to one or more techniques of this disclosure.
[0039] [ Figure 14B ] Figure 14B This is a conceptual diagram illustrating an example of the location of filter coefficients that can be used for cross-component filtering according to one or more techniques of this disclosure.
[0040] [ Figure 14C ] Figure 14C This is a conceptual diagram illustrating an example of the location of filter coefficients that can be used for cross-component filtering according to one or more techniques of this disclosure.
[0041] [ Figure 14D ] Figure 14D This is a conceptual diagram illustrating an example of the location of filter coefficients that can be used for cross-component filtering according to one or more techniques of this disclosure.
[0042] [ Figure 14E ] Figure 14E This is a conceptual diagram illustrating an example of the location of filter coefficients that can be used for cross-component filtering according to one or more techniques of this disclosure.
[0043] [ Figure 14F ] Figure 14F This is a conceptual diagram illustrating an example of the location of filter coefficients that can be used for cross-component filtering according to one or more techniques of this disclosure.
[0044] [ Figure 15A ] Figure 15A This is a conceptual diagram illustrating an example of the location of filter coefficients that can be used for cross-component filtering according to one or more techniques of this disclosure.
[0045] [ Figure 15B ] Figure 15B This is a conceptual diagram illustrating an example of the location of filter coefficients that can be used for cross-component filtering according to one or more techniques of this disclosure.
[0046] [ Figure 15C ] Figure 15C This is a conceptual diagram illustrating an example of the location of filter coefficients that can be used for cross-component filtering according to one or more techniques of this disclosure.
[0047] [ Figure 15D ] Figure 15D This is a conceptual diagram illustrating an example of the location of filter coefficients that can be used for cross-component filtering according to one or more techniques of this disclosure.
[0048] [ Figure 16A ] Figure 16A This is a conceptual diagram illustrating an example of a virtual line buffer that can be used for cross-component filtering according to one or more techniques of this disclosure.
[0049] [ Figure 16B ] Figure 16B This is a conceptual diagram illustrating an example of a virtual line buffer that can be used for cross-component filtering according to one or more techniques of this disclosure.
[0050] [ Figure 16C ] Figure 16C This is a conceptual diagram illustrating an example of a virtual line buffer that can be used for cross-component filtering according to one or more techniques of this disclosure.
[0051] [ Figure 16D ] Figure 16D This is a conceptual diagram illustrating an example of a virtual line buffer that can be used for cross-component filtering according to one or more techniques of this disclosure.
[0052] [ Figure 17 ] Figure 17 This is a block diagram illustrating an example of a video decoder that can be configured to decode video data according to one or more techniques of this disclosure.
[0053] [ Figure 18 ] Figure 18This is a block diagram illustrating an example of a cross-component filter unit that can be configured to encode video data according to one or more techniques of this disclosure.
[0054] [ Figure 19A ] Figure 19A This is a block diagram illustrating an example of a cross-component filter unit that can be configured to reduce reconstruction error according to one or more techniques of this disclosure.
[0055] [ Figure 19B ] Figure 19B This is a block diagram illustrating an example of a cross-component filter unit that can be configured to reduce reconstruction error according to one or more techniques of this disclosure.
[0056] [ Figure 19C ] Figure 19C This is a block diagram illustrating an example of a cross-component filter unit that can be configured to reduce reconstruction error according to one or more techniques of this disclosure. Detailed Implementation
[0057] Generally, this disclosure describes various techniques for encoding video data. Specifically, this disclosure describes techniques for reducing reconstruction errors. It should be noted that although the techniques disclosed herein relate to ITU-T H.264, ITU-T H.265, JEM, JVET-N1001, and JVET-O2001, the techniques disclosed herein are generally applicable to video coding. For example, in addition to those techniques included in ITU-T H.265, JEM, JVET-N1001, and JVET-O2001, the coding techniques described herein can be incorporated into video coding systems (including video coding systems based on future video coding standards), including video block structures, intra-frame prediction techniques, inter-frame prediction techniques, transform techniques, filtering techniques, and / or other entropy coding techniques. Therefore, references to ITU-T H.264, ITU-T H.265, JEM, JVET-N1001, and JVET-O2001 are for descriptive purposes and should not be construed as limiting the scope of the techniques described herein. Furthermore, it should be noted that the inclusion of references in this paper by way of citation is for descriptive purposes and should not be construed as limiting or creating ambiguity regarding the terminology used herein. For example, where a definition of a term is provided in one of the incorporated references that differs from that in another incorporated reference and / or as used herein, the term should be interpreted in a manner that broadly includes each corresponding definition and / or in a manner that includes each particular definition in alternatives.
[0058] In one example, a method includes: receiving reconstructed sample data for a current component of video data, receiving reconstructed sample data for one or more additional components of the video data, deriving a cross-component filter based on data associated with one or more additional components of the video data, and applying the filter to the reconstructed sample data for the current component of the video data based on the derived cross-component filter and the reconstructed sample data for one or more additional components of the video data.
[0059] In one example, a device includes one or more processors configured to perform the following operations: receiving reconstructed sample data for a current component of video data, receiving reconstructed sample data for one or more additional components of the video data, deriving a cross-component filter based on data associated with one or more additional components of the video data, and applying the filter to the reconstructed sample data for the current component of the video data based on the derived cross-component filter and the reconstructed sample data for one or more additional components of the video data.
[0060] In one example, a non-transitory computer-readable storage medium includes instructions stored thereon that, when executed, cause one or more processors of the device to: receive reconstructed sample data for a current component of video data; receive reconstructed sample data for one or more additional components of the video data; derive a cross-component filter based on data associated with one or more additional components of the video data; and apply the filter to the reconstructed sample data for the current component of the video data based on the derived cross-component filter and the reconstructed sample data for one or more additional components of the video data.
[0061] In one example, an apparatus includes: means for receiving reconstructed sample data for one or more additional components of video data; means for deriving a cross-component filter based on data associated with one or more additional components of the video data; and means for applying the filter to reconstructed sample data for a current component of the video data based on the derived cross-component filter and the reconstructed sample data for one or more additional components of the video data.
[0062] Details of one or more examples are set forth in the following figures and description. Other features, objects, and advantages will become apparent from the description, figures, and claims.
[0063] Video content comprises a sequence of frames (or images). A series of frames may also be referred to as a group of pictures (GOP). Each video frame or image may be divided into one or more regions. Regions may be defined based on basic units (e.g., video blocks) and a set of rules defining regions. For example, a rule defining a region may be that a region must be an integer number of video blocks arranged in a rectangle. Furthermore, video blocks within a region may be ordered according to a scanning mode (e.g., raster scan). As used herein, the term "video block" may generally refer to a region of an image, or more specifically, to the largest array of sample values that can be predictably encoded, its sub-partitions, and / or corresponding structures. Additionally, the term "current video block" may refer to the region of an image that is being encoded or decoded. A video block may be defined as an array of sample values. It should be noted that in some cases, pixel values may be described as sample values comprising the corresponding components of the video data, which may also be referred to as color components (e.g., luminance (Y) and chrominance (Cb and Cr) components or red, green, and blue components). It should be noted that in some cases, the terms "pixel value" and "sample value" are used interchangeably. Furthermore, in some cases, a pixel or sample can be referred to as a pel. A video sampling format (also known as a chroma format) can be defined by the number of chroma samples included in a video block, relative to the number of luminance samples included in the video block. For example, in a 4:2:0 sampling format, the luminance component is sampled at twice the rate of the chroma components in both the horizontal and vertical directions.
[0064] Video encoders perform predictive coding on video blocks and their sub-partitions. Video blocks and their sub-partitions can be referred to as nodes. ITU-T H.264 specifies macroblocks comprising 16×16 luma samples. That is, in ITU-T H.264, pictures are segmented into macroblocks. ITU-T H.265 specifies a similar Code Tree Unit (CTU) structure (which may be referred to as a Maximum Code Unit (LCU)). In ITU-T H.265, pictures are segmented into CTUs. In ITU-T H.265, for pictures, the CTU size can be set to include 16×16, 32×32, or 64×64 luma samples. In ITU-T H.265, a CTU consists of a corresponding Code Tree Block (CTB) for each component of the video data (e.g., luma (Y) and chrominance (Cb and Cr)). It should be noted that a video with one luma component and two corresponding chrominance components can be described as having two channels, namely, a luma channel and a chrominance channel. Furthermore, in ITU-T H.265, CTUs can be partitioned according to a quadtree (QT) partitioning structure, which allows the CTU's CTB to be divided into coded blocks (CBs). That is, in ITU-T H.265, a CTU can be divided into quadtree leaf nodes. According to ITU-T H.265, a luma CB, along with two corresponding chroma CBs and associated syntax elements, is called a coding unit (CU). In ITU-T H.265, the minimum permissible size of a CB can be signaled. In ITU-T H.265, the minimum permissible size of a luma CB is 8×8 luma samples. In ITU-T H.265, the decision to code a picture region using intra-frame prediction or inter-frame prediction is made at the CU level.
[0065] In ITU-T H.265, a CU (Cubic Component Unit) is associated with a Prediction Unit (PU) structure that has its root at the CU. In ITU-T H.265, the PU structure allows the partitioning of the Luminance CB (Cubic Block) and Chromaticity CB to generate corresponding reference samples. That is, in ITU-T H.265, the Luminance CB and Chromaticity CB can be partitioned into corresponding Luminance Prediction Blocks (CBs) and Chromaticity Prediction Blocks (PBs), where each PB comprises a block of sample values to which the same prediction is applied. In ITU-T H.265, a CB can be divided into one, two, or four PBs. ITU-T H.265 supports PB sizes from 64×64 samples down to 4×4 samples. In ITU-T H.265, square PBs are supported for intra-frame prediction, where a CB can form a PB or can be partitioned into four square PBs. In addition to square PBs, ITU-T H.265 also supports rectangular PBs for inter-frame prediction, where a CB can be halved vertically or horizontally to form a PB. Furthermore, it should be noted that in ITU-T H.265, for inter-frame prediction, four asymmetric PB partitions are supported, where the CB is divided into two PBs at one-quarter of the height (top or bottom) or width (left or right) of the CB. Intra-frame prediction data (e.g., intra-frame prediction mode syntax elements) or inter-frame prediction data (e.g., motion data syntax elements) corresponding to the PB are used to generate reference and / or prediction sample values for the PB.
[0066] JEM specifies a CTU with a maximum size of 256×256 luminance samples. JEM specifies a Quadtree Plus Binary Tree (QTBT) block structure. In JEM, the QTBT structure allows the quadtree leaf nodes to be further partitioned by a binary tree (BT) structure. That is, in JEM, the binary tree structure allows the quadtree leaf nodes to be recursively partitioned vertically or horizontally. In JVET-N1001 and JVET-O2001, the CTU is partitioned according to a Quadtree Plus Multi-Type Tree (QTMT or QT+MTT) structure. The QTMT in JVET-N1001 and JVET-O2001 is similar to the QTBT in JEM. However, in JVET-N1001 and JVET-O2001, in addition to indicating binary partitioning, the multi-type tree can also indicate so-called ternary (or ternary tree (TT)) partitioning. Ternary partitioning divides a block vertically or horizontally into three blocks. In the case of a vertical TT split, the block is divided at one-quarter of its width from the left edge and at one-quarter of its width from the right edge; and in the case of a horizontal TT split, the block is divided at one-quarter of its height from the top edge and at one-quarter of its height from the bottom edge. See again. Figure 1 , Figure 1 This illustrates an example where a CTU is partitioned into quadtree leaf nodes, and these quadtree leaf nodes are further partitioned based on either BT or TT partitioning. That is, in Figure 1 In the diagram, dashed lines indicate additional binary and ternary partitions in a quadtree.
[0067] As described above, each video frame or picture can be divided into one or more regions. For example, according to ITU-T H.265, each video frame or picture can be divided into one or more slices, and further divided into one or more tiles, wherein each slice includes a sequence of CTUs (e.g., arranged in raster scan order), and wherein a tile is a sequence of CTUs corresponding to a rectangular area of the picture. It should be noted that, in ITU-T H.265, a slice is a sequence of one or more slice segments that begin with an independent slice segment and include all subsequent subordinate slice segments (if any) preceding the next independent slice segment (if any). A slice segment (such as a piece) is a sequence of CTUs. Therefore, in some cases, the terms "slice" and "slice segment" are used interchangeably to refer to a sequence of CTUs arranged in raster scan order. Furthermore, it should be noted that, in ITU-T H.265, a tile may consist of CTUs contained in more than one slice, and a slice may consist of CTUs contained in more than one tile. However, ITU-T H.265 specifies that one or both of the following conditions must be met: (1) all CTUs in a segment belong to the same tile; and (2) all CTUs in a tile belong to the same slice.
[0068] Regarding JVET-N1001 and JVET-O2001, slices need to consist of an integer number of tiles, not just an integer number of CTUs. In JVET-N1001 and JVET-O2001, a tile is a rectangular row region of CTUs within a specific tile in an image. Furthermore, in JVET-N1001 and JVET-O2001, a tile can be divided into multiple tiles, each tile consisting of one or more rows of CTUs within the tile. Tiles not divided into multiple tiles are also referred to as tiles. However, tiles that are a proper subset of a tile are not referred to as tiles. Therefore, some video coding techniques may or may not support slices comprising a set of CTUs that do not form an image. Additionally, it should be noted that in some cases, slices may need to consist of an integer number of complete tiles, and in such cases, the slice is referred to as a tile group. The techniques described herein are applicable to tiles, slices, tiles, and / or tile groups. Figure 1 This is a concept diagram showing an example of a group of images including slices. Figure 1 In the example shown, Pic3 is depicted as comprising two slices (i.e., slice 0 and slice 1). Figure 1In the example shown, slice 0 includes one brick, namely brick 0, and slice 1 includes two bricks, namely brick 1 and brick 2. It should be noted that in some cases, slice 0 and slice 1 may meet the requirements of a tile and / or a tile group and be classified as a tile and / or a tile group.
[0069] For intra-frame predictive coding, the intra-frame prediction mode can specify the location of a reference sample within the image. In ITU-T H.265, the defined possible intra-frame prediction modes include planar (i.e., surface-fitting) prediction modes, DC (i.e., flat global average) prediction modes, and 33 angular prediction modes (predMode: 2-34). In JEM, the defined possible intra-frame prediction modes include planar prediction modes, DC prediction modes, and 65 angular prediction modes. It should be noted that planar prediction modes and DC prediction modes can be referred to as non-directional prediction modes, and angular prediction modes can be referred to as directional prediction modes. It should be noted that the techniques described herein are generally applicable regardless of the number of defined possible prediction modes.
[0070] For inter-frame predictive coding, a reference picture is determined, and motion vectors (MVs) identify samples in that reference picture used to generate predictions for the current video block. For example, reference sample values located in one or more previously encoded pictures can be used to predict the current video block, and motion vectors are used to indicate the position of the reference block relative to the current video block. Motion vectors can describe, for example, the horizontal displacement component of the motion vector (i.e., MV). x ), the vertical displacement component of the motion vector (i.e., MV) yThe resolution of the motion vectors (e.g., quarter-pixel precision, half-pixel precision, one-pixel precision, two-pixel precision, four-pixel precision) is used. Previously decoded images (which may include images output before or after the current image) can be organized into one or more lists of reference images and identified using reference image index values. Furthermore, in inter-frame predictive coding, single prediction refers to generating a prediction using sample values from a single reference image, while dual prediction refers to generating a prediction using corresponding sample values from two reference images. That is, in single prediction, a single reference image and its corresponding motion vector are used to generate a prediction for the current video block, while in dual prediction, a first reference image and its corresponding first motion vector, and a second reference image and its corresponding second motion vector are used to generate a prediction for the current video block. In dual prediction, the corresponding sample values are combined (e.g., added, rounded, and cropped, or averaged according to weights) to generate a prediction. Images and their regions can be classified based on which types of prediction patterns are available for encoding their video blocks. In other words, for regions of type B (e.g., B slices), dual prediction, single prediction, and intra-prediction modes can be used; for regions of type P (e.g., P slices), single prediction and intra-prediction modes can be used; and for regions of type I (e.g., I slices), only intra-prediction mode can be used. As described above, reference images are identified by reference indices. For example, for P slices, a single reference image list RefPicList0 can exist, and for B slices, in addition to RefPicList0, a second independent reference image list RefPicList1 can exist. It should be noted that for single prediction in B slices, either RefPicList0 or RefPicList1 can be used to generate the prediction. Furthermore, it should be noted that during the decoding process, at the start of decoding an image, a reference image list is generated from previously decoded images stored in the Decoding Image Buffer (DPB).
[0071] Furthermore, the coding standard supports various motion vector prediction modes. Motion vector prediction enables the derivation of motion vector values for the current video block based on another motion vector. For example, a set of candidate blocks with associated motion information can be derived from the spatially and temporally adjacent blocks of the current video block. Additionally, the generated (or default) motion information can be used for motion vector prediction. Examples of motion vector prediction include Advanced Motion Vector Prediction (AMVP), Temporal Motion Vector Prediction (TMVP), the so-called "merge" mode, and "skip" and "direct" motion inference. Other examples of motion vector prediction include Advanced Temporal Motion Vector Prediction (ATMVP) and Spatial-Temporal Motion Vector Prediction (STMVP). For motion vector prediction, both the video encoder and video decoder perform the same process to derive a set of candidates. Therefore, for the current video block, the same set of candidates is generated during encoding and decoding.
[0072] As mentioned above, for inter-frame predictive coding, reference samples from previously encoded images are used to encode video blocks in the current image. The previously encoded image that can be used as a reference when encoding the current image is called the reference image. It should be noted that the decoding order does not necessarily correspond to the image output order, i.e., the temporal order of images in the video sequence. In ITU-T H.265, when an image is decoded, it is stored in a decoded image buffer (DPB) (which may be called a frame buffer, reference buffer, reference image buffer, etc.). In ITU-T H.265, images stored in the DPB are removed from the DPB when output and are no longer needed for encoding subsequent images. In ITU-T H.265, after decoding the slice header, i.e., at the start of image decoding, a determination is made once for each image whether it should be removed from the DPB. For example, the reference... Figure 1 Pic3 is shown with reference to Pic2. Similarly, Pic4 is shown with reference to Pic1. Regarding Figure 1Assuming the number of images corresponds to the decoding order, the DPB will be populated as follows: After decoding Pic1, the DPB will include {Pic1}; at the start of decoding Pic2, the DPB will include {Pic1}; after decoding Pic2, the DPB will include {Pic1, Pic2}; at the start of decoding Pic3, the DPB will include {Pic1, Pic2}. Then, Pic3 will be decoded with reference to Pic2, and after decoding Pic3, the DPB will include {Pic1, Pic2, Pic3}. At the start of decoding Pic4, images Pic2 and Pic3 will be marked for removal from the DPB because they are not required for decoding Pic4 (or any subsequent images, not shown), and assuming Pic2 and Pic3 have already been output, the DPB will be updated to include {Pic1}. Pic4 will then be decoded using reference Pic1. The process of marking images to remove them from the DPB can be called Reference Picture Set (RPS) management.
[0073] As described above, intra-frame prediction data or inter-frame prediction data is used to generate reference sample values for blocks of sample values. The difference between sample values included in the current PB or another type of picture region structure and the associated reference samples (e.g., those generated using prediction) can be referred to as residual data. Residual data can include a corresponding array of differences corresponding to each component of the video data. The residual data may be in the pixel domain. Transformations such as Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), integer transform, wavelet transform, or conceptually similar transforms can be applied to the difference array to generate transform coefficients. It should be noted that in ITU-T H.265, JVET-N1001, and JVET-O2001, the CU is associated with a Transform Unit (TU) structure having its root at the CU level. That is, to generate transform coefficients, the array of differences can be partitioned (e.g., four 8×8 transforms can be applied to a 16×16 residual value array). Such a subdivision of the differences for each component of the video data can be referred to as a Transform Block (TB). It should be noted that in some cases, a core transform and a subsequent second transform can be applied (in a video encoder) to generate transform coefficients. For a video decoder, the order of the transforms is reversed.
[0074] Quantization can be performed directly on transform coefficients or residual sample values (e.g., in the case of palette-encoded quantization). Quantization approximates transform coefficients by limiting the amplitude to a specified set of values. Quantization essentially scales the transform coefficients to change the amount of data needed to represent a set of transform coefficients. Quantization may include dividing the transform coefficient (or the value obtained by adding an offset value to the transform coefficient) by a quantization scaling factor and any associated rounding function (e.g., rounding to the nearest integer). The quantized transform coefficients may be referred to as coefficient bit values. Inverse quantization (or “dequantization”) may include multiplying the coefficient bit value by the quantization scaling factor, and any reciprocal rounding or offset addition operations. It should be noted that, as used herein, the term quantization process may refer in some cases to division by a scaling factor to generate a bit value, and in some cases to multiplication by a scaling factor to recover the transform coefficients. That is, quantization process may refer to quantization in some cases and inverse quantization in others. Furthermore, it should be noted that although some examples below describe quantization processes for arithmetic operations related to decimal notation, such descriptions are for illustrative purposes and should not be construed as limiting. For example, the techniques described herein can be implemented in devices using binary arithmetic, etc. For example, the multiplication and division operations described herein can be implemented using bit shifting operations, etc.
[0075] Entropy coding techniques can be used to entropy-encode quantized transform coefficients and syntax elements (e.g., syntax elements indicating the coding structure of video blocks). The entropy coding process involves encoding the syntax element values using a lossless data compression algorithm. Examples of entropy coding techniques include Content Adaptive Variable Length Coding (CAVLC), Context Adaptive Binary Arithmetic Coding (CABAC), Probability Interval Partition Entropy Coding (PIPE), etc. The entropy-encoded quantized transform coefficients and the corresponding entropy-encoded syntax elements can form a compliant bitstream that can be used to reproduce video data at the video decoder. The entropy coding process, such as CABAC, may include binarizing the syntax elements. Binarization is the process of converting the values of syntax elements into a sequence of one or more bits. These bits may be referred to as "bins". Binarization may include one or a combination of the following coding techniques: fixed-length coding, unary coding, truncated unary coding, truncated Rice coding, Golomb coding, k-order exponential Golomb coding, and Golomb-Rice coding. For example, binarization may include representing the integer value 5 of a syntax element as 00000101 using an 8-bit fixed-length binarization technique, or representing the integer value 5 as 11110 using a unary coding binarization technique. As used herein, the terms fixed-length coding, unary coding, truncated unary coding, truncated Rice coding, Golomb coding, k-order exponential Golomb coding, and Golomb-Rice coding may refer to a general implementation of these techniques and / or a more specific implementation of these coding techniques. For example, a Golomb-Rice coding implementation may be specifically defined according to a video coding standard. In the CABAC example, for a particular bin, the context provides the bin's maximum probability state (MPS) value (i.e., the bin's MPS is either 0 or 1), and the probability value that the bin is in either the MPS or minimum probability state (LPS). For example, the context may indicate that the bin's MPS is 0 and the probability that the bin is 1 is 0.3. It should be noted that the context may be determined based on the values of bins in previous encodings that include the current syntax element and / or bins in previously encoded syntax elements. For example, the value of a syntax element associated with an adjacent video block can be used to determine the context of the current bin.
[0076] Figures 2A to 2B This is a conceptual diagram illustrating an example of encoding video data blocks. (Example:) Figure 2A As shown, bit-order values are generated by subtracting a set of predicted values from the current video data block to produce a residual, performing a transformation on the residual, and quantizing the transform coefficients. This process encodes the current block of video data (e.g., the CB corresponding to a video component). Figure 2B As shown, the current video data block is decoded by performing inverse quantization on the bit-order values, performing an inverse transform, and adding a set of predicted values to the resulting residual. It should be noted that, in Figures 2A to 2BIn the example, the sample values of the reconstructed block are different from the sample values of the current video block being encoded. Specifically, Figure 2B The reconstruction error is shown; it is the difference between the current block and the reconstructed block. Thus, the encoding can be considered lossy. However, the difference in sample values can be considered acceptable or imperceptible to the viewer of the reconstructed video.
[0077] In addition, such as Figures 2A to 2B As shown, a scaling factor array is used to generate coefficient bit values. In ITU-T H.265, the scaling factor array is generated by selecting a scaling matrix and multiplying each entry in the scaling matrix by a quantization scaling factor. In ITU-T H.265, the scaling matrix is selected in part based on the prediction mode and color components, where scaling matrices of the following sizes are defined: 4×4, 8×8, 16×16, and 32×32. It should be noted that in some examples, the scaling matrix may provide the same value for each entry (i.e., scaling all coefficients by a single value). In ITU-T H.265, the value of the quantization scaling factor can be determined by the quantization parameter QP. In ITU-T H.265, for an 8-bit bit depth, QP can take 52 values from 0 to 51, and a change of 1 in QP typically corresponds to a change of approximately 12% in the value of the quantization scaling factor. Furthermore, in ITU-T H.265, a set of transform coefficients' QP values can be derived using predicted quantization parameter values (which may be referred to as predicted QP values or QP prediction values) and optional quantization parameter increment values (which may be referred to as QP increment values or incremental QP values) that are signaled. In ITU-T H.265, quantization parameters can be updated for each CU, and corresponding quantization parameters can be derived for each of the luminance and chrominance channels.
[0078] As mentioned above, relative to Figures 2A to 2BIn the example shown, the sample values of the reconstructed block may differ from the sample values of the currently encoded video block. Additionally, it should be noted that in some cases, encoding video data block-by-block can lead to artifacts (e.g., so-called block artifacts, band artifacts, etc.). For example, block artifacts can cause the boundaries of the encoded blocks in the reconstructed video data to be visually perceptible to the user. Thus, the reconstructed sample values can be modified to minimize the difference between the sample values of the currently encoded video block and the reconstructed block and / or minimize artifacts introduced by the video encoding process. Such modifications are generally referred to as filtering. It should be noted that filtering can occur as part of an in-loop filtering process or a post-loop filtering process. For an in-loop filtering process, the resulting sample values can be used to predict video blocks (e.g., stored in a reference frame buffer for subsequent encoding at the video encoder and subsequent decoding at the video decoder). For a post-loop filtering process, the resulting sample values are output only as part of the decoding process (e.g., not used for subsequent encoding). For example, in the case of a video decoder, during an in-loop filtering process, the sample values produced by the filtering reconstruction block will be used for subsequent decoding (e.g., stored in a reference buffer) and will be output (e.g., output to a display). During a post-loop filtering process, the reconstruction block will be used for subsequent decoding, and the sample values produced by the filtering reconstruction block will be output.
[0079] Deblocking (or unblocking), deblocking filtering, or applying a deblocking filter refers to the process of smoothing the boundaries of adjacent reconstructed video blocks (i.e., making the boundaries less noticeable to the observer). Smoothing the boundaries of adjacent reconstructed video blocks can include modifying sample values in rows or columns included near the boundary. ITU-T H.265 provides scenarios for applying deblocking filters to reconstructed sample values as part of a loop filtering process. ITU-T H.265 includes two types of deblocking filters that can be used to modify luminance samples: a Strong Filter, which modifies sample values in three rows or columns adjacent to the boundary; and a Weak Filter, which modifies sample values in rows or columns immediately adjacent to the boundary and conditionally modifies sample values in the second row or column starting from the boundary. Additionally, ITU-T H.265 includes one type of filter that can be used to modify chrominance samples: a standard filter.
[0080] In addition to applying unblocking filters as part of the loop filtering process, ITU-T H.265 also provides scenarios for applying Sample Adaptive Offset (SAO) filtering during the loop filtering process. In ITU-T H.265, SAO is a process of modifying unblocked sample values in a region by conditionally adding offset values. ITU-T H.265 provides two types of SAO filters that can be applied to the CTB: band offset or edge offset. For each of band offset and edge offset, the bitstream includes four offset values. For band offset, the applied offset depends on the amplitude of the sample values (e.g., amplitudes are mapped to bands, which are mapped to four offsets already transmitted with the signal). For edge offset, the applied offset depends on the CTB having one of horizontal, vertical, first diagonal, or second diagonal edge classifications (e.g., classifications are mapped to four offsets already transmitted with the signal).
[0081] Another type of filtering process includes the so-called Adaptive Loop Filter (ALF). The JEM specifies the use of a block-based adaptive ALF. In the JEM, the ALF is applied after the SAO filter. It should be noted that the ALF can be applied to the reconstructed samples independently of other filtering techniques. The process of applying the ALF specified in the JEM at the video encoder can be summarized as follows: (1) each 2×2 block for the luminance component of the reconstructed image is classified according to a classification index; (2) a set of filter coefficients for each classification index is derived; (3) a filtering decision is determined for the luminance component; (4) a filtering decision is determined for the chrominance component; and (5) filter parameters (e.g., coefficients and decisions) are sent as a signal.
[0082] According to the ALF specified in JEM, each 2×2 block is classified according to the classification index C, where C is an integer in the range of 0 to 24, including end values. C is derived based on the quantized values of its directionality D and activity  according to the following formula:
[0083]
[0084] Where D and Â, the gradients in the horizontal, vertical, and two diagonal directions, are calculated using the 1-D Laplace, as shown below:
[0085]
[0086] Here, indices i and j refer to the coordinates of the top-left sample in the 2×2 block, and R(i,j) indicates the reconstructed sample with coordinates (i,j).
[0087] The maximum and minimum values of the gradients in the horizontal and vertical directions can be set as follows:
[0088]
[0089] Furthermore, the maximum and minimum values of the gradients in the two diagonal directions can be set as follows:
[0090]
[0091] In JEM, to derive the value of directionality D, the maximum and minimum values are compared with each other and with two thresholds t1 and t2:
[0092] Step 1. If and If both are true, then set D to 0.
[0093] Step 2. If If yes, continue from step 3; otherwise, continue from step 4.
[0094] Step 3. If If so, set D to 2; otherwise, set D to 1.
[0095] Step 4. If If so, set D to 4; otherwise, set D to 3.
[0096] In JEM, the activity value A is calculated as follows:
[0097]
[0098] A is further quantized to a range of 0 to 4, including the end values, and the quantized value is denoted as Ã.
[0099] As described above, applying the ALF specified in JEM at the video encoder involves deriving a set of filter coefficients for each classification index and determining a filtering decision. It should be noted that the derivation of the filter coefficient set and the determination of the filtering decision can be an iterative process. That is, the filter coefficient set can be updated based on the filtering decision, and the filtering decision can be updated based on the updated filter coefficient set, and this can be repeated multiple times. Furthermore, the video encoder can implement various specialized algorithms to determine the filter coefficient set and / or the filtering decision. Regardless of how the filter coefficient set is derived for each classification index or how the filtering decision is determined, the techniques described herein are generally applicable.
[0100] In one example, a set of optimal filter coefficients is derived by initially deriving a set of optimal filter coefficients for each classification index. The optimal filter coefficients are derived by comparing the desired sample values (i.e., sample values from the source video) with the reconstructed sample values after filtering is applied, and by minimizing the sum of squared errors (SSE) between the desired and reconstructed sample values after filtering is performed. The optimal coefficients derived for each set are then used to perform a basic filter on the reconstructed samples to analyze the effect of the ALF. That is, the desired sample values, the reconstructed sample values before applying the ALF, and the reconstructed sample values after applying the ALF can be compared to determine the effectiveness of applying the ALF using the optimal coefficients.
[0101] Based on the specified ALF in JEM, each reconstructed sample R(i,j) is filtered by determining the obtained sample value R'(i,j) according to the following formula, where L represents the filter length and f(k,l) represents the decoded filter coefficients.
[0102]
[0103] It should be noted that JEM defines three filter shapes (5×5 rhombus, 7×7 rhombus, and 9×9 rhombus). It should also be noted that in JEM, geometric transformations are applied to the filter coefficients f(k,l), specifically depending on the gradient value: g v g h g d1 g d2 As provided in Table 1.
[0104]
[0105] The diagonal, vertical flip, and rotation are defined as follows:
[0106] diagonal f D (k, l) = f(l, k),
[0107] Vertical flip: f v (k, l) = f(k,K - l – 1)
[0108] Rotation: f R (k, l) = f(K - l - 1, k)
[0109] Where K is the size of the filter, and 0 ≤ k, 1 ≤ Kl are the coefficient coordinates, such that position (0,0) is located in the upper left corner and position (Kl,Kl) is located in the lower right corner.
[0110] JEM provides a scenario where up to 25 sets of luminance filter coefficients (i.e., one for each possible classification index) can be signaled. Therefore, optimal coefficients can be signaled for each classification index appearing in the corresponding image region. However, to optimize the amount of data required to correlate the signaled filter coefficient sets with the filter effect, rate-distortion (RD) optimization can be performed. For example, JEM provides a scenario where filter coefficients from adjacent classification groups can be combined and signaled using an array that maps a set of filter coefficients to each classification index. Furthermore, JEM provides a scenario where temporal coefficient prediction can be used for signaled coefficients. That is, JEM provides a scenario where the filter coefficient set for the current image is predicted based on the filter coefficient set of the reference image by inheriting a set of filter coefficients used for the reference image. JEM also provides a scenario where a fixed set of 16 filters can be used to predict the filter coefficient set for intra-frame prediction images. As mentioned above, the derivation of the filter coefficient set and the determination of filtering decisions can be an iterative process. That is, for example, the shape of the ALF can be determined based on how many sets of filter coefficients are signaled, and similarly, whether the ALF is applied to a region of the image can be based on the shape of the signaled filter coefficient set and / or the filter. It should be noted that for an ALF filter, each component uses a set of sample values from the corresponding component as input and derives the output sample values. That is, the ALF filter is applied to each component independently of the data in the other component. Furthermore, it should be noted that JVET-N1001 and JVET-O2001 specify deblocking filters, SAO filters, and ALF filters, which can be described as typically based on the deblocking filters, SAO filters, and ALF filters provided in ITU-T H.265 and JEM.
[0111] The number of chroma samples included in a CU can be defined relative to the number of luma samples included in the CU. For example, in a 4:2:0 sampling format, the sampling rate of the luma component is twice the sampling rate of the chroma components in both the horizontal and vertical directions. Therefore, for a CU formatted according to the 4:2:0 format, the width and height of the sample arrays used for the luma components are twice the width and height of each sample array used for the chroma components. Figure 3 This is a conceptual diagram illustrating an example of a coding unit formatted according to the 4:2:0 sample format. Figure 3 This shows the relative positions of the chromaticity samples with respect to the luminance samples within the CU. As mentioned above, the CU is typically defined based on the number of horizontal and vertical luminance samples. Therefore, as... Figure 3 As shown, the 16×16 CU, formatted according to the 4:2:0 sample format, includes 16×16 samples for the luma component and 8×8 samples for each chroma component. Furthermore, in Figure 3The example shown illustrates the relative positions of chroma samples to luma samples for adjacent video blocks of a 16×16 CU. For a CU formatted in 4:2:2 format, the width of the luma component sample array is twice the width of the chroma component sample array, but the height of the luma component sample array is equal to the height of the chroma component sample array. Furthermore, for a CU formatted in 4:4:4 format, the luma component sample array has the same width and height as the chroma component sample array. Reference Figure 3 For luminance samples, the sample line immediately above the video block can be called reference line 0 (RL0), and subsequent sample lines can be called reference line 1 (RL1), reference line 2 (RL2), and reference line 3 (RL3), respectively. Similarly, the sample column to the left of the current video block can be classified as reference lines in a similar manner (i.e., the sample line immediately to the left of the video block can be called reference line 0 (RL0)).
[0112] It should be noted that for sampling formats, such as the 4:2:0 sample format, the chroma position type can be specified. That is, for example, for the 4:2:0 sample format, horizontal and vertical offset values indicating relative spatial positioning can be specified for the chroma samples relative to the luminance samples. Table 2 provides the definitions of HorizontalOffsetC and VerticalOffsetC for the five chroma position types provided in JVET-N1001 and JVET-O2001. Additionally, Figures 4A to 4F The chromaticity position type specified for the 4:2:0 sample format is shown in JVET-N1001 and JVET-O2001.
[0113]
[0114] The following arithmetic operators can be used for the formulas used in this article:
[0115]
[0116] In addition, the following logical operators can be used:
[0117]
[0118] In addition, the following relational operators can be used:
[0119]
[0120] In addition, the following bitwise operators can be used:
[0121]
[0122] In addition, the following assignment operators can be used:
[0123]
[0124] In addition, the following mathematical functions can be used:
[0125]
[0126] Floor(x), the largest integer less than or equal to x.
[0127] Log2(x) is the base-2 logarithm of x.
[0128]
[0129]
[0130]
[0131] Figure 5 This is a block diagram illustrating an example of a system that can be configured to encode (e.g., encode and / or decode) video data according to one or more techniques of this disclosure. System 100 represents an example of a system that can perform video encoding using one or more techniques of this disclosure. Figure 5 As shown, system 100 includes source device 102, communication medium 110, and target device 120. Figure 5 In the example shown, source device 102 may include any device configured to encode video data and transmit the encoded video data to communication medium 110. Target device 120 may include any device configured to receive and decode the encoded video data via communication medium 110. Source device 102 and / or target device 120 may include computing devices equipped for wired and / or wireless communication, and may include set-top boxes, digital video recorders, televisions, desktop computers, laptops or tablets, game consoles, mobile devices including, for example, "smart" phones, cellular phones, personal gaming devices, and medical imaging equipment.
[0132] Communication medium 110 may include any combination of wireless and wired communication media and / or storage devices. Communication medium 110 may include coaxial cable, fiber optic cable, twisted-pair cable, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other device that can be used to facilitate communication between various devices and sites. Communication medium 110 may include one or more networks. For example, communication medium 110 may include a network configured to allow access to the World Wide Web, such as the Internet. The network may operate according to a combination of one or more telecommunications protocols. Telecommunication protocols may include proprietary aspects and / or may include standardized telecommunications protocols. Examples of standardized telecommunications protocols include the Digital Video Broadcasting (DVB) standard, the Advanced Television Systems Committee (ATSC) standard, the Integrated Services Digital Broadcasting (ISDB) standard, the Cable Data Services Interface Specification (DOCSIS) standard, the Global System for Mobile Communications (GSM) standard, the Code Division Multiple Access (CDMA) standard, the 3rd Generation Partnership Project (3GPP) standard, the European Telecommunications Standards Institute (ETSI) standard, the Internet Protocol (IP) standard, the Wireless Application Protocol (WAP) standard, and the Institute of Electrical and Electronics Engineers (IEEE) standard.
[0133] Storage devices can include any type of device or storage medium capable of storing data. Storage media can include tangible or non-transitory computer-readable media. Computer-readable media can include optical discs, flash memory, magnetic storage, or any other suitable digital storage medium. In some examples, a memory device or a portion thereof may be described as non-volatile memory, and in other examples, a portion of a memory device may be described as volatile memory. Examples of volatile memory can include random access memory (RAM), dynamic random access memory (DRAM), and static random access memory (SRAM). Examples of non-volatile memory can include magnetic hard disks, optical discs, floppy disks, flash memory, or electrically programmable memory (EPROM) or electrically erasable and programmable (EEPROM) memory. Storage devices can include memory cards (e.g., secure digital (SD) memory cards), internal / external hard disk drives, and / or internal / external solid-state drives. Data can be stored on the storage device according to defined file formats.
[0134] Refer again Figure 5Source device 102 includes a video source 104, a video encoder 106, and an interface 108. The video source 104 may include any device configured to capture and / or store video data. For example, the video source 104 may include a camera and a storage device operatively coupled thereto. The video encoder 106 may include any device configured to receive video data and generate a compatible bitstream representing the video data. A compatible bitstream may refer to a bitstream from which a video decoder can receive and reproduce video data. Aspects of a compliant bitstream may be defined according to a video coding standard. When generating a compatible bitstream, the video encoder 106 may compress the video data. Compression may be lossy (perceptible or imperceptible) or lossless. The interface 108 may include any device configured to receive a compatible video bitstream and transmit and / or store the compatible video bitstream to a communication medium. The interface 108 may include a network interface card such as an Ethernet card and may include an optical transceiver, an RF transceiver, or any other type of device capable of transmitting and / or receiving information. Furthermore, interface 108 may include a computer system interface that allows compatible video bitstreams to be stored on a storage device. For example, interface 108 may include protocols supporting Peripheral Component Interconnect (PCI) and Peripheral Component Fast Interconnect (PCIe) bus protocols, dedicated bus protocols, Universal Serial Bus (USB) protocols, and I / O protocols. 2 C's chipset or any other logical and physical structure that can be used to interconnect peer devices.
[0135] Refer again Figure 5 The target device 120 includes an interface 122, a video decoder 124, and a display 126. Interface 122 may include any device configured to receive compatible video bitstreams from a communication medium. Interface 108 may include a network interface card such as an Ethernet card, and may include an optical transceiver, an RF transceiver, or any other type of device capable of receiving and / or transmitting information. Furthermore, interface 122 may include a computer system interface that allows retrieval of compatible video bitstreams from a storage device. For example, interface 122 may include protocols supporting PCI and PCIe bus protocols, dedicated bus protocols, USB protocols, and I / O protocols. 2 The chipset of C, or any other logical and physical structure that can be used to interconnect peer devices. The video decoder 124 may include any device configured to receive compatible bitstreams and / or acceptable variations thereof, and reproduce video data from them. The display 126 may include any device configured to display video data. The display 126 may include one of a variety of display devices such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display. The display 126 may include a high-definition display or an ultra-high-definition display. It should be noted that, although in Figure 5In the example shown, video decoder 124 is described as outputting data to display 126, but video decoder 124 can be configured to output video data to various types of devices and / or their sub-components. For example, video decoder 124 can be configured to output video data to any communication medium, as described herein.
[0136] Figure 6 This is a block diagram illustrating an example of a video encoder 200 that can implement the techniques described herein for encoding video data. It should be noted that although the exemplary video encoder 200 is shown as having different functional blocks, such illustrations are intended for descriptive purposes and do not limit the video encoder 200 and / or its sub-components to a particular hardware or software architecture. The functionality of the video encoder 200 can be implemented using any combination of hardware, firmware, and / or software implementations. In one example, the video encoder 200 may be configured to encode video data according to the techniques described herein. The video encoder 200 may perform intra-frame predictive coding and inter-frame predictive coding of picture regions, and thus may be referred to as a hybrid video encoder. Figure 6 In the example shown, video encoder 200 receives a source video block. In some examples, the source video block may include picture regions that have been partitioned according to the coding structure. For example, source video data may include macroblocks, CTUs, CBs, their sub-partitions, and / or additional equivalent coding units. In some examples, video encoder 200 may be configured to perform additional subdivision of the source video block. It should be noted that some of the techniques described herein are generally applicable to video coding, regardless of how the source video data is partitioned before and / or during encoding. Figure 6 In the example shown, the video encoder 200 includes a summer 202, a transform coefficient generator 204, a coefficient quantization unit 206, an inverse quantization / transformation processing unit 208, a summer 210, an intra-frame prediction processing unit 212, an inter-frame prediction processing unit 214, a filter unit 216, and an entropy coding unit 218.
[0137] like Figure 6As shown, video encoder 200 receives source video blocks and outputs a bitstream. Video encoder 200 generates residual data by subtracting a predicted video block from the source video block. Summer 202 represents the component configured to perform this subtraction operation. In one example, the subtraction of the video block occurs in the pixel domain. Transform coefficient generator 204 applies a transform, such as Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), or a conceptually similar transform, to its residual block or sub-partition (e.g., four 8×8 transforms can be applied to a 16×16 residual value array) to generate a set of residual transform coefficients. Transform coefficient generator 204 can be configured to perform any and all combinations of transforms included in the discrete trigonometric transform series. As mentioned above, in ITU-T H.265, TB is limited to the following sizes: 4×4, 8×8, 16×16, and 32×32. In one example, the transform coefficient generator 204 can be configured to perform transforms based on arrays of sizes 4×4, 8×8, 16×16, and 32×32. In another example, the transform coefficient generator 204 can be further configured to perform transforms based on arrays of other sizes. Specifically, in some cases, performing transforms on rectangular arrays composed of different values may be useful. In one example, the transform coefficient generator 204 can be configured to perform transforms based on array sizes of 2×2, 2×4N, 4M×2, and / or 4M×4N. In one example, a two-dimensional (2D) M×N inverse transform can be implemented as a one-dimensional (1D) M-point inverse transform followed by a 1D N-point inverse transform. In one example, a 2D inverse transform can be implemented as a 1D N-point vertical transform followed by a 1D N-point horizontal transform. In another example, a 2D inverse transform can be implemented as a 1D N-point horizontal transform followed by a 1D N-point vertical transform. The transform coefficient generator 204 can output transform coefficients to the coefficient quantization unit 206.
[0138] Coefficient quantization unit 206 can be configured to perform quantization of the transform coefficients. As described above, the degree of quantization can be modified by adjusting the quantization parameters. Coefficient quantization unit 206 can be further configured to determine the quantization parameters and output QP data (e.g., data for determining the quantization group size and / or incremental QP value), which the video decoder can use to reconstruct the quantization parameters to perform inverse quantization during video decoding. It should be noted that in other examples, one or more additional or alternative parameters (e.g., scaling factors) can be used to determine the quantization level. The techniques described herein are generally applicable to determining the quantization level of transform coefficients corresponding to another component of video data based on the quantization level of transform coefficients corresponding to one component of the video data.
[0139] See you again Figure 6The quantized transform coefficients are output to the inverse quantization / transform processing unit 208. The inverse quantization / transform processing unit 208 can be configured to apply inverse quantization and inverse transform to generate reconstructed residual data. Figure 6 As shown, at summer 210, reconstructed residual data can be added to the predicted video block. This allows for the reconstruction of the encoded video block, which can then be used to evaluate the coding quality of a given prediction, transform, and / or quantization. The video encoder 200 can be configured to perform multiple coding rounds (e.g., coding while changing one or more of the prediction, transform, and quantization parameters). The rate-distortion or other system parameters of the bitstream can be optimized based on the evaluation of the reconstructed video block. Furthermore, the reconstructed video block can be stored and used as a reference for predicting subsequent blocks.
[0140] As described above, intra-frame prediction can be used to encode video blocks. The intra-frame prediction processing unit 212 can be configured to select an intra-frame prediction mode for the video block to be encoded. The intra-frame prediction processing unit 212 can be configured to evaluate frames and / or regions thereof and determine the intra-frame prediction mode to be used for encoding the current block. Figure 6 As shown, the intra-frame prediction processing unit 212 outputs intra-frame prediction data (e.g., syntax elements) to the entropy coding unit 218 and the transform coefficient generator 204. As mentioned above, the transform performed on the residual data can depend on the mode. As mentioned above, possible intra-frame prediction modes can include planar prediction mode, DC prediction mode, and angle prediction mode. Furthermore, in some examples, predictions for the chrominance components can be inferred from intra-frame predictions used for the luma prediction mode. The inter-frame prediction processing unit 214 can be configured to perform inter-frame prediction coding for the current video block. The inter-frame prediction processing unit 214 can be configured to receive a source video block and calculate the motion vector of the PU of the video block. The motion vector can indicate the displacement of the PU (or similar coding structure) of the video block within the current video frame relative to the prediction block within a reference frame. Inter-frame prediction coding can use one or more reference pictures. Furthermore, motion prediction can be unidirectional prediction (using one motion vector) or bidirectional prediction (using two motion vectors). Inter-frame prediction processing unit 214 can be configured to select prediction blocks by calculating pixel differences determined by, for example, sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. As described above, motion vectors can be determined and specified based on motion vector prediction. As described above, inter-frame prediction processing unit 214 can be configured to perform motion vector prediction. Inter-frame prediction processing unit 214 can be configured to generate prediction blocks using motion prediction data. For example, inter-frame prediction processing unit 214 can locate prediction video blocks within a frame buffer ( Figure 6(Not shown in the image). It should be noted that the inter-frame prediction processing unit 214 can be further configured to apply one or more interpolation filters to the reconstructed residual block to compute sub-integer pixel values for motion estimation. The inter-frame prediction processing unit 214 can output the motion prediction data of the computed motion vectors to the entropy coding unit 218. Figure 6 As shown, the inter-frame prediction processing unit 214 can receive reconstructed video blocks via the filter unit 216. The entropy coding unit 218 receives quantized transform coefficients and prediction syntax data (i.e., intra-frame prediction data, motion prediction data, and QP data, etc.). It should be noted that in some examples, the coefficient quantization unit 206 can perform a scan of the matrix including the quantized transform coefficients before outputting the coefficients to the entropy coding unit 218. In other examples, the entropy coding unit 218 can perform a scan. The entropy coding unit 218 can be configured to perform entropy coding according to one or more of the techniques described herein. The entropy coding unit 218 can be configured to output a compatible bitstream (i.e., a bitstream from which the video decoder can receive and reproduce video data).
[0141] Refer again Figure 6 Filter unit 216 can be configured to perform deblocking filtering, sample adaptive offset (SAO) filtering, and / or ALF filtering as described above. Furthermore, filter unit 216 can be configured to perform one or more techniques described herein for reducing reconstruction error based on cross-component correlation. As described above, for the ALF filter in JEM, each component uses a set of sample values from the corresponding component as input and derives output sample values in a manner independent of other components. Filtering independently per component may not be ideal because there may be correlations between the components and / or channels of the video data that can be used to minimize reconstruction error. For example, refer to... Figure 8 , Figure 8 An example is shown, illustrating an 8×8 luma source block and its corresponding 4×4 chroma source block (i.e., according to a 4:2:0 sampling format), along with the corresponding reconstruction block and reconstruction error. Figure 8 As shown, these two source blocks include edges around the diagonal, which would be typical in the case of textures, shape edges, etc. However, for reconstructing the chroma component, fidelity is lost compared to the luminance component (e.g., due to high levels of quantization, etc.), and edges are not recovered.
[0142] According to the techniques described herein, filter units can be configured to predict and / or refine information in the first color channel and / or components from information in the second color channel and / or components. This can provide improved coding efficiency for the first color channel and / or components because the fidelity of the color channel and / or components is increased by a small number of bits. Figure 7An example is shown of a cross-component filter unit that can be configured to encode video data according to one or more techniques of this disclosure. Figure 7 As shown, the cross-component filter unit 300 includes a filter determination unit 302 and a sample modification unit 304. It should be noted that the cross-component filter unit 300 illustrates an example of a cross-component filter unit that can exist in a video encoder. A corresponding example of a cross-component filter unit that can exist in a video decoder is described in more detail below. Figure 7 As shown, the filter determination unit 302 and the sample modification unit 304 can receive encoding parameter information (e.g., intra-frame prediction mode) available when the current block is encoded / decoded, and as... Figure 7 As shown, the video block data at the video encoder may include: cross-component source block; cross-component reconstruction block; cross-component reconstruction error; current component source block; current component reconstruction block; and current component reconstruction error. That is, referring to... Figure 8 The example shown illustrates how filtering is applied to a chroma reconstruction block. Figure 8 All information is available at the cross-component filter unit 300. Therefore, the filter determination unit 302 can derive the filters to be used on the chroma reconstruction block based on the video data, and the sample modification unit 304 can perform filtering according to the derived filters. Figure 7 As shown, the sample modification unit 304 can output the modified reconstruction block to the reference image buffer (i.e., as a loop filter) and output the modified reconstruction block to an output terminal (e.g., a display). Furthermore, as... Figure 7 As shown, the filter determination unit 302 can output filter data. That is, the filter data specifying the derived filter can be signaled to the video decoder. Examples of this signaling are described in further detail below. It should be noted that, relative to... Figure 8 There are several ways to reduce reconstruction errors at the video encoder, such as reducing quantization and / or performing improved prediction techniques. Furthermore, in some cases, the video encoder can directly signal the reconstruction error. However, cross-component filtering according to the techniques described herein provides a way to signal a relatively small amount of information while reducing reconstruction errors. That is, for example, the cross-component filtering techniques described herein can provide a way to reduce reconstruction errors at the video decoder while being more efficient than other techniques used to reduce reconstruction errors. For example, signaling filter data may require fewer bits compared to signaling higher-fidelity residual information for the signaling components.
[0143] Therefore, the cross-component filter unit 300 can operate by taking a first color component and one or more second color components as inputs and providing an enhanced first color component as output. It should be noted that although the examples described herein are relative to the luminance, Cb, and Cr components, the techniques described herein are generally applicable to other video formats (e.g., RGB) and other types of video information such as infrared, depth, parallax, or other features.
[0144] The following formula provides an example of a filter model that takes sample values from multiple components as input and outputs filtered sample values f. i (x,y), therefore, in one example, the cross component filter unit 300 can implement the filtering process based on this formula.
[0145]
[0146] in,
[0147] f i (x,y)
[0148] It is the output of component i at sample position (x, y);
[0149] S i,0 S i 1 ; and S i,2 Defines the position of a set of sample values relative to the origin in the corresponding sub-0, 1, 2 sub-0;
[0150] g(x,y, i, 0) and h(x,y, i, 0), g(x,y, i, 1) and h(x,y, i, 1), g(x,y, i, 2) and h(x,y, i, 2): determine the supported origin based on x, y, i and the input components. The functions g() and h() can also depend on the chroma format, chroma position type, color gamut, and filter shape;
[0151] c0(x0,y0), c1(x0,y0), and c2(x0,y0): are the filter coefficient values for the support region of each component;
[0152] I0, I1, and I2: are the input sample values from each component; and
[0153] I i (x, y): is the sample value of component i at sample position (x, y) before filtering.
[0154] Therefore, according to the technique described herein, the cross-component filter unit 300 can be configured to reduce the reconstruction error of the current component by adding refinement to the reconstructed sample values of the current component based on a derived filtering function that takes the reconstructed sample values of other components as input. In one example, the reconstructed sample values of other components used as input may be referred to as filter support. Figures 9A to 9F This is a conceptual diagram illustrating a sample example of a support sample that can be used for cross-component filtering according to one or more techniques of this disclosure. Figures 9A to 9F In the examples shown, for each of the 4:2:0 sample format chroma position types provided in JVET-N1001 and JVET-O2001, the luminance support samples for the chroma samples to be filtered are shown. That is, 5×5, 5×6, 6×5, and / or 6×6 support samples can be used. It should be noted that in Figures 9A to 9F In the example, luminance support is defined as symmetric about chroma sample values. It should be noted that in other examples, luminance support may undergo a phase shift before being input to a filter stage independent of the chroma position type, which specifically depends on the chroma position type. As described further below, for each support sample, filter coefficients can be determined and signaled. In one example, according to the techniques described herein, the relative position of support for each chroma sample included in a video block can be based on the sample format. For example, in one example, according to the techniques described herein, for a 4:2:0 sample format, when the chroma position (x... C , y C The chromaticity sample at position (x) corresponds to the origin at the luminance position (x). L ,y L When support is provided at location (x), it is used for the chroma location (x) in that video block. C +m, y C The origin of the chromaticity sample at (+n) can be found at the luminance position (x). C +2m, y C +2n); for the 4:2:2 sample format, when the chroma position (x) in the video block... C , y C The chromaticity sample at position (x) corresponds to the origin at the luminance position (x). L ,y L When support is provided at location (x), it is used for the chroma location (x) in that video block. C +m, y C The origin of the chromaticity sample at (+n) can be found at the luminance position (x). C +2m, y C +n); and for the 4:4:4 sample format, when the chroma position (x) in the video block is... C , y C The chromaticity sample at position (x) corresponds to the origin at the luminance position (x).L ,y L When support is provided at location (x), it is used for the chroma location (x) in that video block. C +m, y C The origin of the chromaticity sample at (+n) can be found at the luminance position (x). C +m, y C At +n). It should be noted that in this example, the offset of the chroma sample position corresponds to the offset of the luminance position of the supported origin, where the ratio between these two offsets is based on the chroma format.
[0155] In one example, according to the techniques described herein, the application of cross-component filtering can be based on the properties of samples included in the filter support region. For example, in one example, luminance sample values in the support region can be analyzed, and the application of cross-component filtering can be determined based on this analysis. For example, in one example, the variance and / or bias of the samples in the support region can be calculated, and if the variance and / or bias have certain characteristics, such as the region being smooth (i.e., the variance is less than a threshold), cross-component filtering may not be applied to that region. In one example, the selection of cross-component filters (including whether to apply the filter, when to apply the filter, and which filter to apply) can be based on the luminance classification filter index of the luminance sample corresponding to the evaluated chrominance sample. In one example, the classification filter index for the luminance sample can be derived as described in JVET-O2001. In one example, cross-component filtering may not be applied when it is determined that the luminance classification filter index is in a subset of the luminance classification filter index. As described further in detail below, the values of local region control tags and / or syntax elements can be used to indicate / determine whether cross-component filtering is applied to a region, and if so, which cross-component filter to apply. In one example, the application of cross-component filtering can be based on the attributes of the samples included in the filter support region and / or the values of local region control markers and / or syntax elements. That is, for example, how to analyze luminance support samples can be based on local region control markers and / or syntax elements (e.g., if the marker == 0, then calculate / evaluate the variance; otherwise, calculate / evaluate the luminance classification filter index). Furthermore, in one example, filter selection is based on the values of syntax elements and the attributes of the luminance support samples. For example, a syntax element value of 0 can indicate that no cross-component filtering is applied to the region; a syntax element value of 1 and the luminance support variance greater than a threshold can indicate that a filter with a first set of filter coefficients is applied; a syntax element value of 1 and the luminance support variance not greater than a threshold can indicate that a filter with a second set of filter coefficients is applied; a syntax element value of 2 and the luminance support variance greater than a threshold can indicate that a filter with a third set of filter coefficients is applied; a syntax element value of 2 and the luminance support variance not greater than a threshold can indicate that a filter with a fourth set of filter coefficients is applied, and so on.
[0156] The appendix to this document provides examples of datasets corresponding to specific implementations of the cross-component filters described herein. Specifically, in the appendix, the dataset `orgBlock` represents the sample values of the original 32×32U component block; the dataset `preFilteringBlock` represents the sample values of the reconstructed 32×32U component block; the dataset `orgError` represents the reconstruction error between the original and reconstructed 32×32U component blocks; the dataset `bestSupportY` represents the sample values of a 67×68Y component block that provides filter support for the filtered reconstructed 32×32U component block; the dataset `bestSupportU` represents the sample values of a 36×36U component block that provides support for the filtered reconstructed 32×32U component block; and the dataset `bestSupportV` represents the sample values of a 36×36UV component block. The component blocks provide support for the filtered reconstructed 32×32U component blocks; the dataset coeffY represents the filter coefficients in a 5×6 filter used for sample values of the 67×68Y component support block; the dataset coeffU represents the filter coefficients in a 5×5 filter used for sample values of the 36×36U component support block; the dataset coeffU represents the filter coefficients in a 5×5 filter used for sample values of the 36×36V component support block; the dataset bestOutput represents the sample values of the filtered reconstructed 32×32U component blocks; the dataset bestError represents the error between the original 32×32U component blocks and the filtered reconstructed 32×32U component blocks; the dataset signedimprovement equals Abs(orgError) - Abs(bestError) and represents the change in reconstruction error caused by filtering; and the dataset positive improve represents the reconstructed sample values whose reconstruction error is reduced due to filtering. Therefore, according to the technique presented in this paper, the reconstruction error of one or more or most samples can be reduced by applying a cross-component filter. It should be noted that, for a specific type of video content, the amount of reconstruction error improved according to mathematical relationships can have different results depending on how the perceived visual quality of the video is improved. That is, for example, a relatively small signedimprovement value can lead to a relatively significant improvement in visual quality.
[0157] As described above, the cross-component filter unit 300 typically operates by taking a first color component and one or more second color components as inputs, and provides an enhanced first color component as output. That is, the filtering process performed by the cross-component filter unit 300 can take input luminance sample values as input luminance sample values, which can be used to predict the difference between the original corresponding chrominance sample values and output refined chrominance sample values based on this prediction. (Refer to again...) Figure 8The example shown, Figure 10 This is a conceptual diagram illustrating an example of using cross-component filtering to reduce reconstruction error according to one or more techniques of this disclosure. Figure 10 An example is provided where reconstruction error is reduced by: taking the mean of the support samples, and if the mean is greater than 90, dividing the mean by 10 and adding it to the reconstructed samples; and if the mean is not greater than 90, subtracting the mean divided by 10 from the reconstructed samples. That is, in this example, the prediction filter is typically described as: if the support mean is greater than threshold1, adding weight1 multiplied by the support mean; otherwise, adding weight2 multiplied by the support mean. Figure 10 As shown in the example, post-filtered chroma reconstruction error reduces the overall reconstruction error. Therefore, according to the techniques described in this paper, cross-component filtering can reduce reconstruction error by using cross-component filters defined according to logic functions, thresholds, weights, etc.
[0158] As described above, JVET-N1001 and JVET-O2001 include a deblocking filter, a SAO filter, and an ALF filter. The cross-component filter technique described herein can be implemented at various points in the filter chain. That is, for example, it can be implemented at various stages of a loop filter. Figures 11A to 11D This is a block diagram illustrating an example of a cross-component filter unit that can be configured to reduce reconstruction error according to one or more techniques of this disclosure. Figures 11A to 11DIn this context, luminance SAO filter unit 402 represents a filter unit configured to perform SAO filtering (e.g., SAO filtering provided in JVET-N1001 or JVET-O2001) on luminance sample values Y; Cb SAO filter unit 404 represents a filter unit configured to perform SAO filtering (e.g., SAO filtering provided in JVET-N1001 or JVET-O2001) on chrominance sample values Cb; Cb SAO filter unit 406 represents a filter unit configured to perform SAO filtering (e.g., SAO filtering provided in JVET-N1001 or JVET-O2001) on Cb sample values; luminance ALF filter unit 408 represents a filter unit configured to perform ALF filtering (e.g., ALF filtering provided in JVET-N1001 or JVET-O2001) on luminance sample values Y; chrominance ALF filter unit 410 represents a filter unit configured to perform ALF filtering (e.g., ALF filtering provided in JVET-N1001 or JVET-O2001) on chrominance sample values. The units are as follows: Luminance deblocking filter unit 416 represents a filtering unit configured to perform deblocking filtering on luminance sample values Y (e.g., deblocking filtering provided in JVET-N1001 or JVET-O2001); Cb deblocking filter unit 418 represents a filtering unit configured to perform deblocking filtering on Cb sample values (e.g., deblocking filtering provided in JVET-N1001 or JVET-O2001); and Cr deblocking filter unit 420 represents a filtering unit configured to perform deblocking filtering on Cr sample values (e.g., deblocking filtering provided in JVET-N1001 or JVET-O2001). Furthermore, in... Figures 11A to 11D In this document, Cb cross-component filter unit 412 represents an example configured to generate a Cb-refined ΔCb cross-component filter according to one or more of the techniques described herein; and Cr cross-component filter unit 414 represents an example configured to generate a Cr-refined ΔCr cross-component filter according to one or more of the techniques described herein. Therefore, as... Figures 11A to 11D As shown, the cross-component filtering according to the technique described in this paper can be applied to various points in the filter chain. That is, cross-component filtering input can be received at various points in the filter chain, and cross-component filtering refinement can be output at various points in the filter chain. It should be noted that, in Figure 11B In the example shown, the input to the luminance demodulation is used as the input to the filter, which has the advantage of reducing the requirements for the line buffer. It should be noted that in... Figure 11C In the example shown, the output of the luminance deblocking is used as the input to the filter, which has the advantage of reducing line buffer requirements while slightly improving coding efficiency. It should be noted that in... Figure 11DIn the example shown, the output of ALF luminance is used as the input to the filter, which can have the advantage of improved coding efficiency.
[0159] Furthermore, the cross-component filter technique described herein may also include performing clipping operations at various points in the filter chain. That is, for example, performing them at each stage of the loop filter. Figures 12A to 12B This is a block diagram illustrating an example of a cross-component filter unit that can be configured to reduce reconstruction error according to one or more techniques of this disclosure. Figures 12A to 12B In the text above, the components with the same number are relative to... Figures 11A to 11B As described, clipping units 422A to 422D can be configured to perform a clipping function based on the output bit depth of the corresponding component, such as Clip3(0, 2). BitDepthC - 1, *). It should be noted that clipping units 422A to 422D can be selectively enabled based on whether a specific type of filtering is performed.
[0160] Figures 13A to 13B This is a block diagram illustrating an example of a cross-component filter unit that can be configured to reduce reconstruction error according to one or more techniques of this disclosure. Figures 13A to 13B This further illustrates that cross-component filtering according to the technique presented herein can be applied to various points in the filter chain. Figures 13A to 13C In this context, the commonly numbered elements are as described above.
[0161] Furthermore, it should be noted that in some cases, there may be more than three video data components, such as YUV + depth. The cross-component filtering technique described in this article is generally applicable to these situations. In some cases, preprocessing of the input sample values from each component can be performed before the filtering operation. For example, the input sample values can be clipped. Moreover, in one example, the clipping range can vary for each coefficient, and this can be signaled in the bitstream. It should be noted that in some examples, the following formula provides options for preprocessing the input sample values:
[0162]
[0163] In addition, another option for preprocessing input sample values is as follows:
[0164]
[0165] in,
[0166] I' j (u,v): Sample value at position (u,v) before processing.
[0167] derivedValue: is a value derived from a subset of values in I', (*,*), for example, (a) at the origin. j (g(x,y,i,j) + 0,h(x,y,i,j) + 0), the area around the origin is
[0168] a: is the value received in the bitstream / the value inferred from the data received in the bitstream / from I' j Values derived from a subset of values in (*,*), and
[0169] b: is the value received in the bitstream / the value inferred from the data received in the bitstream / from I' j Values derived from a subset of values in (*,*)
[0170] In one example, b can be derived from a (e.g., b = -a) to reduce the amount of signaling required.
[0171] Furthermore, in one example, the generalization of the input used in the cross-component filter operation can be as follows:
[0172]
[0173] Where Gn() is used to combine sample values from the components and obtain the corresponding value with index (x). cj ,y cj A function that derives the value of each coefficient of ). Function G i () can depend on the chroma format, chroma location type, color gamut, and filter shape.
[0174] In one example, cross-component filtering can be performed based on the following: defining a support region for luminance; upsampling the 2X chrominance components for 4:2:0 as input; subtracting the derived values (e.g., 512 for 10-bit chrominance, or local averages) from the support for the corresponding chrominance components; then taking the sampled product of the luminance and chrominance sample values corresponding to the defined support region; and using this product as one of the inputs to the filtering operation.
[0175] Furthermore, it should be noted that in some examples, the cross-component filtering technique described herein can be performed on either the prediction or the residual. In one example, if domain coding is used instead of progressive coding, then for luminance support samples: in one example, sample values from one of the corresponding luminance domains can be used, and in another example, sample values from both luminance domains can be used.
[0176] As described above, for each supporting sample, filter coefficients can be determined and signaled to those filter coefficients. That is, for example, 5×5, 5×6, 6×6 and / or 6×6 filter coefficients can be signaled. Figures 14A to 14C An example of signaling the filter coefficients used for a 5×5 filter is shown. Figures 14D to 14F An example of signaling the filter coefficients for a 5×6 (and similarly, 6×5) filter is shown. Figures 15A to 15D An example of signaling the filter coefficients used for a 6×6 filter is shown. Figures 14A to 15D In each figure, the corresponding filter coefficients of the filter are given by C. N Instructions. Therefore, in the same C When the same filter is provided at multiple locations, the filter coefficients are identical, i.e., shared. This reduces the number of filter coefficients that need to signal the filter. For example, in Figure 14D In this process, signals are sent to 14 filter coefficients for 18 support locations.
[0177] In one example, it might be desirable to limit the number of line buffers within an architecture that processes samples per CTU. That is, for example, a virtual line boundary provides the location of samples above the horizontal VB that can be processed before a lower CTU becomes available, but samples below the horizontal VB cannot be processed until a lower CTU becomes available. JVET-N1001 and JVET-O2001 define horizontal virtual line boundaries (VBs) for the luma ALF and luma SAO. According to the techniques described herein, this VB can be reused in the luma input-chroma output filter defined herein. Furthermore, the vertical VB can be reused in the vertical luma input-chroma output filter defined herein, and / or a subset of the VB can be reused in the luma input-chroma output filter defined herein. Additionally, two cases are defined for which supporting samples in the luma component can be derived / modified: when a predetermined luma sample (corresponding to, for example, a chroma sample decoded based on chroma position type) is above the VB and supports crossing the VB; and when a predetermined luma sample (corresponding to, for example, a chroma sample decoded based on chroma position type) is below the VB and supports crossing the VB. In one example, the predetermined sample is in the case of... Figure 14D The sample at the location of coefficient C6, which supports 5×6 brightness. Figures 16A to 16D This is a conceptual diagram illustrating an example of a virtual line buffer that can be used for cross-component filtering according to one or more techniques of this disclosure. Figure 16A In this context, the sample below the horizontal VB is obtained by copying the sample above and closest to the virtual line boundary and within the same column. Figure 16BIn this example, samples below the horizontal virtual boundary (VB) are obtained by copying samples below and closest to the virtual boundary, and within the same column. In one example, the luminance VB consists of four samples from the horizontal CTU boundary. In another example, each CTU, SAO, and ALF can process samples to the left of the vertical VB before the right CTU enters, but cannot process samples to the right of the vertical VB before the right CTU becomes available. Exemplary modifications when supporting crossing the vertical virtual boundary (VB) are described in... Figures 16C to 16D As shown in [the image]. Figure 16C In the example, the sample to the right of the vertical VB is obtained by copying the sample to the left, closest to the virtual line boundary, and within the same row. Figure 16D In this example, the sample to the left of the vertical VB is obtained by copying the sample to the right, closest to the virtual line boundary, and in the same row. In one example, the luminance VB is four samples from the vertical CTU boundary. In one example, according to the technique of this invention, a vertical axis and a horizontal axis passing through the center of the support area can be considered relative to the sample generated for the horizontal VB. The copied sample can be obtained by copying samples that are equidistant from the vertical axis but on the opposite side of the vertical axis in a column. In one example, the copied sample can be copied from a row that is equidistant from the horizontal axis but on the opposite side. In one example, according to the technique of this invention, a vertical axis and a horizontal axis passing through the center of the support area can be considered relative to the sample generated for the vertical VB. The copied sample can be obtained by copying samples that are equidistant from the vertical axis but on the opposite side of the horizontal axis in a row. In one example, the copied sample can be copied from a column that is equidistant from the vertical axis but on the opposite side. In one example, according to the technique of this invention, the sample can be obtained by symmetrical filling relative to the sample generated for the VB. In other words, samples can be copied from the same column at the same sample distance from the VB relative to the horizontal VB and relative to the vertical VB, and samples can be copied from the same row at the same sample distance from the VB. That is, the sample values are mirrored with respect to the VB. It should be noted that Section 8.8.5.2 of JVET-O2001, “Coded Tree Block Filtering Procedure for Luminance Samples,” provides a padding scheme for luminance samples across virtual boundaries for use relative to the ALF procedure. In one example, a similar padding scheme can be used for cross-component filtering according to the techniques described herein.
[0178] In one example, according to the techniques described herein, cross-component filtering involves scaling the output of the cross-component filter before adding the output to the corresponding chroma ALF output.
[0179] That is, scaling operations can be used to convert filter coefficients to integers, for example as follows:
[0180]
[0181] In one example, factor = 2 BitDepthC In another example, factor = 2 (BitDepthC -1) In another example, factor = 2 (8-l)
[0182] In one example, the scaling factor can be used to adjust the output of the cross-component filter as follows:
[0183]
[0184] It should be noted that if factor = 2 X This corresponds to right-shifted integer rounding (f i (x,y) +2 (x-1) )>>x
[0185] As described above, filter data specifying the exported filters can be signaled to the video decoder. In one example, signaling the filter data can have three main aspects: turning the filter on / off; local control of the tool, e.g., enabling the tool in some spatial regions but not others; and signaling for a specific filter. In one example, the parameter set (e.g., the sequence parameter set) can conditionally include flags to enable / disable the filter. In one example, this flag can indicate whether one or more filters, such as an ALF filter and a cross-component filter, are enabled.
[0186] In one example, slice-level signaling of filter coefficients can be used. Tables 3 through 5 show examples of syntax that may be included in the slice header used to signal filter coefficients. It should be noted that in Table 3, for the syntax elements `lice_cross_component_alf_cb_log2_control_size_minus4` and `slice_cross_component_alf_cr_log2_control_size_minus4`, in one example, truncated unary codes can be used instead of `ue(v)`, where the maximum value of the truncated unary is based on the valid range.
[0187]
[0188] Table 3
[0189]
[0190] Table 4
[0191]
[0192] Table 5
[0193] For Tables 3 through 5, in one example, the semantics could be based on the following:
[0194] A slice_cross_component_alf_cb_enabled_flag value of 0 indicates that the cross-component adaptive loop filter is not applied to the Cb color component. A slice_cross_component_alf_cb_enabled_flag value of 1 indicates that the cross-component adaptive loop filter is applied to the Cb color component.
[0195] A slice_cross_component_alf_cr_enabled_flag value of 0 indicates that the cross-component adaptive loop filter is not applied to the Cr color component. A slice_cross_component_alf_cb_enabled_flag value of 1 indicates that the cross-component adaptive loop filter is applied to the Cr color component.
[0196] The following values specify the size of the square blocks in units of the number of samples: slice_cross_component_alf_cb_log2_control_size_minus4
[0197] AlfCCSamplesCbW=AlfCCSamplesCbH= 2( slice _cross_component_alf_cb_log2_control_size_minus4+4 )
[0198] The slice_cross_component_alf_cb_log2_control_size_minus4 should be in the range of 0 to 3 (inclusive).
[0199] The slice_cross_component_alf_cr_log2_control size minus4 specifies the value of the square block size in units of the number of samples:
[0200] AlfCCSamplesCrW=AlfCCSamplesCrH= 2( slice _cross_component_alf_cr_log2_control_size_minus4+4 )
[0201] slice_cross_component_alf_cr_log2_control_size_minus4 should be in the range of 0 to 3 (inclusive).
[0202] It should be noted that, in one example, the orientation of slice_cross_component_alf_cb_log2_control_size_minus4 and / or slice_cross_component_alf_cr_log2_control_size_minus4 can be defined in a parameter set such as SPS.
[0203] In one example, the minusX encoding depends on the minimum value defined for the valid range; for example, if the valid range is 2 to 5, then X=2.
[0204] In one example, slice_cross_component_alf_cb_log2_control_size_minus4 and / or slice_cross_component_alf_cr_log2_control_size_minus4 can be signaled in SPS or derived from the CTU size (e.g., the same as the CTU size in the chroma sample).
[0205] `alf_cross_component_cb_filter_signal_flag` equal to 1 specifies that the cross component Cb filter bank is signaled. `alf_cross_component_cb_filter_signal_flag` equal to 0 specifies that the cross component Cb filter bank is not signaled.
[0206] The increment of 1 in `alf_cross_component_cb_min_eg_order_minus1` specifies the minimum order of the exponential Golomb code used for signaling the Cb filter coefficients of the cross-component component. The value of `alf_cross_component_cb_min_eg_order_minus1` should be in the range of 0 to 9 (inclusive). It should be noted that this range may vary in some examples.
[0207] `alf_cross_component_cb_eg_order_increase_flag[i]` equal to 1 specifies that the minimum order of the exponential Golomb code used for the cross-component Cb filter coefficient signaling is incremented by 1. `alf_cross_component_cb_eg_order_increase_flag[i]` equal to 0 specifies that the minimum order of the exponential Golomb code used for the cross-component Cb filter coefficient signaling is not incremented by 1.
[0208] The exponent of the Golomb code used to decode the value of alf_cross_component_cb_coeff_delta_abs[j], expGoOrderCb[i], is derived as follows:
[0209] expGoOrderCb[ i ] = (i = = 0 ? alf_cross_component_cb_min_eg_order_minus1 + 1 :
[0210] expGoOrderCb[ i - 1 ]) + alf_cross_component_cb_eg_order_increase_flag[ i ]
[0211] alf_cross_component_cb_coeff_delta_abs[j] specifies the absolute value of the j-th coefficient delta of the cross component Cb filter that is signaling the signal. If alf_luma_cross_component_cb_coeff_delta_abs[j] does not exist, it is assumed to be equal to 0.
[0212] The order k of the exponential Golomb binarization uek(v) is derived as follows:
[0213] golombOrderIdxCb[ ] = {0,2,2,2,1,2,2,2,2,2,2,1,2,1}
[0214] k = expGoOrderCb[ golombOrderIdxCb[ j ] ]
[0215] alf_cross_component_cb_coeff_sign[j] specifies the sign of the Cb filter coefficient for the j-th cross component as follows:
[0216] - If alf_cross_component_cb_coeff_sign[j] is equal to 0, then the corresponding chromaticity filter coefficient has a positive value.
[0217] Otherwise (alf_cross_component_cb_coeff_sign[j] equals 1), the corresponding chroma filter coefficients have negative values.
[0218] If alf_cross_component_cb_coeff_sign[j] does not exist, it is inferred that it is equal to 0.
[0219] The cross-component filter coefficients AlfCCCoeffcb with elements AlfCCCoeffCb[j], j = 0..13 are derived as follows:
[0220] AlfCCCoeffCb[j] = alf_cross_component_cb_coeff_abs[j]*
[0221] (1 - 2 * alf_cross_component_cb_coeff_sign[ j ])
[0222] The bitstream compliance requirement is that the value of AlfCCCoeffcb[j], j = 0..13, should be in the range of -2. 10 - 1 to 2 10 -1 is the range (inclusive). It should be noted that in some examples, this range may depend on the bit depth of the luminance / chrominance or a subset thereof.
[0223] `alf_cross_component_cr_filter_signal_flag` equal to 1 specifies that the cross component (Cr) filter bank is signaled. `alf_cross_component_cb_filter_signal_flag` equal to 0 specifies that the cross component (Cr) filter bank is not signaled.
[0224] The increment of 1 in `alf_cross_component_cr_min_eg_order_minus1` specifies the minimum order of the exponential Golomb code used for signaling the coefficients of the cross-component (Cr) filter. The value of `alf_cross_component_cb_min_eg_order_minus1` should be in the range of 0 to 9 (inclusive). It should be noted that this range may vary in some examples.
[0225] alf_cross_component_cr_eg_order_increase_flag[i] equal to 1 specifies that the minimum order of the exponential Golomb code used for signaling of the cross component Cr filter coefficients is incremented by 1.
[0226] The value of alf_cross_component_cr_eg_order_increase_flag[i] equal to 0 indicates that the minimum order of the exponential Golomb code used for signaling the coefficients of the cross component Cr filter is not incremented by 1.
[0227] The exponent of the Golomb code used to decode the value of alf_cross_component_cb_coeff_delta_abs[j], expGoOrderCr[i], is derived as follows:
[0228] expGoOrderCr[ i ] = (i = = 0 ? alf_cross_component_cr_min_eg_order_minus1 + 1 :
[0229] expGoOrderCrf i - 1 ]) + alf_cross_component_cr_eg_order_increase_flag[ i ]
[0230] alf_cross_component_cr_coeff_delta_abs[j] specifies the absolute value of the j-th coefficient delta of the cross component Cr filter that is signaling the signal. If alf_luma_cross component_cr_coeff_delta_abs[j] does not exist, it is assumed to be equal to 0.
[0231] The order k of the exponential Golomb binarization uek(v) is derived as follows:
[0232] golombOrderIdxCr[ ] = {0,1,2,1,0,1,2,2,2,2,2,1,2,1}
[0233] k = expGoOrderCr[ golombOrderIdxCr[ j ] ]
[0234] alf_cross_component_cr_coeff_sign[j] specifies the sign of the filter coefficient for the j-th cross component Cr as follows:
[0235] - If alf_cross_component_cr_coeff_sign[j] is equal to 0, then the corresponding chromaticity filter coefficient has a positive value.
[0236] Otherwise (alf_cross_component_cr_coeff_sign[j] equals 1), the corresponding chroma filter coefficients have negative values.
[0237] If alf_cross_component_cr_coeff_sign[j] does not exist, it is inferred that it is equal to 0.
[0238] Having element AlfCCCoeff Cr [j], j = 0..13, cross-component Cr filter coefficients AlfCCCoeff Cr Export as follows:
[0239] AlfCCCoeff Cr [j] = alf_cross_component_cr_coeff_abs[j]*
[0240] (1 - 2 * alf_cross_component_cr_coeff_sign[ j ])
[0241] Bitstream compliance requirements AlfCCCoeff Cr [j], the value of j = 0..13 should be in the range of -2. 10 - 1 to 2 10 -1 is the range (inclusive). It should be noted that in some examples, this range may depend on the bit depth of the luminance / chrominance or a subset thereof.
[0242] In another example, one or more pointers to the APS containing the corresponding filter coefficient data can be sent in the slice header. Tables 6 and 7 show examples of syntax that can be included in the slice header used to signal the filter coefficients, according to this example.
[0243]
[0244] Table 6
[0245]
[0246]
[0247] Table 7
[0248] For Tables 6 and 7, in one example, the semantics could be based on the following:
[0249] A slice_cross_component_alf_cb_enabled_flag value of 0 indicates that the cross component Cb filter is not applied to the Cb color component. A slice_cross_component_alf cb_enabled flag value of 1 indicates that the cross component adaptive loop filter is applied to the Cb color component.
[0250] A slice_cross_component_alf_cr_enabled_flag value of 0 indicates that the cross-component Cr filter is not applied to the Cr color component. A slice_cross_component_alf_cb_enabled_flag value of 1 indicates that the cross-component adaptive loop filter is applied to the Cr color component.
[0251] slice_cross_component_alf_cb_aps_id specifies the adaptation_parameter_set_id indexed by the Cb color component of the slice. When slice_cross_component_alf_cb_aps_id does not exist, it is inferred to be equal to slice_alf_aps_id_luma[0]. The Temporalld of the ALF APS NAL cell with an adaptation_parameter_set_id equal to slice_cross_component_alf_cb_aps_id should be less than or equal to the Temporalld of the encoded slice NAL cell.
[0252] slice_cross_component_alf_cr_aps_id specifies the adaptation_parameter_set_id indexed by the Cr color component of the slice. When slice_cross_component_alf_cr_aps_id does not exist, it is inferred to be equal to slice_alf_aps_id_luma[0]. The Temporalld of the ALF APS NAL cell with an adaptation_parameter_set_id equal to slice_cross_component_alf_cr_aps_id should be less than or equal to the Temporalld of the encoded slice NAL cell.
[0253] The following values specify the size of the square blocks in units of the number of samples: slice_cross_component_alf_cb_log2_control_size_minus4
[0254] AlfCCSamplesCbW = AlfCCSamplesCbH = 2( slice _cross_component_alf_cb_log2_control_size_minus4+4 )
[0255] The slice_cross_component_alf_cb_log2_control_size_minus4 should be in the range of 0 to 3 (inclusive).
[0256] The following values specify the size of the square blocks in units of the number of samples: slice_cross_component_alf_cr_log2_control_size_minus4
[0257] AlfCCSamplesCrW = AlfCCSamplesCrH = 2( slice _cross_component_alf_cb_log2_control_size_minus4+4 )
[0258] slice_cross_component_alf_cr_log2_control_size_minus4 should be in the range of 0 to 3 (inclusive).
[0259] `alf_luma_filter_signal_flag` equal to 1 specifies that a signal is sent to notify the luminance filter bank. `alf_luma_filter_signal_flag` equal to 0 specifies that no signal is sent to notify the luminance filter bank. If `alf_luma_filter_signal_flag` does not exist, it is assumed to be equal to 0.
[0260] `alf_chroma_filter_signal_flag` equal to 1 specifies that a signal is sent to notify the chroma filter. `alf_chroma_filter_signal_flag` equal to 0 specifies that no signal is sent to notify the chroma filter. If `alf_chroma_filter_signal_flag` does not exist, it is inferred that it is equal to 0.
[0261] `alf_cross_component_cb_filter_signal_flag` equal to 1 indicates that the cross component Cb filter bank is signaled. `alf_cross_component_cb_filter_signal_flag` equal to 0 indicates that the cross component Cb filter bank is not signaled. If `alf_cross_component_cb_filter_signal_flag` does not exist, it is inferred that it is equal to 0.
[0262] `alf_cross_component_cr_filter_signal_flag` equal to 1 specifies that the cross-component (Cr) filter bank should be signaled. `alf_cross_component_cb_filter_signal_flag` equal to 0 specifies that the cross-component (Cr) filter bank should not be signaled. If `alf_cross_component_cr_filter_signal_flag` does not exist, it is assumed to be equal to 0.
[0263] The increment of 1 in `arf_cross_component_cb_min_eg_order_minus1` specifies the minimum order of the exponential Golomb code used for signaling the Cb filter coefficients of the cross-component component. The value of `arf_cross_component_cb_min_eg_order_minus1` should be in the range of 0 to 9 (inclusive). It should be noted that this range may vary in some examples.
[0264] The increment of 1 in `alf_cross_component_cr_min_eg_order_minus1` specifies the minimum order of the exponential Golomb code used for signaling the coefficients of the cross-component (Cr) filter. The value of `alf_cross_component_cb_min_eg_order_minus1` should be in the range of 0 to 9 (inclusive). It should be noted that this range may vary in some examples.
[0265] `alf_cross_component_cb_eg_order_increase_flag[i]` equal to 1 specifies that the minimum order of the exponential Golomb code used for the cross-component Cb filter coefficient signaling is incremented by 1. `alf_cross_component_cb_eg_order_increase_flag[i]` equal to 0 specifies that the minimum order of the exponential Golomb code used for the cross-component Cb filter coefficient signaling is not incremented by 1.
[0266] The exponent of the Golomb code used to decode the value of alf_cross_component_cb_coeff_delta_abs[j], expGoOrderCb[i], is derived as follows:
[0267] expGoOrderCb[ i ] = (i = = 0 ? alf_cross_component_cb_min_eg_order_minus1 + 1 :
[0268] expGoOrderCb[ i - 1 ]) + alf_cross_component_cb_eg_order_increase_flag[ i ]
[0269] `alf_cross_component_cr_eg_order_increase_flag[i]` equal to 1 specifies that the minimum order of the exponential Golomb code used for the signaling of the cross-component Cr filter coefficients is incremented by 1. `alf_cross_component_cr_eg_order_increase_flag[i]` equal to 0 specifies that the minimum order of the exponential Golomb code used for the signaling of the cross-component Cr filter coefficients is not incremented by 1.
[0270] The exponent of the Golomb code used to decode the value of alf_cross_component_cb_coeff_delta_abs[j], expGoOrderCr[i], is derived as follows:
[0271] expGoOrderCr[ i ] = (i = = 0 ? alf_cross_component_cr_min_eg_order_minus1 + 1 :
[0272] expGoOrderCr[ i - 1 ]) + alf_cross_component_cr_eg_order_increase_flag[ i ]
[0273] alf_cross_component_cb_coeff_delta_abs[j] specifies the absolute value of the j-th coefficient delta of the cross component Cb filter that is signaling the signal. If alf_luma_cross_component_cb_coeff_delta_abs[j] does not exist, it is assumed to be equal to 0.
[0274] The order k of the exponential Golomb binarization uek(v) is derived as follows:
[0275] golombOrderIdxCb[ ] = {0,2,2,2,1,2,2,2,2,2,2,1,2,1} [These can classify the coefficients into 3 categories, each using the same k-order exponent Golomb code]
[0276] k = expGoOrderCb[ golombOrderIdxCb[ j ] ]
[0277] alf_cross_component_cr_coeff_delta_abs[j] specifies the absolute value of the j-th coefficient delta of the cross component Cr filter that sends the signal. If alf_luma_cross_component_cr_coeff_delta_abs[j] does not exist, it is assumed to be equal to 0.
[0278] The order k of the exponential Golomb binarization uek(v) is derived as follows:
[0279] golombOrderIdxCr[ ] = {0,1,2,1,0,1,2,2,2,2,2,1,2,1 [These can classify the coefficients into 3 categories, each using the same k-order exponent Golomb code]}
[0280] k = expGoOrderCr[ golombOrderldxCr[ j ] ]
[0281] alf_cross_component_cb_coeff_sign[j] specifies the sign of the Cb filter coefficient for the j-th cross component as follows:
[0282] - If alf_cross_component_cb_coeff_sign[j] is equal to 0, then the corresponding cross component Cb filter coefficients have positive values.
[0283] Otherwise (alf_cross_component_cb_coeff_sign[j] equals 1), the corresponding cross component Cb filter coefficients have negative values.
[0284] If alf_cross_component_cb_coeff_sign[j] does not exist, it is inferred that it is equal to 0.
[0285] Having element AlfCCCoeff cb [adaptation_parameter_set__id][j], cross-component Cb filter coefficients AlfCCCoeff, j = 0..13 cb [adaptation_parameter_set_id] is exported as follows:
[0286] AlfCCCoeff cb[ adaptation_parameter_set_id ][ j ] = alf_cross_component_cb_coeff_abs[ j ] *
[0287] (1 - 2 * alf_cross_component_cb_coeff_sign[ j ])
[0288] Bitstream compliance requirements AlfCCCoeff cb [adaptation_parameter_set_id][j], where j = 0.. 13, should have a value in the range of -2. 10 - 1 to 2 10 -1 (inclusive).
[0289] alf_cross_component_cr_coeff sign[j] specifies the sign of the filter coefficient for the j-th cross component Cr as follows:
[0290] - If alf_cross_component_cr_coeff_sign[j] is equal to 0, then the corresponding cross component Cr filter coefficients have positive values.
[0291] Otherwise (alf_cross_component_cr_coeff_sign[j] equals 1), the corresponding cross component Cr filter coefficients have negative values.
[0292] If alf_cross_component_cr_coeff_sign[j] does not exist, it is inferred that it is equal to 0.
[0293] Having element AlfCCCoeff Cr [adaptation_parameter_set_id] [j], cross-component Cr filter coefficients AlfCCCoeff, j = 0..13 cr [adaptation_parameter_set_id] is exported as follows:
[0294] AlfCCCoeff Cr [ adaptation_parameter_set_id ] [ j ] = alf_cross_component_cr_coeff_abs[ j ] *
[0295] (1 - 2 * alf_cross_component_cr_coeff_sign[ j ])
[0296] Bitstream compliance requirements AlfCCCoeff Cr [adaptation_parameter_set_id][j], where j = 0.. 13, should have a value in the range of -2. 10 - 1 to 2 10 -1 (inclusive).
[0297] It should be pointed out that -2 10 - 1 to 2 10 The range of -1 can vary. It should be noted that in some examples, this range may depend on the bit depth of the luminance / chrominance or a subset thereof.
[0298] For Table 6, in one example, the alf_data() syntax structure provided in Table 8A can be used.
[0299]
[0300]
[0301] Table 8A
[0302] For Table 8A, in one example, the semantics could be based on the following:
[0303] `alf_luma_filter_signal_flag` equal to 1 specifies that a signal is sent to notify the luminance filter bank. `alf_luma_filter_signal_flag` equal to 0 specifies that no signal is sent to notify the luminance filter bank. If `alf_luma_filter_signal_flag` does not exist, it is assumed to be equal to 0.
[0304] `alf_chroma_filter_signal_flag` equal to 1 specifies that a signal is sent to the chroma filter. `alf_chroma_filter_signal_flag` equal to 0 specifies that no signal is sent to the chroma filter. If `alf_chroma_filter_signal_flag` does not exist, it is inferred that it is equal to 0.
[0305] `alf_cross_component_cb_filter_signal_flag` equal to 1 indicates that the cross component Cb filter bank is signaled. `alf_cross_component_cb_filter_signal_flag` equal to 0 indicates that the cross component Cb filter bank is not signaled. If `alf_cross_component_cb_filter_signal_flag` does not exist, it is inferred that it is equal to 0.
[0306] `alf_cross_component_cr_filter_signal_flag` equal to 1 specifies that the cross-component (Cr) filter bank should be signaled. `alf_cross_component_cb_filter_signal_flag` equal to 0 specifies that the cross-component (Cr) filter bank should not be signaled. If `alf_cross_component_cr_filter_signal_flag` does not exist, it is assumed to be equal to 0.
[0307] The increment of 1 in `alf_cross_component_cb_min_eg_order_minus1` specifies the minimum order of the exponential Golomb code used for signaling the Cb filter coefficients of the cross-component component. The value of `alf_cross_component_cb_min_eg_order_minus1` should be in the range of 0 to 9 (inclusive). It should be noted that this range may vary in some examples.
[0308] The increment of 1 in `alf_cross_component_cr_min_eg_order_minus1` specifies the minimum order of the exponential Golomb code used for signaling the coefficients of the cross-component (Cr) filter. The value of `alf_cross_component_cb_min_eg_order_minus1` should be in the range of 0 to 9 (inclusive). It should be noted that this range may vary in some examples.
[0309] `alf_cross_component_cb_eg_order_increase_flag[i]` equal to 1 specifies that the minimum order of the exponential Golomb code used for the cross-component Cb filter coefficient signaling is incremented by 1. `alf_cross_component_cb_eg_order_increase_flag[i]` equal to 0 specifies that the minimum order of the exponential Golomb code used for the cross-component Cb filter coefficient signaling is not incremented by 1.
[0310] The exponent of the Golomb code used to decode the value of alf_cross_component_cb_coeff_abs[j], the order expGoOrderCb[i], is derived as follows:
[0311] expGoOrderCb[ i ] = (i = = 0 ? alf_cross_component_cb_min_eg_order_minus1 + 1 : expGoOrderCb[ i - 1 ]) + alf_cross_component_cb_eg_order_increase_flag[ i ]
[0312] In one example, a predetermined value can be used that corresponds to the minimum order of the exponential Golomb code used for signaling the Cb filter coefficients of the cross component.
[0313] In one example, a single value for the minimum order of the exponential Golomb code corresponding to the cross components used for all Cb filter coefficients can be signaled in the bitstream.
[0314] `alf_cross_component_cr_eg_order_increase_flag[i]` equal to 1 specifies that the minimum order of the exponential Golomb code used for the signaling of the cross-component Cr filter coefficients is incremented by 1. `alf_cross_component_cr_eg_order_increase_flag[i]` equal to 0 specifies that the minimum order of the exponential Golomb code used for the signaling of the cross-component Cr filter coefficients is not incremented by 1.
[0315] The exponent of the Golomb code used to decode the value of alf_cross_component_cb_coeff_abs[j], the order of expGoOrderCr[i], is derived as follows:
[0316] expGoOrderCr[ i ] = (i = = 0 ? alf_cross_component_cr_min_eg_order_minus1 + 1 :
[0317] expGoOrderCr[ i - 1 ]) + alf_cross_component_cr_cg order_increase_flag[ i ]
[0318] In one example, a predetermined value is used that corresponds to the minimum order of the exponential Golomb code used for signaling the coefficients of the cross-component Cr filter.
[0319] In one example, a single value is signaled in the bitstream to indicate the minimum order of the exponential Golomb code corresponding to the cross-components used for all Cr filter coefficients.
[0320] alf_cross_component_cb_coeff_abs[j] specifies the absolute value of the j-th coefficient of the cross component Cb filter that signals the signal. If alf_cross_component_cb_coeff_abs[j] does not exist, it is assumed to be equal to 0.
[0321] The order k of the exponential Golomb binarization uek(v) is derived as follows:
[0322] golombOrderIdxCb[ ] = {0,2,2,2,1,2,2,2,2,2,2,1,2,1} [These can classify the coefficients into 3 categories, each using the same k-order exponent Golomb code]
[0323] k = expGoOrderCb[ golombOrderIdxCb[ j ] ]
[0324] In one example, the ue(v) encoding can be used to signal the value of the syntax element alf_cross_component_cb_coeff_abs[j].
[0325] alf_cross_component_cr_coeff_abs[j] specifies the absolute value of the j-th coefficient of the cross component Cr filter that signals the signal. If alf_cross_component_cr_coeff_abs[j] does not exist, it is assumed to be equal to 0.
[0326] The order k of the exponential Golomb binarization uek(v) is derived as follows:
[0327] golombOrderIdxCr[ ] = {0,1,2,1,0,1,2,2,2,2,2,1,2,1} [These can classify the coefficients into 3 categories, each using the same k-order exponent Golomb code]
[0328] k = expGoOrderCr[ golombOrderIdxCr[ j ] ]
[0329] In one example, ue(v) encoding is used to signal the value of the syntax element alf_cross_component_cr_coeff_abs[j].
[0330] alf_cross_component_cb_coeff_sign[j] specifies the sign of the Cb filter coefficient for the j-th cross component as follows:
[0331] - If alf_cross_component_cb_coeff_sign[j] is equal to 0, then the corresponding cross component Cb filter coefficients have positive values.
[0332] Otherwise (alf_cross_component_cb_coeff_sign[j] equals 1), the corresponding cross component Cb filter coefficients have negative values.
[0333] If alf_cross_component_cb_coeff_sign[j] does not exist, it is inferred that it is equal to 0.
[0334] Having element AlfCCCoeff Cb [adaptation_parameter_set_id][j], cross-component Cb filter coefficients AlfCCCoeff, j = 0..13 cb [adaptation_parameter_set_id] is exported as follows:
[0335] AlfCCCoeff cb [ adaptation_parameter_set_id ][ j ] = alf_cross_component_cb_coeff_abs[ j ] *
[0336] (1 - 2 * alf_cross_component_cb_coeff_sign[ j ])
[0337] Bitstream compliance requirements AlfCCCoeff cb [adaptation_parameter_set_id] [j], where j = 0.. 13, should have a value in the range of -2. 10 - 1 to 2 10 -1 (inclusive).
[0338] alf_cross_component_cr_coeff_sign[j] specifies the sign of the filter coefficient for the j-th cross component Cr as follows:
[0339] - If alf_cross_component_cr_coeff_sign[j] is equal to 0, then the corresponding cross component Cr filter coefficients have positive values.
[0340] Otherwise (alf_cross_component_cr_coeff_sign[j] equals 1), the corresponding cross component Cr filter coefficients have negative values.
[0341] If alf_cross_component_cr_coeff_sign[j] does not exist, it is inferred that it is equal to 0.
[0342] Having element AlfCCCoeff Cr [adaptation_parameter_set_id] [j], cross-component Cr filter coefficients AlfCCCoeff, j = 0..13 cr [adaptation_parameter_set_id] is exported as follows:
[0343] AlfCCCoeff Cr [adaptation_parameter_set_id][ j ] = alf_cross_component_cr_coeff_abs[ j ] *
[0344] (1 - 2 * alf_cross_component_cr_coeff_sign[ j ])
[0345] Bitstream compliance requirements AlfCCCoeff Cr [adaptation_parameter_set_id][j], where j = 0.. 13, should have a value in the range of -2. 10 - 1 to 2 10 -1 (inclusive).
[0346] In one example, the alf_data() syntax structure provided in Table 8B can be used. It should be noted that in Table 8B, the number of signaling filters is encoded using minus1 when signaling coefficients are sent in the APS.
[0347]
[0348]
[0349] Table 8B
[0350] For Table 8B, in one example, the semantics could be based on the following:
[0351] `alf_cross_component_cb_filter_signal_flag` equal to 1 indicates that the cross component Cb filter bank is signaled. `alf_cross_component_cb_filter_signal_flag` equal to 0 indicates that the cross component Cb filter bank is not signaled. If `alf_cross_component_cb_filter_signal_flag` does not exist, it is inferred that it is equal to 0.
[0352] `alf_cross_component_cr_filter_signal_flag` equal to 1 specifies that the cross-component (Cr) filter bank should be signaled. `alf_cross_component_cb_filter_signal_flag` equal to 0 specifies that the cross-component (Cr) filter bank should not be signaled. If `alf_cross_component_cr_filter_signal_flag` does not exist, it is assumed to be equal to 0.
[0353] The increment of 1 in `alf_cross_component_cb_filters_signalled_minus1` specifies the number of cross-component Cb filter banks that signal their coefficients. The value of `alf_cross_component_cb_filters_signalled_minus1` should be in the range of 0 to `NumCcAlfCbFilters - 1` (inclusive).
[0354] NumCcAlfCbFilters represents the maximum number of cross-component Cb filter banks allowed in a video sequence. In one example, NumCcAlfCbFilters can be set to a predetermined non-negative integer value. In another example, NumCcAlfCbFilters is signaled to a parameter set such as SPS or PPS.
[0355] The increment of 1 in alf_cross_component_cb_min_eg_order_minus1[k] specifies the minimum order of the exponential Golomb code used for the signaling of the coefficient group of the Cb filter for the k-th cross component. The value of alf_cross_component_cb_min_eg_order_minus1[k] should be in the range of 0 to 9 (inclusive).
[0356] `alf_cross_component_cb_eg_order_increase_flag[k][i]` equal to 1 indicates that the minimum order of the exponential Golomb code used for the signaling of the Cb filter coefficient group of the k-th cross component is incremented by 1. `alf_cross_component_cb_eg_order_increase_flag[k][i]` equal to 0 indicates that the minimum order of the exponential Golomb code used for the signaling of the Cb filter coefficient group of the k-th cross component is not incremented by 1.
[0357] The exponent of the Golomb code used to decode the value of alf_cross_component_cb_coeff_abs[k][j], expGoOrderCb[k][i], is derived as follows:
[0358] expGoOrderCb[ k ] [ i ] = (i = = 0 ? alf_cross_component_cb_min_eg_order_minus1[ k ] + 1 : expGoOrderCbf k ][ i - 1 ]) + alf_cross_component_cb_eg_order_increase_flag[ k ] [ i ]
[0359] `alf_cross_component_cb_coeff_abs[k][j]` specifies the absolute value of the j-th coefficient of the k-th cross component Cb filter bank that is being signaled. If `alf cross component_cb_coeff_abs[k][j]` does not exist, it is assumed to be equal to 0.
[0360] The order k of the exponential Golomb binarization uek(v) is derived as follows:
[0361] golombOrderIdxCb[ ] = {0,2,2,2,1,2,2,2,2,2,2,1,2,1} [Note that these can classify the coefficients into 3 categories, each using the same order 1 exponent Golomb code]
[0362] k = expGoOrderCb(golombOrderIdxCb[ j ] ]
[0363] alf_cross_component_cb_coeff_sign[k][j] specifies the sign of the j-th coefficient in the coefficient group of the k-th cross component Cb filter as follows:
[0364] - If alf_cross_component_cb_coeff_sign[k][j] is equal to 0, then the corresponding cross component Cb filter coefficients have positive values.
[0365] Otherwise (alf_cross_component cb_coeff_sign[k][j] equals 1), the corresponding cross component Cb filter coefficients have negative values.
[0366] If alf_cross_component_cb_coeff_sign[k][j] does not exist, it is inferred that it is equal to 0.
[0367] Having element AlfCCCoeff cb [adaptation_parameter_set_id][k][j], cross-component Cb filter coefficients AlfCCCoeff j = 0..13 cb [adaptation_parameter_set_id] [k] is exported as follows:
[0368] AlfCCCoeff cb [ adaptation_parameter_set_id ] [ k ] [ j ] = alf_cross_component_cb_coeff_abs[ k ] [ j]*
[0369] (1 - 2 * alf_cross_component_cb_coeff_sign[ k ][ j ])
[0370] Bitstream compliance requirements AlfCCCoeff cb[adaptation parameier_set_id][k][j], where j = 0.. 13 should have a value between -2. 7 Up to 2 7 -1 (inclusive) in the range.
[0371] The increment of 1 in `alf_cross_component_cr_filters_signalled_minus1` specifies the number of cross-component (Cr) filter banks that signal their coefficients. The value of `alf_cross_component_cr_filters_signalled_minus1` should be in the range of 0 to `NumCcAlfCrFilters - 1` (inclusive).
[0372] NumCcAlfCrFilters represents the maximum number of cross-component (Cr) filter banks allowed in a video sequence. In one example, NumCcAlfCrFilters can be set to a predetermined non-negative integer value. In another example, NumCcAlfCrFilters is signaled to a parameter set such as SPS or PPS.
[0373] The increment of 1 in alf_cross_component_cr_min_eg_order_minus1[k] specifies the minimum order of the exponential Golomb code used for the signaling of the coefficient group of the k-th cross component Cr filter. The value of alf_cross_component_cb_min_eg_order minus1[k] should be in the range of 0 to 9 (inclusive).
[0374] `alf_cross_component_cr_eg_order_increase_flag[k][i]` equal to 1 indicates that the minimum order of the exponential Golomb code used for the signaling of the k-th cross-component Cr filter coefficient group is incremented by 1. `alf_cross_component_cr_eg_order_increase_flag[k][i]` equal to 0 indicates that the minimum order of the exponential Golomb code used for the signaling of the k-th cross-component Cr filter coefficient group is not incremented by 1.
[0375] The exponent of the Golomb code used to decode the value of alf_cross_component_cb_coeff_abs[k][j], expGoOrderCr[k][i], is derived as follows:
[0376] expGoOrderCr[ k ][ i ] = (i = = 0 ? alf_cross_component_cr_min_eg_order_minus1 [ k ] + 1 : expGoOrderCr[ k ][ i - 1 ]) + alf_cross_component_cr_eg_order_increase_flag[ k ][ i ]
[0377] alf_cross_component_cr_coeff_abs[k][j] specifies the absolute value of the j-th coefficient of the k-th cross component Cr filter bank that is being signaled. If alf_cross_component_cr_coeff_abs[k][j] does not exist, it is assumed to be equal to 0.
[0378] The order k of the exponential Golomb binarization uek(v) is derived as follows:
[0379] golombOrderIdxCr[ ] = {0,1,2,1,0,1,2,2,2,2,2,1,2,1} [Note that these can classify the coefficients into 3 categories, each using the same k-order exponent Golomb code]
[0380] k = expGoOrderCr[ golombOrderIdxCr[ j ] ]
[0381] alf_cross_component_cr_coeff_slgn[k][i] specifies the sign of the j-th coefficient in the coefficient group of the k-th cross component Cr filter as follows:
[0382] - If alf_cross_component_cr_coeff_sign[k][j] is equal to 0, then the corresponding cross component Cr filter coefficients have positive values.
[0383] Otherwise (alf_cross_component_cr_coeff_sign[k][j] equals 1), the corresponding cross component Cr filter coefficients have negative values.
[0384] If alf_cross_component_cr_coeff_sign[k][j] does not exist, it is inferred that it is equal to 0.
[0385] Having element AlfCCCoeff Cr[adaptation_parameter_set_id][k][j], cross-component Cr filter coefficients AlfCCCoeff, j = 0..13 Cr [adaptation_parameter_set_id] [k] is exported as follows:
[0386] AlfCCCoeff Cr [ adaptation_parameter_set_id ][ k ][j ] = alf_cross_component_cr_coeff_abs[ k ][ j ]*
[0387] (1 - 2 * alf_cross_componcnt_cr coeff_sign[ k ] [ j ])
[0388] Bitstream compliance requirements AlfCCCoeff Cr The value of [adaptation_parameter_set_id][k][j], where j = 0.. 13, should be in the range of -2. 7 Up to 2 7 -1 (inclusive).
[0389] For Table 8B, in one example, for the syntax elements alf_cross_component_cb_coeff_abs[k][i] and alf_cross_component_cr_coeff_abs[k][i], it indicates that k of the k-th order exponent golomb encoded uek(v) can correspond to a predetermined value. Therefore, it is not necessary to signal "k" in the bitstream. In one example, the alf_data() syntax structure provided in Table 8C can be used in this case.
[0390]
[0391]
[0392] Table 8C
[0393] For Table 8C, in one example, the semantics may be based on the semantics provided above relative to Table 8B. For the syntax elements alf_cross_component_cb_coeff_abs[k][i] and alf_cross_component_cr_coeff_abs[k][i], in one example, the semantics may be based on the following:
[0394] alf_cross_component_cb_coeff_abs[k][j] specifies the absolute value of the j-th coefficient of the k-th cross component Cb filter bank that is being signaled. If alf_cross_component_cb_coeff_abs[k][j] does not exist, it is assumed to be equal to 0.
[0395] Set the order k of the exponential Golomb binarization uek(v) to be equal to 3.
[0396] alf_cross_component_cr_coeff_abs[k][j] specifies the absolute value of the j-th coefficient of the k-th cross component Cr filter bank that is signaling the signal. If alf_cross_component_cr_coeff_abs[k][j] does not exist, it is assumed to be equal to 0.
[0397] Set the order k of the exponential Golomb binarization uek(v) to be equal to 3.
[0398] Furthermore, for Tables 8A to 8C, in one example, the slice_header() syntax structure provided in Table 8D can be used. It should be noted that in Table 8D, the syntax element slice_cross_component_alf_cb_aps_id specifies that the adaptation_parameter_set_id of the Cb color component index of the slice (in one example, the APS ID (e.g., for cross-component filters, luminance filters, chrominance filters) is not received, and values inferred for, for example, a set of slice types (e.g., I slices) or for images with a set of NALU types (e.g., corresponding to IRAP).
[0399]
[0400] Table 8D
[0401] For Table 8D, in one example, the semantics could be based on the following:
[0402] A slice_cross_component_alf_cb_enabled_flag value of 0 indicates that the cross component Cb filter is not applied to the Cb color component. A slice_cross_component_alf_cb_enabled_flag value of 1 indicates that the cross component Cb filter is applied to the Cb color component.
[0403] A slice_cross_component_alf_cr_enabled_flag value of 0 indicates that the cross component Cr filter is not applied to the Cr color component. A slice_cross_component_alf_cb_enabled_flag value of 1 indicates that the cross component Cr filter is applied to the Cr color component.
[0404] For Table 8A:
[0405] The `slice_cross_component_alf_cb_reuse_temporal_layer_filter` value being equal to 1 specifies that the cross-component Cb filter coefficients (where j = 0..13, inclusive) are set to equal to `AlfCCTemporalCoeff`. cb [Temporalld][j].
[0406] The condition that slice_cross_component_alf_cb_reuse_temporal_layer_filter is equal to 0 and slice_cross_component_alf_cb_enabled__flag is equal to 1 indicates that the syntax element slice_cross_component_alf_cb_aps_id exists in the slice header.
[0407] When slice_cross_component_alf_cb_enabled_flag equals 1 and slice_cross_component_alf_cb_reuse_temporal_layer_filter equals 0, the following export of AlfCCTemporalCoeff is performed. cb [Temporalld][j], elements j=0..13:
[0408] AlfCCTemporalCoeff Cb [Temporalld][j] = AlfCCCoeff Cb [ slice_cross_component_alf_cb_aps_id ][ j ]
[0409] In one example, the cross component Cb coefficient (where j = 0..13) can be set to the corresponding APSAlfCCCoeff. Cb The coefficients received in [slice_cross_component_alf_cb_aps_id][j].
[0410] `slice_cross_component_alf_cr_reuse_temporal_layer_filter` equal to 1 specifies that the cross-component Cr filter coefficients (where j = 0..13, inclusive) are set to equal to `AlfCCTemporalCoeff`. Cr [TemporalId][j].
[0411] The syntax element slice_cross_component_alf_cr_reuse_temporal_layer_filter being equal to 0 and slice_cross_component_alf_cr_enabled_flag being equal to 1 indicates that the slice_cross_component_alf_cr_aps_id exists in the slice header.
[0412] When slice_cross_component_alf_cr_enabled_flag equals 1 and slice_cross_component_alf_cr_reuse__temporal_layer_filter equals 0, the following export of AlfCCTemporalCoeff is performed. Cr [TemporalId][j], elements where j = 0..13:
[0413] AlfCCTemporalCoeff Cr [TemporalId][j] = AlfCCCoeff Cr [ slice_cross__component_alf_cr_aps_id ][ j ]
[0414] In one example, the cross component Cr coefficient (where j = 0..13) can be set to the corresponding APSALfCCCoeff. Cr The coefficients received in [slice_cross_component_alf_cr_aps_id][j].
[0415] For Tables 8B and 8C:
[0416] The `slice_cross_component_alf_cb_reuse_temporal_layer_filter` being equal to 1 specifies that the cross-component Cb filter coefficients (where j = 0..13, k = 0..(NumCcAlfCbFilters - 1), including end values) are set to equal to `AlfCCTemporalCoeff`. Cb [TemporalId] [k] [j].
[0417] The syntax element slice_cross_component_alf_cb_reuse_temporal_layer_filter being equal to 0 and slice_cross_component_alf_cb_enabled_flag being equal to 1 indicates that the slice_cross_component_alf_cb_aps_id exists in the slice header.
[0418] When slice_cross_component_alf_cb_enabled_flag equals 1 and slice_cross_component_alf_cb_reuse_temporal_layer_filter equals 0, the following export of AlfCCTemporalCoeff is performed. cb Elements of [TemporalId][k][j], j=0..13, k=0..(NumCcAlfCbFilters -1):
[0419] AlfCCTemporalCoeff Cb [ TemporalId ][ k ][ j ] = AlfCCCoeff Cb [slice_cross_component_alf_cb_ap s_id][ k][j]
[0420] The slice_cross_component_alf_cr_reuse_temporal_layer_filter being equal to 1 specifies that the cross-component Cr filter coefficients (where j = 0..13, k = 0..(NumCcAlfCrFilters - 1), including end values) are set to equal AlfCCTemporalCoeff. Cr [TemporalId][k][j].
[0421] The syntax element slice_cross_component_alf_cr_reuse_temporal_layer_filter being equal to 0 and slice_cross_component_alf_cr_enabled_flag being equal to 1 indicates that the slice_cross_component_alf_cr_aps_id exists in the slice header.
[0422] When slice_cross_component_alf_cr_enabled_flag equals 1 and slice_cross_component alf_cr_reuse_tcmporal_layer_filter equals 0, the following export of AlfCCTemporalCoeff is performed. Cr Elements of [TemporalId][k][j], j=0..13, k=0..(NumCcAlfCrFilters -1):
[0423] AlfCCTemporalCoeff Cr [ TemporalId ][ k ][ j ] = AlfCCCoeff Cr [slice_cross_component_alf_cr_aps id][k][j]
[0424] It should be noted that when using filter coefficient sets from APS, the corresponding time filter coefficient set buffer is filled with filter coefficient sets from APS.
[0425] `slice_cross_component_alf_cb_aps_id` specifies the `adaptation_parameter_set_id` indexed for the cross-component Cb filter of the slice's Cb color component. If `slice_cross_component_alf_cb_aps_id` does not exist, it is inferred to be equal to `slice_alf_aps_id_luma[0]`. The `TemporalId` of an ALF APS NAL cell with an `adaptation_parameter_set_id` equal to `slice_cross_component_alf_cb_aps_id` should be less than or equal to the `TemporalId` of the encoded slice NAL cell.
[0426] In one example, the APS can be reset, for example, for a set of NALU types (e.g., corresponding to random access points such as IRAP) or a set of slice types (e.g., I slices). For such images / slices, `lice_cross_component_alf_cb_aps_id` is not signaled. In one example, it can be inferred as a predetermined value. In another example, it can be inferred as an derived value (e.g., based on slice type, TemporalID, NALU type, QP).
[0427] In one example, slice_cross_component_alf_cb_aps_id is not received, and values can be inferred for a set of slices (e.g., the first slice) of an image that has a set of NALU types (e.g., corresponding to IRAP).
[0428] `slice_cross_component_alf_cr_aps_id` specifies the `adaptation_parameter_set_id` indexed for the cross-component Cb filter of the slice's Cr color component. If `slice_cross_component_alf_cr_aps_id` does not exist, it is inferred to be equal to `slice_alf_aps_id_luma[0]`. The `TemporalId` of an ALF APS NAL unit with `slice_cross_component_alf_cr_aps_id` should be less than or equal to the `TemporalId` of the encoded slice NAL unit.
[0429] In one example, the APS can be reset, for example, for a set of NALU types (e.g., corresponding to random access points such as IRAP) and for a set of slice types (e.g., I slices). For such images / slices, instead of signaling and inferring slice_cross_component_alf_cr_aps_jd, it is inferred as a predetermined value. In one example, it can be inferred as a predetermined value. In another example, it can be inferred as an derived value (e.g., based on slice type, TemporalID, NALU type, QP).
[0430] In one example, slice_cross_component_alf_cr_aps_id was not received, and values were inferred for a set of slices (e.g., the first slice) of an image with a set of NALU types (e.g., corresponding to IRAP).
[0431] In one example, an APS reset operation may imply that an APS encoded before the access cell (meeting a predetermined set of conditions, such as NALU type indicating IRAP, slice type equal to I slice) may not be available for that access cell and subsequently encoded access cells.
[0432] The following values specify the size of the square blocks in units of the number of samples: slice_cross_component_alf_cb_log2_control_size_minus4
[0433] AlfCCSamplesCbW = AlfCCSamplesCbH = 2 (slice _cross_component_alf_cb_log2_control_size_minus4+4)
[0434] The slice_cross_component_alf_cb_log2_control_size_minus4 should be in the range of 0 to 3 (inclusive).
[0435] The following values specify the size of the square blocks in units of the number of samples: slice_cross_component_alf_cr_log2_control_size_minus4
[0436] AlfCCSamplesCrW = AlfCCSamplesCrH = 2 (slice _cross_component_alf_cr_log2_control_size_minus4+4)
[0437] slice_cross_component_alf_cr_log2_control_size_minus4 should be in the range of 0 to 3 (inclusive).
[0438] TemporalId is the time identifier of the current NAL cell.
[0439] It should be noted that in the above examples, the coefficient values are indicated using syntax elements that indicate the sign of the coefficient (e.g., `alf_cross_component_cr_coeff_sign`) and syntax elements that indicate the absolute value of the coefficient (e.g., `alf_cross_component_cr_coeff_abs`). In one example, according to the techniques described herein, the absolute value of a coefficient can be indicated using one or more flags indicating whether the absolute value of the coefficient is greater than a specific value (e.g., greater than 1, greater than 2, etc.) and syntax elements that indicate the absolute value of the coefficient based on the value of those flags. For example, in one example, the encoding of a specific coefficient value can be based on the following semantics:
[0440] `alf_cross_component_coeff_abs_greater_than_N_flag[k][j]` specifies whether the absolute value of the j-th coefficient of the k-th cross-component filter bank, which is being signaled, is greater than N. In one example, the coefficient should be in the range of -2. 7 Up to 2 7 The range is -1 (inclusive), and N can be 64.
[0441] `alf_cross_component_coeff_abs[k][j]` specifies the absolute value of the `j`-th coefficient of the `k`-th cross-component filter bank that is being signaled. If `alf_cross_component_coeff_abs[k][j]` does not exist, it is assumed to be equal to 0.
[0442] alf_cross_component_coeff_sign[k][i] specifies the sign of the j-th coefficient in the k-th cross component filter coefficient group as follows:
[0443] - If alf_cross_component_coeff_sign[k][j] equals 0, then the corresponding cross component filter coefficients have positive values.
[0444] Otherwise (alf_cross_component_coeff_sign[k][j] equals 1), the corresponding cross component filter coefficients have negative values.
[0445] If alf_cross_component_coeff_sign[k][j] does not exist, it is inferred that it is equal to 0.
[0446] The cross-component filter coefficients AlfCCCoeff[adaptation_parameter_set_id][k][j] with elements AlfCCCoeff[adaptation_parameter_set_id][k] are derived as follows:
[0447] AlfCCCoeff [ adaptation_parameter_set_id ] [ k ] [ j ] = (N*alf_cross_component_coeff_abs_greater than N flag[ k ][ j ] +alf_cross_component_coeff_abs[ k ][ j ]) *
[0448] (1 - 2 * alf_cross_component_coeff_sign[ k ] [ j ])
[0449] In one example, a set of cross-component filter coefficients can be signaled for the cross-component color (Cb and / or Cr) filters. Each set of cross-component filter coefficients can be assigned a cross-component filter index. In one example, this set of cross-component filter coefficients can be signaled in the APS. In one example, sample values can be partitioned (e.g., using a technique similar to that used in control tag signaling or any suitable alternative to determine partitioning). Filter indexes can be signaled for each partition of sample values, where the filter index identifies the cross-component filter to be applied to the samples in the partition. Partitioning can be transmitted using parameters such as block size (which can be the same as the control block size parameter or can be independent). Partition regions can be derived using parameter values, picture / slice / tile group / MCTS size. In one example, only one cross-component filter coefficient in the set can be signaled in the APS for each color component. An APS identifier can be signaled for each sample value partition that identifies the cross-component filter to be applied to the samples in the partition.
[0450] In one example, buffers for a subset or all of the temporal layers can be reset, for example, for a set of NALU types (e.g., corresponding to random access points such as IRAP) and for a set of slice types (e.g., I-slices). In one example, a reset may imply a clear operation. In one example, a reset may imply setting the buffer to a predetermined set of values (e.g., 0, a fixed set of values for each TemporalID). A buffer reset operation may imply that values stored in the buffer prior to the access unit (satisfying a predetermined set of conditions, e.g., NALU type indicating IRAP, slice type equal to I-slice) may not be available for that access unit and subsequent encoded access units. In one example, when the buffer does not contain any coefficients, syntax elements indicating whether filter coefficients from the buffer will be used can be inferred without signaling. For example, when clearing the buffer for an IRAP picture, the first slice in the IRAP picture does not need to signal slice_cross_component_alf_cb_reuse_temporal_layer_filter and / or slice_cross_component_alf_cr_reuse_temporal_layer_filter. In one example, when the buffer contains no coefficients, the syntax elements indicating whether filter coefficients from the buffer will be used can be signaled to be restricted to predetermined values. For example, when the buffer is cleared for an IRAP image, the syntax elements slice_cross_component_alf_cb_reuse_temporal_layer_filter and / or slice_cross_component_alf_cr_reuse_temporal_layer_filter in the first slice of the IRAP image need to be 0.
[0451] Furthermore, in one example, a signal is sent indicating that the filter used for the preamble image should not be used by the trailing image of the associated IRAP image. This is because the preamble image can be discarded from the bitstream during random access operations. As described above, according to the techniques described herein, the cross-component filter can be signaled by reusing the filter in the corresponding temporal sublayer buffer and / or the filter in the APS. It should be noted that in some cases, the APS can be signaled out of band, and therefore in some cases, bitstream compliance may only require that the indexed APS be available. Thus, in one example, the bitstream compliance requirement can be implemented using the syntax elements slice_cross_component_alf_cb_reuse_temporal_layer_filter and slice_cross_component_alf_cr_reuse_temporal_layer_filter. In one example, the syntax elements slice_cross_component_alf_cb_reuse_temporal_layer_filter and slice_cross_component_alf_cr_reuse_temporal_layer_filter can have the following constraints:
[0452] When the filter coefficients in the temporal sublayer buffer AlfCCTemporalCoeffCb[TemporalId][k][j] with temporal identifier TemporalId are exported for the leading image of an associated IRAP image, the bitstream conforms to the requirement that the value of the syntax element slice_cross_component_alf_cb_reuse_temporal_layer_filter should be 0 for the trailing image of an associated IRAP image with temporal identifier TemporalId.
[0453] When the filter coefficients in the temporal sublayer buffer AlfCCTemporalCoeffCr[TemporalId][k][j] with temporal identifier TemporalId are exported for the leading image of an associated IRAP image, the bitstream conforms to the requirement that the value of the syntax element slice_cross_component_alf_cr_reuse_temporal_layer_filter should be 0 for the trailing image of an associated IRAP image with temporal identifier TemporalId.
[0454] As mentioned above Figures 9A to 9FThe filter shape can be determined based on the chroma location type. It should be noted that the filter shape can include filter shapes for various types of filters. Furthermore, the filter shape can generally be described as representing the support and the origin relative to the filtered sample. In one example, the filter shape can be signaled for each cross-component filter in the parameter set: for example, SPS, PPS, VUI. In one example, the filter shape can also be signaled in the APS or slice header, for example, when the filter shape affects the number of encoded filter coefficients. In one example, the filter shape can be used to determine the number of filter coefficients. In one example, filter coefficients can be reused. That is, in one example, a first-in-first-out (FIFO) buffer can be maintained for each component Cb / Cr for each time layer. In one example, the size of the FIFO buffer is 1 for each component and each time layer. In such an example, a signaling flag can be sent to indicate whether to use the filter coefficients in the corresponding (time layer) FIFO buffer or to receive a new set of coefficients. When a new set of filter coefficients is received, they replace the contents of the corresponding (time layer) FIFO buffer. In one example, the size of the FIFO buffer is 1 for each component and each time layer. In such an example, filter coefficients belonging to the same or lower time layer in a FIFO buffer can be reused. The received syntax element (e.g., a flag) indicates whether a filter coefficient in one of the FIFO buffers is reused, and if it is reused (e.g., the flag value is 1), the index of the FIFO buffer containing the coefficient to be reused is received. The range of valid values for that index (e.g., 0 to the current time layer ID) is determined based on the current time layer ID. ue(v) encoding can be used to signal values such as -(FIFO buffer time layer - current time layer). In another example, a truncated unary (with a maximum TU value based on the current time layer ID) can be used.
[0455] In one example, signaling for local control of the cross-component filter may include sending all cross-component Cb and Cr block-level control flags for the slice in the first CTU of the slice. In one example, when the control block size is larger than the CTU size, control flags may be sent in the first coded CTU of the control block. The remaining CTUs in the control block can infer the same value as in the first coded CTU of the control block. When the first coded CTU in the control block does receive a control flag, the value of the control flag can be inferred to be 0. In one example, control flags are sent for each CTU. In one example, control flags may be signaled in the first CTU of the slice / tile group. Furthermore, in one example, four control flags may be signaled for each CTU of the tile group / slice. In one example, when there are four control blocks within a CTU, a different number of control block flags may exist for some CTUs (e.g., at the boundaries) compared to the full CTU. In one example, one control flag may be signaled for each CTU of the tile group / slice. In one example, control flags may be signaled for the first CTU in a group of CTUs, and the same flag value can be inferred for the remaining CTUs in that group.
[0456] Table 9 shows the syntax of the `coding_tree_unit()` syntax structure used to signal notification in the first CTU of the control block, and examples of multiple markers when the control block is smaller than the CTU. It should be noted that in Table 9, the function `isInSliceCb()` returns TRUE when the passed Cb chroma position represents a Cb sample within the slice being encoded, and FALSE otherwise; similarly, the function `isInSliceCr()` returns TRUE when the passed Cr chroma position represents a Cr sample within the slice being encoded, and FALSE otherwise.
[0457]
[0458] Table 9
[0459] Table 10 shows the syntax in the coding_tree_unit() syntax structure for signaling notification in the first CTU of the control block, and examples of multiple tags when the control block is smaller than the CTU.
[0460]
[0461] Table 10
[0462] For Tables 9 and 10, in one example, the semantics could be based on the following:
[0463] cross_component_alf_cb_control_flag[x][y] equal to 1 specifies that the cross component Cb filter is applied to the block of color component Cb located at the chromaticity position (xCtb / SubWidthC + x * AlfCCSamplesCbW, yCtb / SubHeightC + y * AlfCCSamplesCbH).
[0464] cross_component_alf_cb_flag[x][y] equal to 0 specifies that the cross component Cb adaptive loop filter is not applied to the color component Cb block located at the chromaticity position (xCtb / SubWidthC + x *AlfCCSamplesCbW, yCtb / SubHeightC + y * AlfCCSamplesCbH).
[0465] If cross_component_alf_cb_flag[x][y] does not exist, it is inferred that it is equal to 0.
[0466] The cross_component_alf_cr_control_flag[x][y] being equal to 1 specifies that the cross component Cr filter is applied to the block of color component Cr located at the chromaticity position (xCtb / SubWidthC + x * AlfCCSamplesCrW, yCtb ISubHeightC + y * AlfCCSamplesCrH).
[0467] cross_component_alf_cr_flag[x][y] equal to 0 specifies that the cross component Cr adaptive loop filter is not applied to the block of color component Cr located at the chromaticity position (xCtb / SubWidthC + x *AlfCCSamplesCrW, yCtb / SubHeightC + y * AlfCCSamplesCrH).
[0468] If cross_component_alf_cr_flag[x][y] does not exist, it is inferred that it is equal to 0.
[0469] In one example, according to the techniques described herein, the values of spatially adjacent flags are used to select the context for encoding the control flag. In one example, there are a total of three contexts for `cross_component_alf_cb_control_flag` and `cross_component_alf_cr_control_flag`, which can be determined as follows:
[0470] int ctxt = 0;
[0471] / / Left side
[0472] if (the cross component Cb control flag on the left is available) / / One of the availability conditions is x>0, other examples could be that it should not cross slice boundaries, should not cross image boundaries, should not cross MCTS boundaries
[0473] {
[0474] ctxt+= cross_component_alf_cb_control_flag[xl]|y] ? 1 : 0;
[0475] }
[0476] / / top
[0477] If (the top cross component Cb control flag is available) / / One of the availability conditions is y>0; other examples could be that it should not cross CTU boundaries, should not cross slice boundaries, should not cross image boundaries, or should not cross MCTS boundaries.
[0478] {
[0479] ctxt+= cross_component_alf__cb_control_flag[x][y-1] ? 1 : 0;
[0480] }
[0481] int ctxt = 0;
[0482] / / Left side
[0483] If the cross component Cr control flag on the left is available.
[0484] {
[0485] ctxt+= cross_component_alf_cr_control_flag[xl][y] ? 1 : 0;
[0486] }
[0487] / / top
[0488] If (the top cross component Cr control flag is available)
[0489] {
[0490] ctxt+= cross_component_alf_cr_control_flag[x][yl] ? 1 : 0;
[0491] }
[0492] In one example, the cross-component filter coefficients used for the cross-component filter carry their own independent parameter set, such as the APS. In another example, the cross-component filter coefficients carry a parameter set different from that of the non-cross-component filter, such as the APS (e.g., the ALF in JVET-N1001-v8). In one example, instead of a 5×6 filter, an implementation can use a 7×7 ALF filtering process described in the section on "Coded Tree Block Filtering for Luminance Samples" in JVET-N1001-v8. This will further reduce the number of coefficients that need to be signaled. In the above description, filtering operations are applied to refine samples in the color components and / or channels, and signaling is provided to enable / disable the operations frame-by-frame and position-by-position. In one example, filtering operations can also be implemented such that multiple filtering operations are available at each frame and / or position, and the signaling can be used to transmit these multiple filters and select among the available filters.
[0493] In one example, according to the techniques described herein, local region control flags can be shared across different chroma channels. That is, for example, the syntax elements `cross_component_alf_cb_control_flag` and `cross_component_alf_cr_control_flag` can be replaced with the syntax element `cross_component_chroma_alf_control_flag`, where, in one example, the semantics of `cross_component_chroma_alf_control_flag` are based on the following:
[0494] `cross_component_chroma_alf_control_flag[x][y]` equals 1 and `slice_cross_component_alf_cb_enabled_flag` equals 1, specifying that the adaptive loop filter for the cross component Cb is applied to samples of the color component Cb located within the block at the chromaticity position (xCtb / SubWidthC + x * AlfCCSamplesCbW, yCtb / SubHeightC + y * AlfCCSamplesCbH). `cross_component_chroma_alf_flag[x][y]` equals 1 and `slice_cross_component_alf_cr_enabled_flag` equals 1, specifying that the adaptive loop filter for the cross component Cr is applied to samples of the color component Cr located within the block at the chromaticity position (xCtb / SubWidthC + x * AlfCCSamplesCbW, yCtb / SubHeightC + y * AlfCCSamplesCbH). A cross_component_chroma_alf_flag[x][y] equal to 0 specifies that the cross-component chroma adaptive loop filter is not applied to the sample block of color components Cb and Cr located at the chroma position (xCtb / SubWidthC + x * AlfCCSamplesCbW, yCtb / SubHeightC + y * AlfCCSamplesCbH).
[0495] If cross component_chroma_alf_flag[x][y] does not exist, it is inferred that it is equal to 0.
[0496] Furthermore, in one example, the syntax elements slice_cross_component_alf_cb_enabled_flag and slice_cross_component_alf_cr_enabled_flag can be replaced with the syntax element slice_cross_component_chroma_alf_enabled_flag, and the syntax elements cross_component_alf_cb_control_flag and cross_component_alf_cr_control_flag can be replaced with the syntax element cross_component_chroma_alf_control_flag. In one example, the semantics of slice_cross_component_chroma_alf_enabled_flag and cross_component_chroma_alf_control_flag are based on the following:
[0497] The fact that `cross_component_chroma_alf_control_flag[x][y]` equals 1 and `slice_cross_component_alf_chroma_enabled_flag` equals 1 specifies that the cross component Cb adaptive loop filter is applied to the samples of color component Cb located in the block at the chromaticity position (xCtb / SubWidthC + x * AlfCCSamplesCbW, yCtb / SubHeightC + y * AlfCCSamplesCbH).
[0498] The condition that `cross_component_chroma_alf_flag[x][y]` equals 1 and `slice_cross_component_alf_chroma_enabled_flag` equals 1 specifies that the cross component Cr adaptive loop filter is applied to the samples of the color component Cr located in the block at the chromaticity position (xCtb / SubWidthC + x * AlfCCSamplesCbW, yCtb / SubHeightC + y * AlfCCSamplesCbH).
[0499] A cross_component_chroma_alf_flag[x][y] equal to 0 indicates that the cross-component chroma adaptive loop filter is not applied to the sample blocks of color components Cb and Cr located at chroma positions (xCtb / SubWidthC + x *AlfCCSamplesCbW, yCtb / SubHeightC + y * AlfCCSamplesCbH).
[0500] If cross_component_chroma_alf_flag[x][y] does not exist, it is inferred that it is equal to 0.
[0501] For these examples, Table 11 shows examples of syntax included in the coding_tree_unit() syntax structure for the first CTU used for slicing, and Table 12 shows examples of syntax included in the coding_tree_unit() syntax structure for the corresponding CTU.
[0502]
[0503] Table 11
[0504]
[0505] Table 12
[0506] As described above, the application of cross-component filtering can be based on the properties of the samples included in the filter support region (e.g., variance and / or bias). In one example, a shared control flag (e.g., cross_component_chroma_alf_control_flag) and the properties of the samples included in the filter support region can be used to determine whether cross-component filtering is applied to one or both chroma channels. For example, in one example, the value of cross_component_chroma_alf_control_flag can determine whether cross-component filtering is applied for Cb, and for Cr, the value of cross_component_chroma_alf_control_flag can additionally determine whether cross-component filtering is applied.
[0507] As described above, in JVET-N1001 and JVET-O2001, CTUs are partitioned according to a quadtree plus multi-type tree (QTMT or QT+MTT) structure. In JVET-N1001 and JVET-O2001, for I-slices, each CTU can be partitioned into coding units with 64×64 luma samples using implicit quadtree partitioning, and these coding units can be the roots of two separate coding tree syntax structures, one for the luma channel and one for the chroma channel. In either case, local area control flag values can be signaled / determined within the syntax of the partitioning tree signaling. That is, for example, the syntax and semantics for partitioning chroma channels into CUs (i.e., the chroma coding tree) can include signaling (implicit and / or explicit) indicating local area control flag values. For example, JVET-N1001 and JVET-O2001 define the variable cbSubdiv = 2 * cqtDepth, where cqtDepth is the current coding quadtree depth, and the depth at the CTU is 0. In one example, according to the techniques described herein, the smallest tree node whose cbSubdiv is less than or equal to a given threshold can represent the parent node control flag signaling group, where all blocks resulting from further segmentation belong to the same control flag signaling group. That is, the local area control flag at the parent CU can be used to infer the value of the local area control flag at any resulting child CU. Tables 13 and 14 show examples of the syntax and variable assignments used to determine the value of the local area control flag. That is, in Table 13, the smallest tree node whose cbSubdiv is less than or equal to a given threshold can represent the parent node control flag signaling group. Table 14 provides the corresponding coding unit syntax.
[0508]
[0509] Table 13
[0510]
[0511] Table 14
[0512] In the example shown in Table 13, the CC ALF control signaling depth used for slicing represents a threshold. In one example, this threshold can be signaled for a slice, sequence, or image. For Table 13, the luminance position ((xCtrlBlk, yCtrlBlk) specifies the top-left luminance sample of the current chroma control block relative to the top-left luminance sample of the current image. The horizontal position xCtrlBlk and the vertical position yCtrlBlk are set to equal CuCcAlfTopLeftX and CuCcAlfTopLeftY, respectively. Furthermore, the current chroma control block is a rectangular region within the coding tree block that shares the same CC ALF control flag value (shared or independent for chroma components). Its width and height are equal to the width and height of the coding tree node, whose top-left luminance sample position is assigned to the variables CuCcAlfTopLeftX and CuAlfCcTopLeftY. It should be noted that in one example, each chroma component may have an independent set of variables.
[0513] For Table 14, in one example, when `cross_component_chroma_alf_control_flag` (shared or independent) does not exist, it is inferred to be equal to 0, and when `cross_component_chroma_alf_control_flag` (shared or independent) exists, the variable `IsCuCcAlfControlFlagCoded` is set to 1. Furthermore, the Cu CC ALF control flag `Vai` is set to `cross_component_chroma_alf_control_flag`. Therefore, in this example, as long as the condition for starting a new control flag signaling group is true (i.e., CC ALF is enabled for the slice and `cuSubdiv` is not higher than the limit), the internal flag `IsCuCcAlfControlFlagCoded` is set to 0, and the origin of the current tree node is saved as the control flag signaling group origin in the variables `CuCcAlfTopLeftX` and `CuCcAlfTopLeftY`. Later, in the coding unit syntax, if "IsCuCcAlfControlFlagCoded" is zero, the control signaling group flag is encoded, and the "IsCuCcAlfControlFlagCoded" flag is set to 1, preventing other control block flags from being encoded until a new control flag signaling group is found. The coding unit can inherit its control flag value from the origin of the last coding tree node until a new control flag signaling group is found. In one example, when a local control area syntax element exists in the coding unit, IsCuCcAlfControlFlagCoded is set to 1. It should be noted that JVET-N1001 and JVET-O2001 provide quantization parameter groups indicating the lowest depth of the signaling QP value. In one example, the CC ALF control signaling can be an aligned quantization parameter group. That is, child nodes within the quantization parameter group share the QP value and the CC ALF control value. In addition, in one case, the value of the CC ALF control value can be based on the QP value or derived entirely from the QP value (e.g., if the QP value is less than a threshold, the CC ALF control flag is not signaled to the group but is inferred to be 0).
[0514] As described above, the signaling for cross-component filtering can include signaling for a specific filter (e.g., filter shape and / or filter coefficients). In one example, according to the techniques described herein, the value of a syntax element can indicate whether cross-component filtering is applied to a region, and when cross-component filtering is applied to a region, it indicates the specific filter used for that region. For example, a value of 0 can indicate that cross-component filtering is not applied to a region, a value of 1 can indicate that a filter with a first set of filter coefficients is applied, a value of 2 can indicate that a filter with a second set of filter coefficients is applied, and so on. In one example, the region can be a CTU. Table 15 shows an example of syntax included in the coding_tree_unit() syntax structure according to the techniques described herein. That is, in the example shown in Table 15, the syntax elements alf_cross_component_cb_idc and alfcross_component_cb_idc indicate whether cross-component filtering is applied, and when cross-component filtering is applied, they indicate the filter.
[0515]
[0516]
[0517]
[0518] Table 15
[0519] For Table 15, in one example, the semantics can be based on the following:
[0520] PicWidthlnChromaSamples = pic_width_in_luma__samples / SubWidthC
[0521] PicHeightlnChromaSamples = pic_height_in_luma_samples / SubHeightC
[0522] AlfCCSamplesCbLog2W = AlfCCSamplesCbLog2H = slice_cross_component_alf_cb_log2_control_size_minus4 + 4
[0523] AlfCCSamplesCrLog2W = AlfCCSamplesCrLog2H = slice_cross_component_alf_cr_log2_control_size_minus4 + 4
[0524] Furthermore, it should be pointed out that:
[0525] In one example, xStartC corresponds to the top edge of the CTU. In another example, yStartC corresponds to the left edge of the CTU.
[0526] When the local control region does not cross the CTU, the check (xCtbC == xStartC && yCtbC == yStartC) can be omitted.
[0527] In one example, (xEndC >= PicWidthlnChromaSamples) and (yEndC >= PicHeightlnChromaSamples), the upper limits provided by PicWidthlnChromaSamples and PicHeightlnChromaSample can alternatively correspond to the right and bottom slice / tile / brick / CTU boundaries, respectively.
[0528] `alf_cross_component_cb_idc[ xC>> AlfCCSamplesCbLog2W ][ yC>> AlfCCSamplesCbLog2 H ]` equal to 0 indicates that the cross-component Cb filter is not applied to the block with the Cb chromaticity sample at the top-left chromaticity position (xC, yC). `alf_cross_component_cb_idc[ xC>> AlfCCSamplesCbLog2W ][yC>> AlfCCSamplesCbLog2H ]` equal to m, where m is greater than 0, indicates that the k=(ml)th cross-component Cb filter bank is applied to the block with the Cb chromaticity sample at the top-left chromaticity position (xC, yC).
[0529] `alf_cross_component_cr_idc[ xC>> AlfCCSamplesCrLog2W ][ yC>> AlfCCSamplesCrLog2H]` equal to 0 indicates that the cross-component Cr filter is not applied to the block with the Cr chromaticity sample at the top-left chromaticity position (xC, yC). `alf_cross_component_cr_idc[ xC >> AlfCCSamplesCrLog2W ][ yC >> AlfCCSamplesCrLog2H ]` equal to m, where m is greater than 0, indicates that the k=(ml)th cross-component Cr filter group is applied to the block with the Cr chromaticity sample at the top-left chromaticity position (xC, yC).
[0530] The function UnavailableCb(xC,yC) returns TRUE when the alf_cross_component_cb_idc[ xC >> AlfCCSamplesCbLog2W ][ yC >> AlfCCSamplesCbLog2 H ] corresponding to the chroma sample position (xC,yC) is unavailable, for example when (xC,yC) is located in a different slice / tile compared to the current CTU; otherwise, it returns FALSE.
[0531] The function UnavailableCr(xC,yC) returns TRUE when the alf_cross_component_cr_idc[ xC >> AlfCCSamplesCrLog2W ][ yC >> AlfCCSamplesCrLog2H ] corresponding to the chroma sample position (xC,yC) is unavailable, for example when (xC,yC) is located in a different slice / tile compared to the current CTU; otherwise, it returns FALSE.
[0532] In one example, according to the techniques described herein, the binarization of alf_cross_component_cb_idc and / or alf_cross_component_cr_idc can be a truncated Rice (TR) binarization with a maximum value cMax (based on the number of filter coefficient groups notified by the signal) and cRiceParam equal to 0. Table 16 shows examples of truncated Rice binarization of alf_cross_component_cb_idc and / or alf_cross_component_cr_idc.
[0533]
[0534] in,
[0535] For CC ALF Cb, cMax is set to the corresponding value (alf_cross component cb filters_signalled_minus1 plus 1).
[0536] For CC ALF Cr, cMax is set to the corresponding value (alf_cross_component cr_filterssignalled_minus1 plus 1).
[0537] In one example, for `alf_cross_component_cb_idc` and / or `alf_cross_component_cr_ide`, context encoding can be performed only on the first bin (i.e., bypass encoding can be performed on the other bins). In one example, the context of the first bin can be derived as follows:
[0538] The inputs to this process are the brightness position (x0, y0) of the current luminance block relative to the top-left luminance sample of the current image, the color component cIdx, the current coding quadtree depth cqtDepth, the dual-tree channel type chType, the width and height cbWidth and cbHeight of the current coding block in the luminance sample, and the variables allowSplitBtVer, allowSplitBtHor, allowSplitTtVer, allowSplitTtHor, and allowSplitQt derived from the coding tree semantics.
[0539] The output of this process is ctxInc.
[0540] The position (xNbL, yNbL) is set to equal to (x0-l, y0), and the derived procedure for the availability of the specified neighboring block is invoked, where the position (xCurr, yCurr) is set to equal to (x0, y0), the neighboring position (xNbY, yNbY) is set to equal to (xNbL, yNbL), checkPredModeY is set to equal to FALSE, and cIdx is taken as input, and the output is assigned to availableL.
[0541] The position (xNbA, yNbA) is set to equal to (x0, y0 - 1), and the derived procedure for the availability of the specified neighboring block is invoked, where the position (xCurr, yCurr) is set to equal to (x0, y0), the neighboring position (xNbY, yNbY) is set to equal to (xNbA, yNbA), checkPredModeY is set to equal to FALSE, and cIdx is taken as input, and the output is assigned to availableA.
[0542] The allocation of ctxInc is specified as follows, where condL and condA are specified in Table 17:
[0543] For the syntax elements alf_cross_component_cb_idc[x0][y0] and alf_cross_component_cr_idc[x0][y0]:
[0544] ctxInc = (condL && availableL) + (condA && availableA) + ctxSetIdx *3
[0545]
[0546] In one example, for alf_cross_component_cb_idc and / or alf_cross_component_cr_idc, context encoding can be performed on all bins, and each bin can use a separate set of contexts as shown in Table 18.
[0547]
[0548] For Table 18, in one example, the context can be derived as follows:
[0549] The inputs to this process are the brightness position (x0, y0) of the current luminance block relative to the top-left luminance sample of the current image, the color component cIdx, the current coding quadtree depth cqtDepth, the dual-tree channel type chType, the width and height cbWidth and cbHeight of the current coding block in the luminance sample, and the variables allowSplitBtVer, allowSplitBtHor, allowSplitTtVer, allowSplitTtHor, and allowSplitQt derived from the coding tree semantics.
[0550] The output of this process is ctxInc.
[0551] The position (xNbL, yNbL) is set to equal to (x0-1, y0), and the derived procedure for the availability of the specified neighboring block is invoked, where the position (xCurr, yCurr) is set to equal to (x0, y0), the neighboring position (xNbY, yNbY) is set to equal to (xNbL, yNbL), checkPredModeY is set to equal to FALSE, and cIdx is taken as input, with the output assigned to availableL. The position (xNbA, yNbA) is set to equal to (x0, y0 - 1), and the derived procedure for the availability of the specified neighboring block is invoked, where the position (xCurr, yCurr) is set to equal to (x0, y0), the neighboring position (xNbY, yNbY) is set to equal to (xNbA, yNbA), checkPredModeY is set to equal to FALSE, and cIdx is taken as input, with the output assigned to availableA. The allocation of ctxInc is specified as follows, where condL and condA are specified in Table 19:
[0552] For the syntax elements alf_cross_component_cb_idc[x0][y0] and alf_cross_component_cr_idc[x0][y0]:
[0553] ctxInc = (condL && availableL) + (condA && availableA) + ctxSetIdx *3
[0554]
[0555] In one example, for alf_cross_component_cb_idc and / or alf_cross_component_cr_idc, context encoding can be performed on all bins, and all bins use the same context, i.e., each binIdx uses the same context. In some cases, subsets of binIdx are not applicable, for example, when NumCcAlfCbFilters is 3, binIdx>3 is not applicable to alf_cross_component_cb_idc[ ][ ], and when NumCcAlfCrFilters is 3, binIdx>3 is not applicable to alf_cross_component_cr_idc[ ][ ].
[0556] In one example, for alf_cross_component_cb_idc and / or alf_cross_component_cr_idc, context encoding can be performed on all bins, and each bin can use the same set of contexts as shown in Table 20. Regarding Table 20, it should be noted that in some cases, subsets of binIdx are not applicable; for example, when NumCcAlfCbFilters is 3, binIdx > 3 is not applicable to alf_cross_component_cb_idc[ ][ ], and when NumCcAlfCrFilters is 3, binIdx > 3 is not applicable to alf_cross_component_cr_idc[ ][ ].
[0557]
[0558] For Table 20, in one example, the context can be derived as follows:
[0559] The inputs to this process are the brightness position (x0, y0) of the current luminance block relative to the top-left luminance sample of the current image, the color component cIdx, the current coding quadtree depth cqtDepth, the dual-tree channel type chType, the width and height cbWidth and cbHeight of the current coding block in the luminance sample, and the variables allowSplitBtVer, allowSplitBtHor, allowSplitTtVer, allowSplitTtHor, and allowSplitQt derived from the coding tree semantics.
[0560] The output of this process is ctxInc.
[0561] The position (xNbL, yNbL) is set to equal to (x0-l, y0), and the derived procedure for the availability of the specified neighboring block is invoked, where the position (xCurr, yCurr) is set to equal to (x0, y0), the neighboring position (xNbY, yNbY) is set to equal to (xNbL, yNbL), checkPredModeY is set to equal to FALSE, and cIdx is taken as input, and the output is assigned to availableL.
[0562] The location (xNbA, yNbA) is set to equal to (x0, y0 - 1), and the derived procedure for the availability of the specified neighboring block is invoked, where the location (xCurr, yCurr) is set to equal to (x0, y0), the neighboring location (xNbY, yNbY) is set to equal to (xNbA, yNbA), checkPredModeY is set to equal to FALSE, and cIdx is taken as input, with the output assigned to availableA. The assignment of ctxInc is specified as follows, where condL and condA are specified in Table 21:
[0563] For the syntax elements alf_cross_component_cb_idc[x0][y0] and alf_cross_component_cr_idc[x0][y0]:
[0564] ctxInc = (condL && availableL) + (condA && availableA) + ctxSetIdx *3
[0565]
[0566] In one example, a video encoder represents an example of a device configured to perform the following operations: receive reconstructed sample data for the current component of video data, receive reconstructed sample data for one or more additional components of video data, derive a cross-component filter based on data associated with one or more additional components of video data, and apply the filter to the reconstructed sample data for the current component of video data based on the derived cross-component filter and the reconstructed sample data for one or more additional components of video data.
[0567] Figure 17 This is a block diagram illustrating an example of a video decoder configured to decode video data according to one or more techniques described herein. In one example, the video decoder 500 may be configured to reconstruct video data based on one or more of the techniques described above. That is, the video decoder 500 may operate in a manner reversible from the video encoder 200 described above. The video decoder 500 may be configured to perform intra-frame predictive decoding and inter-frame predictive decoding, and may therefore be referred to as a hybrid decoder. Figure 18In the example shown, the video decoder 500 includes an entropy decoding unit 502, an inverse quantization unit 504, an inverse transform processing unit 506, an intra-frame prediction processing unit 508, an inter-frame prediction processing unit 510, a summer 512, a filter unit 514, and a reference buffer 516. The video decoder 500 can be configured to decode video data in a manner consistent with a video coding system that implements one or more aspects of a video coding standard. It should be noted that although the exemplary video decoder 500 shown has different functional blocks, such illustrations are intended for descriptive purposes and do not limit the video decoder 500 and / or its sub-components to a particular hardware or software architecture. The functionality of the video decoder 500 can be implemented using any combination of hardware, firmware, and / or software implementations.
[0568] like Figure 17 As shown, the entropy decoding unit 502 receives an entropy-encoded bitstream. The entropy decoding unit 502 can be configured to decode the quantization syntax elements and quantization coefficients from the bitstream according to a process that is the inverse of the entropy encoding process. The entropy decoding unit 502 can be configured to perform entropy decoding according to any of the entropy encoding techniques described above. The entropy decoding unit 502 can parse the encoded bitstream in a manner consistent with video coding standards. The video decoder 500 can be configured to parse the encoded bitstream, wherein the encoded bitstream is generated based on the techniques described above.
[0569] Refer again Figure 17 The inverse quantization unit 504 receives quantization transform coefficients (i.e., bit values) and quantization parameter data from the entropy decoding unit 502. The quantization parameter data may include any and all combinations of the aforementioned incremental QP values and / or quantization group size values. The video decoder 500 and / or the inverse quantization unit 504 may be configured to determine the QP value for inverse quantization based on the value notified by the signal sent by the video encoder and / or through video attributes and / or encoding parameters. That is, the inverse quantization unit 504 may operate in a manner inverse of the aforementioned coefficient quantization unit 206. For example, the inverse quantization unit 504 may be configured to infer predetermined values, allowed quantization group sizes, etc., according to the aforementioned techniques. The inverse quantization unit 504 may be configured to apply inverse quantization. The inverse transform processing unit 506 may be configured to perform an inverse transform to generate reconstructed residual data. The techniques performed by the inverse quantization unit 504 and the inverse transform processing unit 506 may be similar to the techniques performed by the aforementioned inverse quantization / transform processing unit 208. The inverse transform processing unit 506 can be configured to apply inverse DCT, inverse DST, inverse integer transform, indivisible quadratic transform (NSST), or conceptually similar inverse transform procedures to transform the coefficients in order to generate residual blocks in the pixel domain. Furthermore, as mentioned above, whether a specific transform (or the type of specific transform) is performed can depend on the intra-frame prediction mode. Figure 17As shown, the reconstructed residual data can be provided to the summer 512. The summer 512 can add the reconstructed residual data to the predicted video block and generate reconstructed video data. The predicted video block can be determined based on the predicted video technique (i.e., intra-frame prediction and inter-frame prediction).
[0570] Intra-prediction processing unit 508 can be configured to receive intra-prediction syntax elements and retrieve predicted video blocks from reference buffer 516. Reference buffer 516 may include a memory device configured to store one or more video data frames. The intra-prediction syntax elements can identify intra-prediction modes, such as those described above. In one example, intra-prediction processing unit 508 may use one or more techniques from the intra-prediction coding techniques described herein to reconstruct the video block. Inter-prediction processing unit 510 can receive inter-prediction syntax elements and generate motion vectors to identify predicted blocks in one or more reference frames stored in reference buffer 516. Inter-prediction processing unit 510 may generate motion-compensated blocks, possibly performing interpolation based on interpolation filters. Identifiers for interpolation filters used for motion estimation with sub-pixel precision may be included in the syntax elements. Inter-prediction processing unit 510 may use interpolation filters to compute interpolated values for sub-integer pixels of the reference blocks.
[0571] Filter unit 514 can be configured to perform filtering on the reconstructed video data. For example, filter unit 514 can be configured to perform deblocking and / or SAO filtering, as described above with respect to filter unit 216. In the example, filter unit 514 may include cross-component filter unit 600 as described below. Furthermore, it should be noted that in some examples, filter unit 514 can be configured to perform dedicated arbitrary filtering (e.g., visual enhancement). Figure 17 As shown, the video decoder 500 can output reconstructed video blocks.
[0572] As mentioned above, Figure 7 An example of a cross-component filter unit that can be configured to encode video data according to one or more techniques of this disclosure is shown. Figure 18 An example of a cross-component filter unit configured to decode video data according to one or more techniques of this disclosure is shown. That is, the cross-component filter unit 600 can operate in a manner inverse of the cross-component filter unit 300. For example... Figure 18 As shown, the component filter unit 600 includes a filter determination unit 602 and a sample modification unit 604. The sample modification unit 604 can operate in a manner similar to that of the sample modification unit 304. That is, the sample modification unit 604 can perform filtering based on a derived filter, including one or more filters described herein. Figure 18As shown, the sample modification unit 604 can output the modified reconstructed block to the reference image buffer (i.e., as a loop filter) and output the modified reconstructed block to the output terminal (e.g., a display). The filter determination unit 602 can receive coding parameter information (e.g., intra-frame prediction) and available video block data when the current block is decoded, such as... Figure 18 As shown, at the video decoder, available video block data can include: cross-component reconstruction blocks and current component reconstruction blocks. However, as... Figure 18 As shown, the filter determination unit 602 can receive filter data. That is, filter data specifying the derived filter can be signaled to the filter determination unit 602. An example of such signaling has been described above. Therefore, the filter determination unit 602 can derive the filter to be used on the chroma reconstruction block based on video data, coding parameters, and / or filter data.
[0573] As described above, the cross-component filtering technique presented in this paper can typically be applied to each component of video data. Therefore, one or more combinations of video data components can be used to reduce the reconstruction error of one or more other components of the video data. Figures 19A to 19C This is a block diagram illustrating an example of a cross-component filter unit that can be configured to reduce reconstruction error according to one or more techniques of this disclosure. That is, Figures 19A to 19C An example of a loop filter that may be included in filter unit 514 is shown. Figures 19A to 19C In this context, the commonly numbered elements are as described above. The Cr cross-component filter unit 702 is an example of a filter unit configured to filter the Cr component based on the luminance component, Cb component, filter data, and encoding parameters. The luminance cross-component filter unit 704 is an example of a filter unit configured to filter the luminance component based on the luminance component, Cb component, Cr component, filter data, and encoding parameters. Therefore, the video decoder 500 represents an example of a device configured to perform the following operations: receiving reconstructed sample data for the current component of video data; receiving reconstructed sample data for one or more additional components of video data; deriving a cross-component filter based on data associated with one or more additional components of video data; and applying the filter to the reconstructed sample data for the current component of video data based on the derived cross-component filter and the reconstructed sample data for one or more additional components of video data.
[0574] In one or more examples, the functionality may be implemented by hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or code on or transmitted over a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a propagation medium that includes, for example, any medium facilitating the transfer of a computer program from one place to another according to a communication protocol. Thus, a computer-readable medium may generally correspond to: (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products may include computer-readable media.
[0575] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other medium that can be used to store required program code in the form of instructions or data structures and that is accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transient tangible storage media. As used herein, disks and optical discs include compact optical discs (CDs), laser discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, where disks typically copy data magnetically, while optical discs use lasers to copy data optically. Combinations of the above should also be included within the scope of computer-readable media.
[0576] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the term "processor" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Moreover, these techniques can be implemented entirely within one or more circuit or logic elements.
[0577] The techniques disclosed herein can be implemented in various devices or apparatuses, including wireless handsets, integrated circuits (ICs), or IC sets (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but they do not necessarily need to be implemented through different hardware units. Rather, as described above, various units can be combined in a codec hardware unit, or provided through an interoperable hardware unit comprising a collection of one or more processors as described above, combined with suitable software and / or firmware.
[0578] Furthermore, each functional block or feature of the base station equipment and terminal equipment used in each of the above embodiments can be implemented or executed by circuitry (typically one or more integrated circuits). Circuitry designed to perform the functions described in this specification may include general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, or combinations thereof. The general-purpose processor may be a microprocessor, or alternatively, it may be a conventional processor, controller, microcontroller, or state machine. The general-purpose processor or each of the above circuitry may be configured by digital circuitry or by analog circuitry. Furthermore, when advancements in semiconductor technology lead to the development of technologies for manufacturing integrated circuits that replace current integrated circuits, integrated circuits produced using such technologies can also be used.
[0579] Various examples have been described. These and other examples are within the scope of the following claims.
[0580] <Summary of the Invention>
[0581] In one example, a method for reducing reconstruction errors in video data is provided, the method comprising: receiving reconstruction sample data for a current component of the video data; receiving reconstruction sample data for one or more additional components of the video data; deriving a cross-component filter based on data associated with one or more additional components of the video data; and applying the filter to the reconstruction sample data for the current component of the video data based on the derived cross-component filter and the reconstruction sample data for one or more additional components of the video data.
[0582] In one example, the method also includes signaling information associated with the derived cross-component filter.
[0583] In one example, a method is provided in which deriving the cross component filter includes parsing the signaling to determine the cross component filter parameters.
[0584] In one example, a method is provided in which deriving a cross-component filter based on data associated with one or more additional components of the video data includes deriving the cross-component filter based on a known reconstruction error.
[0585] In one example, a method is provided to specify the cross component filter based on the filter coefficients.
[0586] In one example, an apparatus for encoding video data is provided, the apparatus including one or more processors configured to perform any and all combinations of these steps.
[0587] In one example, a device is provided that includes a video encoder.
[0588] In one example, a device is provided that includes a video decoder.
[0589] In one example, a system is provided that includes: a device including a video encoder; and the device including a video decoder.
[0590] In one example, an apparatus for encoding video data includes means for performing any and all combinations of steps.
[0591] In one example, a non-transitory computer-readable storage medium includes instructions stored thereon that, when executed, cause one or more processors of a device for encoding video data to perform any and all combinations of steps.
[0592] In one example, a method for filtering reconstructed video data is provided, the method comprising: inputting reconstructed luma component sample values; deriving filtered sample values by using cross-component filter coefficients and reconstructed luma component sample values prior to an adaptive loop filtering process; deriving refined values for chroma components by using the filtered sample values; and deriving refined chroma sample values by using the sample values of the chroma components and the sum of the refined values for the chroma components.
[0593] In one example, the method also includes performing a clipping function using the sum and bit depth.
[0594] In one example, the method further includes: decoding a first filter coefficient syntax element that specifies the absolute value of the cross-component filter coefficients for the Cb component in the adaptive loop filter data syntax structure, and decoding a second filter coefficient syntax element that specifies the sign of the cross-component filter coefficients for the Cb component in the adaptive loop filter data syntax structure.
[0595] In one example, the method further includes: decoding a first cross component syntax element that specifies whether to apply a cross component filter to the Cb color component; and decoding a second cross component syntax element that specifies whether to apply a cross component filter to the Cr color component.
[0596] In one example, a decoder for decoding encoded data is provided, the decoder including: a processor and memory associated with the processor; wherein the processor is configured to perform the following steps: inputting reconstructed luminance component sample values; deriving filtered sample values by using cross-component filter coefficients and the reconstructed luminance component sample values prior to an adaptive loop filtering process; deriving refined values for chrominance components using the filtered sample values; and deriving refined chrominance sample values by using the sample values of the chrominance components and the sum of the refined values for the chrominance components.
[0597] In one example, an encoder for encoding video data is provided, the encoder including: a processor and memory associated with the processor; wherein the processor is configured to perform the following steps: inputting reconstructed luma component sample values; deriving filtered sample values by using cross-component filter coefficients and the reconstructed luma component sample values prior to an adaptive loop filtering process; deriving refined values for the chroma component using the filtered sample values; and deriving refined chroma sample values by using the sample values of the chroma component and the sum of the refined values for the chroma component.
[0598] <Cross-reference>
[0599] This non-provisional patent application claims priority to provisional application 62 / 865,933, filed June 24, 2019; provisional application 62 / 870,752, filed July 4, 2019; and provisional application 62 / 886,891, filed August 14, 2019, pursuant to 35 USC § 119, the entire contents of which are incorporated herein by reference.
Claims
1. A device including one or more processors, said one or more processors being configured to: Receive slice header; Based on the second syntax element in the sequence parameter set, the first syntax element in the slice header is conditionally parsed, wherein... The first syntax element is a flag that specifies whether to enable cross component filtering for the color components; Based on the first syntax element, the third syntax element in the slice header is conditionally parsed, wherein the third syntax element specifies an adaptive parameter set identifier for the color component reference of the slice. Receive encoding tree unit syntax structure; Based on the first syntax element, the fourth syntax element in the coding tree unit syntax structure is conditionally parsed, wherein the fourth syntax element equal to 0 indicates that the cross component filter is not applied to the coding tree block of the color component, and when the fourth syntax element is not equal to 0, the fourth syntax element is related to the index of the cross component filter applied to the coding tree block of the color component.
2. The device according to claim 1, wherein, The device includes a video decoder.
3. A device including one or more processors, said one or more processors being configured to: Send a signal to notify the bit stream, where, The bit stream includes: Slice header and encoding tree syntax structure, The slice header conditionally includes the first syntax element based on the second syntax element in the sequence parameter set. The first syntax element is a flag that specifies whether to enable cross-component adaptive loop filter for the color components. The slice header conditionally includes a third syntax element based on the first syntax element. The third syntax element specifies the adaptive parameter set identifier for the color component reference of the slice; The encoding tree syntax structure conditionally includes a fourth syntax element based on the first syntax element. Wherein, the fourth syntax element being equal to 0 indicates that the cross component filter is not applied to the coding tree block of the color component, and when the fourth syntax element is not equal to 0, the fourth syntax element is related to the index of the cross component filter applied to the coding tree block of the color component.
4. The device according to claim 3, wherein, The device includes a video encoder.
5. A non-transitory computer-readable recording medium storing a bit stream, the bit stream comprising: Slice header and encoding tree syntax structure, The slice header conditionally includes the first syntax element based on the second syntax element in the sequence parameter set. The first syntax element is a flag that specifies whether to enable cross-component adaptive loop filter for the color components. The slice header conditionally includes a third syntax element based on the first syntax element. The third syntax element specifies the adaptive parameter set identifier for the color component reference of the slice; The encoding tree syntax structure conditionally includes a fourth syntax element based on the first syntax element. Wherein, the fourth syntax element being equal to 0 indicates that the cross component filter is not applied to the coding tree block of the color component, and when the fourth syntax element is not equal to 0, the fourth syntax element is related to the index of the cross component filter applied to the coding tree block of the color component.
6. A method for decoding video data, the method comprising: Receive slice header; Based on the second syntax element in the sequence parameter set, the first syntax element in the slice header is conditionally parsed, wherein the first syntax element is a flag and specifies whether to enable the cross component filter for the color component. Based on the first syntax element, the third syntax element in the slice header is conditionally parsed, wherein the third syntax element specifies an adaptive parameter set identifier for the color component reference of the slice. Receive encoding tree unit syntax structure; Based on the first syntax element, the fourth syntax element in the coding tree unit syntax structure is conditionally parsed, wherein the fourth syntax element equal to 0 indicates that the cross component filter is not applied to the coding tree block of the color component, and when the fourth syntax element is not equal to 0, the fourth syntax element is related to the index of the cross component filter applied to the coding tree block of the color component.
7. A method for encoding video data, the method comprising: Signaling a bitstream, wherein the bitstream includes: Slice header and encoding tree syntax structure, The slice header conditionally includes the first syntax element based on the second syntax element in the sequence parameter set. The first syntax element is a flag that specifies whether to enable cross-component adaptive loop filter for the color components. The slice header conditionally includes a third syntax element based on the first syntax element. The third syntax element specifies the adaptive parameter set identifier for the color component reference of the slice; The encoding tree syntax structure conditionally includes a fourth syntax element based on the first syntax element. Wherein, the fourth syntax element being equal to 0 indicates that the cross component filter is not applied to the coding tree block of the color component, and when the fourth syntax element is not equal to 0, the fourth syntax element is related to the index of the cross component filter applied to the coding tree block of the color component.