Cross-component adaptive loop filtering-based image coding apparatus and method

Cross-component adaptive loop filtering addresses the need for efficient image/video compression by filtering reconstructed chroma samples based on luma samples, enhancing compression efficiency and visual quality in high-resolution media.

JP2025168425APending Publication Date: 2025-11-07LG ELECTRONICS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025139458
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-08-29
Filing Date
2025-08-25
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality images/videos, particularly in immersive media formats like VR and AR, has led to a need for highly efficient image/video compression technologies to manage the increased data transmission and storage costs.

Method used

Implementing cross-component adaptive loop filtering (CCALF) procedures within the in-loop filtering process, allowing for efficient filtering of reconstructed chroma samples based on luma samples, and signaling information about filter coefficients and availability in the sequence parameter set (SPS) for improved image/video coding.

Benefits of technology

Enhances overall image/video compression efficiency and subjective/objective visual quality by adaptively applying ALF and CCALF on a picture, slice, and/or coding block basis, improving coding accuracy and filtering performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025168425000001_ABST
    Figure 2025168425000001_ABST
Patent Text Reader

Abstract

To provide a cross-component adaptive loop filtering-based image coding apparatus and method.SOLUTION: According to one embodiment of the present document, an in-loop filtering procedure in an image / video coding procedure includes a cross-component adaptive loop filtering procedure. CCALF according to the present embodiment increases the accuracy of in-loop filtering.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This document relates to a cross-component adaptive loop filtering based video coding apparatus and method. [Background technology]

[0002] In recent years, demand for high-resolution, high-quality images / videos such as 4K or 8K or higher UHD (Ultra High Definition) images / videos has been increasing in various fields. As the resolution and quality of image / video data increases, the amount of information or bits to be transmitted increases relatively compared to existing image / video data. Therefore, when transmitting image data using existing media such as wired or wireless broadband lines or storing image / video data using existing storage media, transmission costs and storage costs increase.

[0003] In addition, interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content and holograms has been increasing in recent years, and the broadcast of images / videos with different image characteristics from real images, such as game images, is increasing.

[0004] Therefore, there is a need for highly efficient image / video compression technology to effectively compress and transmit, store, and play back high-resolution, high-quality image / video information that has the various characteristics described above.

[0005] Cross-component adaptive loop filtering (CCALF) is a procedure performed within the in-loop filtering procedure to improve filtering accuracy, and there is ongoing discussion regarding the transmission of information used in the CCALF procedure. Summary of the Invention [Means for solving the problem]

[0006] According to one embodiment of the present document, a method and apparatus for improving the efficiency of image / video coding is provided.

[0007] According to one embodiment of the present document, an efficient method and apparatus for applying filtering is provided.

[0008] According to one embodiment of the present document, an efficient method and apparatus for applying ALF is provided.

[0009] According to one embodiment of this document, a filtering procedure can be performed on the reconstructed chroma samples based on the reconstructed luma samples.

[0010] According to one embodiment of this document, the filtered reconstructed chroma samples can be modified based on the reconstructed luma samples.

[0011] According to one embodiment of this document, information regarding the availability of CCALF can be signaled in the SPS.

[0012] According to one embodiment of this document, information about the values ​​of the cross-component filter coefficients can be derived from ALF data (general ALF data or CCALF data).

[0013] According to one embodiment of this document, APS identifier (ID) information including ALF data for deriving cross-component filter coefficients in a slice can be signaled.

[0014] According to one embodiment of this document, information regarding filter set indexes for CCALF can be signaled in units of CTUs (blocks).

[0015] According to one embodiment of the present document, there is provided a video / image decoding method performed by a decoding device.

[0016] According to one embodiment of the present document, a decoding device for video / image decoding is provided.

[0017] According to one embodiment of the present document, there is provided a video / image encoding method performed by an encoding device.

[0018] According to one embodiment of the present document, an encoding device for performing video / image encoding is provided.

[0019] According to one embodiment of the present document, there is provided a computer-readable digital storage medium having encoded video / image information generated by the video / image encoding method disclosed in at least one of the embodiments of the present document stored thereon.

[0020] According to one embodiment of the present document, a computer-readable digital storage medium is provided that stores encoded information or encoded video / image information that causes a decoding device to perform the video / image decoding method disclosed in at least one of the embodiments of the present document. [Effects of the Invention]

[0021] According to one embodiment of this document, the overall image / video compression efficiency can be improved.

[0022] According to one embodiment of this document, subjective / objective visual quality can be enhanced through efficient filtering.

[0023] According to one embodiment of this document, the ALF procedure can be performed efficiently and the filtering performance can be improved.

[0024] According to one embodiment of this document, the reconstructed chroma samples filtered based on the reconstructed luma samples are modified to improve the image quality and coding accuracy of the chroma components of the decoded picture.

[0025] According to one embodiment of this document, the CCALF procedure can be efficiently performed.

[0026] According to one embodiment of this document, ALF-related information can be signaled efficiently.

[0027] According to one embodiment of this document, CCALF-related information can be signaled efficiently.

[0028] According to one embodiment of this document, ALF and / or CCALF can be adaptively applied on a picture, slice, and / or coding block basis.

[0029] According to one embodiment of this document, when CCALF is used in an encoding and decoding method and apparatus for still or moving images, the filter coefficients for CCALF and the on / off transmission method in block or CTU units are improved, thereby increasing coding efficiency. [Brief explanation of the drawings]

[0030] [Figure 1] 1 illustrates schematically an example of a video / image coding system that can be applied to embodiments of the present document. [Figure 2] 1 is a diagram illustrating the configuration of a video / image encoding device that can be applied to an embodiment of the present document. [Figure 3] 1 is a diagram illustrating the configuration of a video / image decoding device that can be applied to an embodiment of the present document. [Figure 4] 1 shows an exemplary hierarchical structure for coded images / video. [Figure 5] 10 is a flowchart illustrating an intra-prediction-based block reconstruction method in a decoding device. [Figure 6] 10 is a flow chart illustrating an inter-prediction-based block reconstruction method in a decoding device. [Figure 7]Examples of ALF filter shapes are shown below. [Figure 8] 1 is a diagram illustrating a virtual boundary applied to a filtering procedure according to one embodiment of the present document. [Figure 9] An example of an ALF procedure using a virtual boundary is provided according to an embodiment of this document. [Figure 10] 1 is a diagram illustrating a cross-component adaptive loop filtering (CC-ALF (CCALF)) procedure according to one embodiment of the present document. [Figure 11] 1 illustrates an example of a video / image encoding method and associated components according to embodiment(s) of the present document; [Figure 12] 1 illustrates an example of a video / image encoding method and associated components according to embodiment(s) of the present document; [Figure 13] 1 illustrates an example of a schematic representation of a picture / video decoding method and related components according to embodiment(s) of the present document; [Figure 14] 1 illustrates an example of a schematic representation of a picture / video decoding method and related components according to embodiment(s) of the present document; [Figure 15] 1 illustrates an example of a content streaming system to which the embodiments disclosed herein can be applied. DETAILED DESCRIPTION OF THE INVENTION

[0031] The disclosure of this document may be modified in various ways and may have various embodiments. Specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit the disclosure to the specific embodiments. The terms used in this document are used merely to describe specific embodiments and are not intended to limit the technical ideas of the embodiments of this document. The singular expressions include the plural expressions unless the context clearly dictates otherwise. In this document, the terms "comprise" or "have" are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the document, and should be understood not to preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0032] Meanwhile, the components in the drawings described in this document are illustrated independently for the convenience of explaining the different characteristic functions, and do not mean that the components are realized by separate hardware or software. For example, two or more of the components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which the components are integrated and / or separated are also included within the scope of this document.

[0033] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document may be applied to methods disclosed in the versatile video coding (VVC) standard. The methods / embodiments disclosed in this document may also be applied to methods disclosed in the essential video coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation audio video coding standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268).

[0034] This document presents various embodiments relating to video / image coding, and unless otherwise stated, the embodiments may be implemented in combination with each other.

[0035] FIG. 1 illustrates schematically an example of a video / image coding system to which this document can be applied.

[0036] As shown in Figure 1, a video / image coding system may include a source device and a receiving device. The source device may transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or a network.

[0037] The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / video encoding device, and the decoding device may be referred to as a video / video decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may be a separate device or an external component.

[0038] A video source may acquire video / images through a video / image capture, synthesis, or generation process. A video source may include a video / image capture device and / or a video / image generation device. A video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device may include, for example, a computer, a tablet, a smartphone, etc., and may (electronically) generate video / images. For example, a virtual video / image may be generated via a computer, etc., in which case the video / image capture process may be replaced by a process in which related data is generated.

[0039] An encoding device can encode input video / images. The encoding device can perform a series of procedures such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0040] The transmitter may transmit the encoded video / image information or data output in the form of a bitstream to a receiver of a receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include elements for generating a media file in a predetermined file format and elements for transmission via a broadcast / communication network. The receiver may receive / extract the bitstream and transmit it to a decoding device.

[0041] The decoding device can decode the video / image by performing a series of steps such as inverse quantization, inverse transform, and prediction, which correspond to the operations of the encoding device.

[0042] The renderer can render the decoded video / image, and the rendered video / image can be displayed via a display unit.

[0043] In this document, video may refer to a collection of a series of images over time. A picture generally refers to a unit that represents one image at a specific time period, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile may include one or more coding tree units (CTUs). One picture may consist of one or more slices / tiles. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs, and the rectangular region has the same height as the picture, and the width can be specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the width of the picture.A tile scan may indicate a specific sequential ordering of CTUs partitioning a picture, in which the CTUs are ordered consecutively in a CTU raster scan in a tile, whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice includes an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that may be exclusively contained in a single NAL unit.

[0044] On the other hand, a picture can be divided into two or more sub-pictures, each of which is a rectangular region of one or more slices within a picture.

[0045] A pixel or a pel may refer to the smallest unit constituting one picture (or image). A term corresponding to a pixel may also be used: "sample." A sample may generally refer to a pixel or a pixel value, may refer to only a pixel / pixel value of a luma component, or may refer to only a pixel / pixel value of a chroma component.

[0046] A unit may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. The term unit may be used interchangeably with terms such as block or area. In general, an M×N block may include samples (or a sample array) consisting of M columns and N rows, or a set (or an array) of transform coefficients.

[0047] In this document, "A or B" can mean "only A," "only B," or "both A and B." Alternatively, in this document, "A or B" can be interpreted as "A and / or B." For example, in this document, "A, B or C" can mean "only A," "only B," "only C," or "any combination of A, B, and C."

[0048] A slash ( / ) or a comma (comma) used in this document can mean "and / or." For example, "A / B" can mean "A and / or B." This allows "A / B" to mean "only A," "only B," or "both A and B." For example, "A,B,C" can mean "A, B, or C."

[0049] In this document, "at least one of A and B" can mean "only A," "only B," or "both A and B." Also, in this document, the expressions "at least one of A or B" and "at least one of A and / or B" can be interpreted as "at least one of A and B."

[0050] Also, in this document, "at least one of A, B and C" can mean "only A," "only B," "only C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" can mean "at least one of A, B and C."

[0051] Furthermore, parentheses used in this document may mean "for example." Specifically, when "prediction (intra prediction)" is used, "intra prediction" is proposed as an example of "prediction." In other words, "prediction" in this document is not limited to "intra prediction," and "intra prediction" is proposed as an example of "prediction." Furthermore, when "prediction (i.e., intra prediction)" is used, "intra prediction" is proposed as an example of "prediction."

[0052] In this document, technical features individually described in one drawing may be embodied individually or simultaneously.

[0053] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. Hereinafter, the same components in the drawings will be designated by the same reference numerals, and duplicate descriptions of the same components will be omitted.

[0054] 2 is a diagram illustrating a schematic configuration of a video / image encoding device to which this document can be applied. Hereinafter, the term "video encoding device" may include a video encoding device.

[0055] As shown in FIG. 2, the encoding apparatus 200 may include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter predictor 221 and an intra predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. The image dividing unit 210, the predicting unit 220, the residual processing unit 230, the entropy encoding unit 240, the adding unit 250, and the filtering unit 260 may be configured as one or more hardware components (e.g., an encoder chipset or a processor) depending on the embodiment. Also, the memory 270 may include a decoded picture buffer (DPB) and may be configured as a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.

[0056] The image division unit 210 may divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quad-tree, binary-tree, ternary-tree (QTBTTT) structure. For example, one coding unit may be divided into multiple coding units of deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and then the binary tree structure and / or ternary structure may be applied. Alternatively, the binary tree structure may be applied first. The coding procedure according to the present disclosure may be performed based on a final coding unit that is not further divided. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to image characteristics, or the coding unit may be recursively divided into coding units of lower depths as needed, and a coding unit of an optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be divided or partitioned from the final coding unit.The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0057] The term "unit" may be used interchangeably with terms such as "block" or "area." In general, an M×N block may refer to a set of samples or transform coefficients consisting of M columns and N rows. A sample generally refers to a pixel or a pixel value, and may refer to only a pixel / pixel value of a luma component, or may refer to only a pixel / pixel value of a chroma component. A sample can be used as a term corresponding to one pixel or pel of a picture (or image).

[0058] The subtraction unit 231 may subtract a prediction signal (predicted block, prediction sample, or prediction sample array) output from the prediction unit 220 from an input video signal (original block, original sample, or original sample array) to generate a residual signal (residual block, residual sample, or residual sample array), and the generated residual signal is transmitted to the conversion unit 232. The prediction unit 220 may perform prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit 220 may determine whether intra prediction or inter prediction is applied in units of the current block or CU. The prediction unit may generate various information related to prediction, such as prediction mode information, and transmit the information to the entropy encoding unit 240, as will be described later in the description of each prediction mode. The prediction information may be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0059] The intra prediction unit 222 may predict the current block by referring to samples in the current picture. The referenced samples may be located adjacent to or distant from the current block depending on the prediction mode. Prediction modes in intra prediction may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, DC mode and planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the granularity of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 222 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.

[0060] The inter prediction unit 221 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (such as L0 prediction, L1 prediction, or Bi prediction). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (col CU), etc., and the reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter predictor 221 may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of a skip mode or a merge mode, the inter predictor 221 may use motion information of neighboring blocks as motion information of the current block. In the case of the skip mode, unlike the merge mode, a residual signal may not be transmitted.In the case of motion vector prediction (MVP) mode, the motion vector of the neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0061] The prediction unit 220 may generate a prediction signal based on various prediction methods, which will be described later. For example, the prediction unit may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as combined inter and intra prediction (CIIP). The prediction unit may also perform intra block copy (IBC) for predicting a block. The intra block copy may be used for content image / video coding, such as games, such as screen content coding (SCC). IBC basically performs prediction within a current picture, but may be similar to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described herein.

[0062] The prediction signal generated via the inter prediction unit 221 and / or the intra prediction unit 222 may be used to generate a reconstructed signal or a residual signal. The transform unit 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include a discrete cosine transform (DCT), a discrete sine transform (DST), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, GBT refers to a transform obtained from a graph representing relationship information between pixels. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. The transform process may be applied to pixel blocks having the same square size or non-square blocks of variable size.

[0063] The quantization unit 233 quantizes the transform coefficients and transmits the quantized signal to the entropy encoding unit 240. The entropy encoding unit 240 encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantization unit 233 may rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 may perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. In addition to the quantized transform coefficients, the entropy encoding unit 240 may also encode information required for video / image restoration (e.g., values ​​of syntax elements, etc.) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / picture information) may be transmitted or stored in the form of a bitstream in network abstraction layer (NAL) unit units. The video / picture information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / picture information may also include general constraint information. Signaling / transmitted information and / or syntax elements described later in this document may be encoded through the above-mentioned encoding procedure and included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcasting network and / or a communication network, and the digital storage medium can include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting the signal output from the entropy encoding unit 240 and / or a storage unit (not shown) for storing the signal can be configured as an internal / external element of the encoding apparatus 200, or the transmitter can be included in the entropy encoding unit 240.

[0064] The quantized transform coefficients output from the quantization unit 233 may be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) may be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transform unit 235. The adder 155 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample, or reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the prediction unit 220. When there is no residual for the current block, such as when skip mode is applied, the predicted block may be used as the reconstructed block. The generated reconstructed signal may be used for intra prediction of the next block to be processed in the current picture, and may also be used for inter prediction of the next picture after filtering, as described below.

[0065] Meanwhile, luma mapping with chroma scaling (LMCS) can be applied during picture encoding and / or reconstruction.

[0066] The filtering unit 260 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 260 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset (SAO), an adaptive loop filter, a bilateral filter, etc. The filtering unit 260 may generate various information related to filtering and transmit it to the entropy encoding unit 240, as will be described later in connection with each filtering method. The filtering information may be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0067] The modified reconstructed picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. When inter prediction is applied through this, the encoding apparatus can avoid prediction mismatch between the encoding apparatus 100 and the decoding apparatus, and can also improve coding efficiency.

[0068] The DPB of the memory 270 may store a modified reconstructed picture for use as a reference picture in the inter predictor 221. The memory 270 may store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter predictor 221 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 270 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 222.

[0069] FIG. 3 is a diagram illustrating the configuration of a video / image decoding device to which this document can be applied.

[0070] As shown in FIG. 3, the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter predictor 331 and an intra predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. The entropy decoding unit 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be implemented as a single hardware component (e.g., a decoder chipset or processor) according to an embodiment. The memory 360 may include a decoded picture buffer (DPB) or may be implemented as a digital storage medium. The hardware components may further include a memory 360 as an internal / external component.

[0071] When a bitstream including video / image information is input, the decoding apparatus 300 can reconstruct an image corresponding to the process in which the video / image information was processed by the encoding apparatus of FIG. 3. For example, the decoding apparatus 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding apparatus 300 can perform decoding using a processing unit applied by the encoding apparatus. Therefore, the processing unit for decoding can be, for example, a coding unit, and the coding unit can be divided from a coding tree unit or a maximal coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. The reconstructed image signal decoded and output by the decoding apparatus 300 can be played back via a playback device.

[0072] The decoding apparatus 300 may receive a signal output from the encoding apparatus of FIG. 3 in the form of a bitstream, and the received signal may be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 may parse the bitstream to derive information (e.g., video / video information) necessary for video restoration (or picture restoration). The video / video information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / video information may also include general constraint information. The decoding apparatus may further decode pictures based on the information on the parameter sets and / or the general constraint information. Signaled / received information and / or syntax elements, which will be described later in this document, may be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 may decode information in a bitstream based on a coding method such as Exponential Golomb Coding, CAVLC, or CABAC, and output values ​​of syntax elements required for image restoration and quantized values ​​of transform coefficients related to residuals. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using information on the syntax element to be decoded and decoding information on neighboring and current blocks or information on symbols / bins decoded in previous steps, predicts the occurrence probability of the bins based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values ​​of each syntax element.In this case, after determining a context model, the CABAC entropy decoding method can update the context model using information on the decoded symbol / bin for the context model of the next symbol / bin. Prediction-related information from the information decoded by the entropy decoding unit 310 is provided to the prediction unit 330, and residual information entropy-decoded by the entropy decoding unit 310, i.e., quantized transform coefficients and related parameter information, can be input to the inverse quantization unit 321. In addition, filtering-related information from the information decoded by the entropy decoding unit 310 can be provided to the filtering unit 350. Meanwhile, a receiving unit (not shown) that receives a signal output from the encoding apparatus can be further configured as an internal / external element of the decoding apparatus 300, or the receiving unit can be a component of the entropy decoding unit 310. Meanwhile, the decoding device according to this document may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, the inverse transform unit 322, the prediction unit 330, the addition unit 340, the filtering unit 350, and the memory 360.

[0073] The inverse quantization unit 321 may inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 321 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding apparatus. The inverse quantization unit 321 may inverse quantize the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.

[0074] The inverse transform unit 322 performs inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0075] The prediction unit may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on information about the prediction output from the entropy decoding unit 310, and may determine a specific intra / inter prediction mode.

[0076] The predictor may generate a prediction signal based on various prediction methods, which will be described later. For example, the predictor may apply intra prediction or inter prediction for prediction of a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as combined inter and intra prediction (CIIP). The predictor may also perform intra block copy (IBC) for prediction of a block. The intra block copy may be used for content image / video coding, such as games, such as screen content coding (SCC). IBC basically performs prediction within a current picture, but may be performed similarly to inter prediction in deriving a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described herein. Palette mode may be considered an example of intra coding or intra prediction.

[0077] The intra prediction unit 331 can predict a current block by referring to samples in a current picture. The referenced samples can be located adjacent to or distant from the current block depending on the prediction mode. In intra prediction, prediction modes can include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 can also determine a prediction mode to be applied to the current block using prediction modes applied to neighboring blocks.

[0078] The inter prediction unit 332 may derive a predicted block for the current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter prediction unit 332 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes, and the prediction information may include information indicating the inter prediction mode for the current block.

[0079] The adder 340 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit. When there is no residual for the current block, such as when skip mode is applied, the predicted block may be used as the reconstructed block.

[0080] The adder 340 may be referred to as a reconstruction unit or a reconstruction block generator. The generated reconstruction signal may be used for intra prediction of a next block to be processed in the current picture, may be output after filtering as described below, or may be used for inter prediction of a next picture.

[0081] Meanwhile, LMCS (luma mapping with chroma scaling) can be applied during picture decoding.

[0082] The filtering unit 350 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 350 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may transmit the modified reconstructed picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.

[0083] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter predictor 332. The memory 360 can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information can be transmitted to the inter predictor 332 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 360 can store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 331.

[0084] In this specification, the embodiments described for the prediction unit 330, inverse quantization unit 321, inverse transform unit 322, and filtering unit 350 of the decoding device 300 can also be applied identically or correspondingly to the prediction unit 220, inverse quantization unit 234, inverse transform unit 235, and filtering unit 260 of the encoding device 200, respectively.

[0085] As described above, prediction is performed to improve compression efficiency when performing video coding. Through this, a predicted block including predicted samples for a current block, which is a block to be coded, can be generated. Here, the predicted block includes predicted samples in the spatial domain (or pixel domain). The predicted block is derived in the same way by an encoding device and a decoding device. The encoding device signals information (residual information) regarding the residual between the original block and the predicted block, rather than the original sample values ​​of the original block, to the decoding device, thereby improving video coding efficiency. The decoding device derives a residual block including residual samples based on the residual information, combines the residual block with the predicted block to generate a reconstructed block including reconstructed samples, and generates a reconstructed picture including the reconstructed block.

[0086] The residual information may be generated through a transform and quantization procedure. For example, an encoding device may derive a residual block between the original block and the predicted block, perform a transform procedure on residual samples (residual sample array) included in the residual block to derive transform coefficients, and perform a quantization procedure on the transform coefficients to derive quantized transform coefficients, and then signal the related residual information (via a bitstream) to a decoding device. Here, the residual information may include information such as value information, position information, transform technique, transform kernel, and quantization parameter of the quantized transform coefficients. The decoding device may derive residual samples (or residual blocks) by performing an inverse quantization / inverse transform procedure based on the residual information. The decoding device may generate a reconstructed picture based on the predicted block and the residual block. The encoding device may also derive a residual block by inverse quantizing / inverse transforming quantized transform coefficients for reference for inter-prediction of a future picture, and generate a reconstructed picture based on the residual block.

[0087] In this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When the quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When the transform / inverse transform is omitted, the transform coefficients may also be referred to as coefficients or residual coefficients, or may still be referred to as transform coefficients for the sake of uniformity of expression.

[0088] In this document, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficient(s), and the information about the transform coefficient(s) may be signaled via residual coding syntax. Transform coefficients may be derived based on the residual information (or information about the transform coefficient(s), and scaled transform coefficients may be derived through an inverse transform (scaling) of the transform coefficients. Residual samples may be derived based on an inverse transform (transform) of the scaled transform coefficients. This may be similarly applied / expressed in other parts of this document.

[0089] A prediction unit of an encoding / decoding device may perform inter prediction on a block-by-block basis to derive prediction samples. Inter prediction may refer to prediction derived in a manner dependent on data elements (e.g., sample values ​​or motion information) of pictures other than the current picture. When inter prediction is applied to a current block, a predicted block (prediction sample array) for the current block may be derived based on a reference block (reference sample array) identified by a motion vector in a reference picture indicated by a reference picture index. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, motion information of the current block may be predicted in block, sub-block, or sample units based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction type (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). When inter prediction is applied, neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporally neighboring block may be the same or different. The temporally neighboring block may be called a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporally neighboring block may be called a collocated picture (colPic). For example, a candidate list of motion information may be constructed based on the neighboring blocks of the current block, and flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled.Inter prediction is performed based on various prediction modes. For example, in skip mode and merge mode, the motion information of the current block may be the same as the motion information of the selected neighboring block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected neighboring block is used as a motion vector predictor, and a motion vector difference may be signaled. In this case, the motion vector of the current block may be derived using the sum of the motion vector predictor and the motion vector difference.

[0090] The motion information may include L0 motion information and / or L1 motion information depending on the inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction may be referred to as an L0 motion vector or MVL0, and a motion vector in the L1 direction may be referred to as an L1 motion vector or MVL1. Prediction based on an L0 motion vector may be referred to as L0 prediction, prediction based on an L1 motion vector may be referred to as L1 prediction, and prediction based on both the L0 motion vector and the L1 motion vector may be referred to as bi-prediction (Bi) prediction. Here, an L0 motion vector may indicate a motion vector associated with a reference picture list L0 (L0), and an L1 motion vector may indicate a motion vector associated with a reference picture list L1 (L1). The reference picture list L0 may include pictures that are earlier in output order than the current picture as reference pictures, and the reference picture list L1 may include pictures that are later in output order than the current picture. The previous picture may be called a forward (reference) picture, and the subsequent picture may be called a backward (reference) picture. The reference picture list L0 may further include, as reference pictures, pictures that are subsequent to the current picture in output order. In this case, the previous picture may be indexed first in the reference picture list L0, and the subsequent picture may be indexed thereafter. The reference picture list L1 may further include, as reference pictures, pictures that are subsequent to the current picture in output order. In this case, the subsequent picture may be indexed first in the reference picture list L1, and the previous picture may be indexed thereafter. Here, the output order may correspond to a picture order count (POC) order.

[0091] FIG. 4 shows an exemplary hierarchical structure for coded images / video.

[0092] Referring to Figure 4, the coded image / video is divided into a VCL (video coding layer) that handles the image / video decoding process and itself, a lower system that transmits and stores the coded information, and a NAL (network abstraction layer) that exists between the VCL and the lower system and is responsible for network adaptation functions.

[0093] The VCL can generate VCL data including compressed video data (slice data), or can generate parameter sets including information such as a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), and a Video Parameter Set (VPS), or an SEI (Supplemental Enhancement Information) message that is additionally required for the video decoding process.

[0094] In NAL, NAL units can be generated by adding header information (NAL unit header) to RBSP (Raw Byte Sequence Payload) generated by VCL. In this case, RBSP refers to slice data, parameter sets, SEI messages, etc. generated by VCL. The NAL unit header can contain NAL unit type information identified by the RBSP data included in the corresponding NAL unit.

[0095] As shown in the drawing, NAL units can be classified into VCL NAL units and non-VCL NAL units according to the RBSP generated by the VCL. A VCL NAL unit can refer to a NAL unit containing information about a video (slice data), and a non-VCL NAL unit can refer to a NAL unit containing information necessary for decoding a video (parameter set or SEI message).

[0096] The VCL NAL unit and non-VCL NAL unit can be transmitted over a network with header information according to the data standard of the lower system. For example, the NAL unit can be transformed into a data format of a predetermined standard such as H.266 / VVC file format, RTP (Real-time Transport Protocol), TS (Transport Stream), etc. and transmitted over various networks.

[0097] As mentioned above, the NAL unit type of an NAL unit can be identified by the RBSP data structure included in the corresponding NAL unit, and information about such NAL unit type can be stored and signaled in the NAL unit header.

[0098] For example, NAL units can be broadly classified into VCL NAL unit types and non-VCL NAL unit types depending on whether the NAL unit contains information about a video (slice data). The VCL NAL unit types can be classified according to the nature and type of pictures included in the VCL NAL unit, and the non-VCL NAL unit types can be classified according to the type of parameter set.

[0099] The following is an example of a NAL unit type identified by the type of parameter set included in the non-VCL NAL unit type.

[0100] -APS (Adaptation Parameter Set) NAL unit: Type for NAL units including APS

[0101] -DPS (Decoding Parameter Set) NAL unit: Type for NAL units containing DPS

[0102] -VPS (Video Parameter Set) NAL unit: Type for NAL unit containing VPS

[0103] -SPS (Sequence Parameter Set) NAL unit: Type for NAL unit including SPS

[0104] -PPS (Picture Parameter Set) NAL unit: Type for NAL unit including PPS

[0105] -PH (Picture header) NAL unit: Type for NAL unit including PH

[0106] The above-mentioned NAL unit type has syntax information for the NAL unit type, and the syntax information can be stored in a NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified by the nal_unit_type value.

[0107] Meanwhile, as described above, one picture may include multiple slices, and one slice may include a slice header and slice data. In this case, one picture header may be added to multiple slices (slice header and slice data set) in one picture. The picture header (picture header syntax) may include information / parameters commonly applicable to the picture. In this document, slices may be combined with or replaced by tile groups. Also, in this document, slice headers may be combined with or replaced by type group headers.

[0108] The slice header (slice header syntax, slice header information) can include information / parameters commonly applicable to the slices. The APS (APS syntax) or PPS (PPS syntax) can include information / parameters commonly applicable to one or more slices or pictures. The SPS (SPS syntax) can include information / parameters commonly applicable to one or more sequences. The VPS (VPS syntax) can include information / parameters commonly applicable to multiple layers. The DPS (DPS syntax) can include information / parameters commonly applicable to video in general. The DPS can include information / parameters related to concatenation of coded video sequences (CVSs). In this document, the term "High level syntax (HLS)" refers to at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, and slice header syntax.

[0109] In this document, image / video information encoded from an encoding device to a decoding device and signaled in the form of a bitstream may include not only intra-picture partitioning-related information, intra / inter prediction information, residual information, in-loop filtering information, etc., but also information included in the slice header, information included in the picture header, information included in the APS, information included in the PPS, information included in the SPS, information included in the VPS, and / or information included in the DPS. In addition, the image / video information may further include information in a NAL unit header.

[0110] Meanwhile, in order to compensate for differences between an original image and a reconstructed image due to errors that occur during the compression encoding process, such as quantization, an in-loop filtering procedure may be performed on the reconstructed samples or pictures, as described above. As described above, the in-loop filtering may be performed in a filter unit of an encoding device and a filter unit of a decoding device, and a deblocking filter, SAO, and / or adaptive loop filter (ALF) may be applied. For example, the ALF procedure may be performed after the deblocking filtering procedure and / or SAO procedure are completed. However, in this case, the deblocking filtering procedure and / or SAO procedure may be omitted.

[0111] A detailed description of picture reconstruction and filtering will be given below. In image / video coding, reconstructed blocks may be generated based on intra prediction / inter prediction for each block, and a reconstructed picture including the reconstructed blocks may be generated. If a current picture / slice is an I picture / slice, blocks included in the current picture / slice may be reconstructed based only on intra prediction. On the other hand, if the current picture / slice is a P or B picture / slice, blocks included in the current picture / slice may be reconstructed based on intra prediction or inter prediction. In this case, intra prediction may be applied to some blocks in the current picture / slice, and inter prediction may be applied to the remaining blocks.

[0112] Intra prediction may refer to a prediction that generates prediction samples for a current block based on reference samples in a picture to which the current block belongs (hereinafter referred to as the current picture). When intra prediction is applied to the current block, neighboring reference samples used for intra prediction of the current block may be derived. The neighboring reference samples of the current block may include samples adjacent to the left boundary and bottom-left neighboring samples of the current block having a size of nW×nH, a total of 2×nH samples adjacent to the top boundary and top-right neighboring samples of the current block, and one sample adjacent to the top-left neighboring sample of the current block. Alternatively, the neighboring reference samples of the current block may include upper neighboring samples of multiple columns and left neighboring samples of multiple rows. In addition, the neighboring reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block, which has a size of nW x nH, a total of nW samples adjacent to the bottom boundary of the current block, and one sample adjacent to the bottom-right of the current block.

[0113] However, some of the neighboring reference samples of the current block may not yet be decoded or may not be available. In this case, the decoder may construct neighboring reference samples to be used for prediction by substituting unavailable samples as available samples, or may construct neighboring reference samples to be used for prediction through interpolation of available samples.

[0114] When neighboring reference samples are derived, (i) a predicted sample may be derived based on an average or interpolation of neighboring reference samples of the current block, or (ii) the predicted sample may be derived based on a reference sample of the neighboring reference samples of the current block that exists in a specific (prediction) direction relative to the predicted sample. (i) This may be referred to as a non-directional mode or a non-angular mode, and (ii) this may be referred to as a directional mode or an angular mode. Furthermore, the predicted sample may be generated by interpolating the first and second neighboring samples, which are located in the opposite direction to the prediction direction of the intra-prediction mode of the current block, based on the predicted sample of the current block among the neighboring reference samples. This may be referred to as linear interpolation intra-prediction (LIP). Alternatively, a chroma predicted sample may be generated based on a luma sample using a linear model. This may be referred to as LM mode. In addition, a temporary predicted sample of the current block may be derived based on a filtered neighboring reference sample, and a predicted sample of the current block may be derived by weighting the temporary predicted sample and at least one reference sample derived according to the intra prediction mode from the existing neighboring reference samples, i.e., non-filtered neighboring reference samples. The above-mentioned case may be called Position Dependent Intra Prediction (PDPC). In addition, intra prediction coding may be performed by selecting a reference sample line with the highest prediction accuracy from multiple neighboring reference sample lines of the current block, deriving a predicted sample using a reference sample located in the prediction direction of the corresponding line, and signaling the reference sample line used to a decoding device.The above-described case may be referred to as multi-reference line (MRL) intra prediction or MRL-based intra prediction. Furthermore, the current block may be divided into vertical or horizontal sub-partitions, and intra prediction may be performed based on the same intra prediction mode. Neighboring reference samples may be derived and used for each sub-partition. That is, in this case, the intra prediction mode for the current block is equally applied to the sub-partitions, and neighboring reference samples may be derived and used for each sub-partition, thereby improving intra prediction performance in some cases. This prediction method may be referred to as intra sub-partitions (ISP) or ISP-based intra prediction. The above-described intra prediction method may be referred to as an intra prediction type, distinguished from the intra prediction modes in Tables 1 and 2. The intra prediction type may be referred to by various terms, such as an intra prediction technique or an additional intra prediction mode. For example, the intra prediction type (or additional intra prediction mode, etc.) may include at least one of the LIP, PDPC, MRL, and ISP. A general intra prediction method excluding specific intra prediction types such as LIP, PDPC, MRL, and ISP may be referred to as a normal intra prediction type. The normal intra prediction type may be generally applied when the specific intra prediction types are not applied, and prediction may be performed based on the intra prediction mode. Meanwhile, post-processing filtering may be performed on the derived prediction samples as needed.

[0115] Specifically, the intra prediction procedure may include an intra prediction mode / type determination step, a neighboring reference sample derivation step, and an intra prediction mode / type-based prediction sample derivation step. If necessary, a post-processing filtering step may be performed on the derived prediction samples.

[0116] The following describes intra prediction in an encoding device. The encoding device performs intra prediction on a current block. The encoding device may derive an intra prediction mode for the current block, derive neighboring reference samples for the current block, and generate predicted samples within the current block based on the intra prediction mode and the neighboring reference samples. Here, the intra prediction mode determination, neighboring reference sample derivation, and predicted sample generation procedures may be performed simultaneously, or one procedure may be performed before the other. For example, the intra prediction unit 222 of the encoding device may include a prediction mode / type determination unit, a reference sample derivation unit, and a predicted sample derivation unit. The prediction mode / type determination unit may determine the intra prediction mode / type for the current block, the reference sample derivation unit may derive neighboring reference samples for the current block, and the predicted sample derivation unit may derive motion samples for the current block. Meanwhile, although not shown, if a predicted sample filtering procedure (described below) is performed, the intra prediction unit 222 may further include a predicted sample filter unit (not shown). The encoding apparatus may determine a mode to be applied to the current block from among a plurality of intra prediction modes, and may determine an optimal intra prediction mode for the current block by comparing RD costs for the intra prediction modes.

[0117] Meanwhile, the encoding apparatus may also perform a prediction sample filtering procedure, which may be called post-filtering. The prediction sample filtering procedure may filter some or all of the prediction samples. In some cases, the prediction sample filtering procedure may be omitted.

[0118] The encoding apparatus may derive residual samples for the current block based on the predicted samples, and may derive the residual samples by comparing the predicted samples with original samples of the current block based on a phase.

[0119] The encoding apparatus transforms / quantizes the residual samples to derive quantized transform coefficients, and then performs inverse quantization / inverse transform on the quantized transform coefficients to derive (modified) residual samples. The reason for performing inverse quantization / inverse transform after transform / quantization is to derive residual samples that are the same as the residual samples derived by the decoding apparatus, as described above.

[0120] The encoding apparatus may generate a reconstructed block including reconstructed samples for the current block based on the predicted samples and the (corrected) residual samples, and may generate a reconstructed picture for the current picture based on the reconstructed block.

[0121] As described above, the encoding apparatus may encode video information including prediction information related to the intra prediction (e.g., prediction mode information indicating a prediction mode) and residual information related to the intra and residual samples, and output the encoded video information in the form of a bitstream. The residual information may include a residual coding syntax. The encoding apparatus may transform / quantize the residual samples to derive quantized transform coefficients. The residual information may include information on the quantized transform coefficients.

[0122] 5 is a flow chart illustrating an intra-prediction-based block reconstruction method in a decoding device. The method of FIG. 5 may include steps S500, S510, S520, S530, and S540. The decoding device may perform operations corresponding to those performed in the encoding device.

[0123] Steps S500 to S520 may be performed by an intra prediction unit 331 of a decoding device, and the prediction information of S500 and the residual information of S530 may be obtained from a bitstream by an entropy decoding unit 310 of the decoding device. The residual processing unit 320 of the decoding device may derive residual samples for the current block based on the residual information. Specifically, the inverse quantization unit 321 of the residual processing unit 320 may perform inverse quantization on quantized transform coefficients derived based on the residual information to derive transform coefficients, and the inverse transform unit 322 of the residual processing unit may perform inverse transform on the transform coefficients to derive residual samples for the current block. Step S540 may be performed by an adder 340 or a reconstruction unit of the decoding device.

[0124] Specifically, the decoding apparatus may derive an intra prediction mode for a current block based on received prediction mode information (S500). The decoding apparatus may derive neighboring reference samples of the current block (S510). The decoding apparatus may generate prediction samples within the current block based on the intra prediction mode and the neighboring reference samples (S520). In this case, the decoding apparatus may perform a prediction sample filtering procedure. The prediction sample filtering may be referred to as post-filtering. Some or all of the prediction samples may be filtered by the prediction sample filtering procedure. In some cases, the prediction sample filtering procedure may be omitted.

[0125] The decoding apparatus generates residual samples for the current block based on the received residual information (S530). The decoding apparatus generates reconstructed samples for the current block based on the predicted samples and the residual samples, and may derive a reconstructed block including the reconstructed samples (S540). A reconstructed picture for the current picture may be generated based on the reconstructed block.

[0126] Here, the intra prediction unit 331 of the decoding device may include a prediction mode / type determination unit, a reference sample derivation unit, and a prediction sample derivation unit, where the prediction mode / type determination unit determines an intra prediction mode for the current block based on prediction mode information acquired from the entropy decoding unit 310 of the decoding device, the reference sample derivation unit derives neighboring reference samples of the current block, and the prediction sample derivation unit derives prediction samples of the current block. Meanwhile, although not shown, if the above-mentioned prediction sample filtering procedure is performed, the intra prediction unit 331 may further include a prediction sample filter unit (not shown).

[0127] The prediction information may include intra prediction mode information and / or intra prediction type information. The intra prediction mode information may include, for example, flag information (e.g., intra_luma_mpm_flag) indicating whether a most probable mode (MPM) or a remaining mode is applied to the current block. If the MPM is applied to the current block, the prediction mode information may further include index information (e.g., intra_luma_mpm_idx) indicating one of the intra prediction mode candidates (MPM candidates). The intra prediction mode candidates (MPM candidates) may be configured as an MPM candidate list or an MPM list. If the MPM is not applied to the current block, the intra prediction mode information may further include remaining mode information (e.g., intra_luma_mpm_remainder) indicating one of the remaining intra prediction modes excluding the intra prediction mode candidates (MPM candidates). A decoding apparatus may determine the intra prediction mode of the current block based on the intra prediction mode information. A separate MPM list can be configured for the above-mentioned MIP.

[0128] Also, the intra prediction type information may be implemented in various forms. For example, the intra prediction type information may include intra prediction type index information indicating one of the intra prediction types. For another example, the intra prediction type information may include at least one of reference sample line information (e.g., intra_luma_ref_idx) indicating whether the MRL is applied to the current block and, if so, which reference sample line is used, ISP flag information (e.g., intra_subpartitions_mode_flag) indicating whether the ISP is applied to the current block, ISP type information (e.g., intra_subpartitions_split_flag) indicating a subpartition split type if the ISP is applied, flag information indicating whether PDCP is applied, or flag information indicating whether LIP is applied. Also, the intra prediction type information may include an MIP flag indicating whether MIP is applied to the current block.

[0129] The intra prediction mode information and / or the intra prediction type information may be encoded / decoded using a coding method described in this document. For example, the intra prediction mode information and / or the intra prediction type information may be encoded / decoded using entropy coding (e.g., CABAC or CAVLC coding) based on a truncated (rice) binary code.

[0130] A prediction unit of an encoding / decoding device may perform inter prediction on a block-by-block basis to derive prediction samples. Inter prediction may refer to a prediction derived in a manner that is dependent on data elements (e.g., sample values ​​or motion information) of picture(s) other than the current picture. When inter prediction is applied to a current block, a predicted block (prediction sample array) for the current block may be derived based on a reference block (reference sample array) identified by a motion vector on a reference picture indicated by a reference picture index. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, motion information of the current block may be predicted on a block, sub-block, or sample-by-block basis based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction type information (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). When inter-prediction is applied, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring blocks may be the same or different. The temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture including the temporal neighboring blocks may be called collocated pictures (colPics).For example, a motion information candidate list may be constructed based on neighboring blocks of the current block, and flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled. Inter prediction may be performed based on various prediction modes. For example, in skip mode and merge mode, the motion information of the current block is the same as that of the selected neighboring block. In skip mode, unlike merge mode, no residual signal is transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected neighboring block is used as a motion vector predictor, and a motion vector difference may be signaled. In this case, the motion vector of the current block may be derived using the sum of the motion vector predictor and the motion vector difference.

[0131] An inter-prediction procedure performed by an encoding apparatus will be described below. The encoding apparatus performs inter-prediction on a current block. The encoding apparatus may derive an inter-prediction mode and motion information for the current block and generate a predicted sample for the current block. Here, the inter-prediction mode determination, motion information derivation, and predicted sample generation procedures may be performed simultaneously, or one procedure may be performed before the other procedures. For example, the inter-prediction unit 221 of the encoding apparatus may include a prediction mode determination unit, a motion information derivation unit, and a predicted sample derivation unit. The prediction mode determination unit may determine an inter-prediction mode for the current block, the motion information derivation unit may derive motion information for the current block, and the predicted sample derivation unit may derive a motion sample for the current block. For example, the inter-prediction unit 221 of the encoding apparatus may search for a block similar to the current block within a certain region (search region) of a reference picture through motion estimation, and derive a reference block whose difference from the current block is minimal or equal to or less than a certain criterion. Based on this, a reference picture index indicating a reference picture in which the reference block is located can be derived, and a motion vector can be derived based on a position difference between the reference block and the current block. The encoding apparatus can determine a mode to be applied to the current block from various prediction modes. The encoding apparatus can compare RD costs for the various prediction modes to determine an optimal prediction mode for the current block.

[0132] For example, when a skip mode or a merge mode is applied to the current block, the encoding apparatus may construct a merge candidate list (described below) and derive a reference block whose difference from the current block is minimum or equal to or less than a certain criterion among reference blocks indicated by merge candidates included in the merge candidate list. In this case, a merge candidate associated with the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and signaled to the decoding apparatus. Motion information of the current block may be derived using motion information of the selected merge candidate.

[0133] As another example, when the (A)MVP mode is applied to the current block, the encoding apparatus may construct an (A)MVP candidate list (described below) and use a motion vector of an MVP (motion vector predictor) candidate selected from the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. In this case, for example, a motion vector pointing to a reference block derived by the motion estimation described above may be used as the motion vector of the current block, and the MVP candidate having the smallest difference from the motion vector of the current block among the MVP candidates may become the selected MVP candidate. A motion vector difference (MVD), which is the difference obtained by subtracting the MVP from the motion vector of the current block, may be derived. In this case, information regarding the MVD may be signaled to the decoding apparatus. Furthermore, when the (A)MVP mode is applied, the value of the reference picture index may be configured as reference picture index information and separately signaled to the decoding apparatus.

[0134] The encoding apparatus may derive residual samples based on the predicted samples, or may derive the residual samples by comparing original samples of the current block with the predicted samples.

[0135] The encoding apparatus transforms / quantizes the residual samples to derive quantized transform coefficients, and then performs inverse quantization / inverse transform on the quantized transform coefficients to derive (modified) residual samples. The reason for performing inverse quantization / inverse transform after transform / quantization is to derive residual samples that are the same as the residual samples derived by the decoding apparatus, as described above.

[0136] The encoding apparatus may generate a reconstructed block including reconstructed samples for the current block based on the predicted samples and the (corrected) residual samples, and may generate a reconstructed picture for the current picture based on the reconstructed block.

[0137] Although not shown, as described above, the encoding apparatus may encode video information including prediction information and residual information. The encoding apparatus may output the encoded video information in the form of a bitstream. The prediction information is information related to the prediction procedure and may include prediction mode information (e.g., a skip flag, a merge flag, or a mode index) and information on motion information. The information on the motion information may include candidate selection information (e.g., a merge index, an MVP flag, or an MVP index) for deriving a motion vector. The information on the motion information may also include the above-mentioned information on MVD and / or reference picture index information. The information on the motion information may also include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information on the residual sample. The residual information may include information on quantized transform coefficients for the residual sample.

[0138] The output bitstream can be stored in a (digital) storage medium and then transmitted to the decoding device, or can be transmitted to the decoding device via a network.

[0139] 6 is a flow chart illustrating an inter-prediction-based block reconstruction method in a decoding device. The method of FIG. 6 may include steps S600, S610, S620, S630, and S640. The decoding device may perform operations corresponding to those performed by the encoding device.

[0140] Steps S600 to S620 may be performed by an inter-prediction unit 332 of a decoding device, and the prediction information of S600 and the residual information of S630 may be obtained from a bitstream by an entropy decoding unit 310 of the decoding device. The residual processing unit 320 of the decoding device may derive residual samples for the current block based on the residual information. Specifically, the inverse quantization unit 321 of the residual processing unit 320 may perform inverse quantization on quantized transform coefficients derived based on the residual information to derive transform coefficients, and the inverse transform unit 322 of the residual processing unit may perform inverse transform on the transform coefficients to derive residual samples for the current block. Step S640 may be performed by an adder 340 or a reconstruction unit of the decoding device.

[0141] Specifically, the decoding apparatus may determine a prediction mode for the current block based on received prediction information (S600). The decoding apparatus may determine which inter-prediction mode is applied to the current block based on prediction mode information in the prediction information.

[0142] For example, the merge flag may determine whether the merge mode is applied to the current block or whether the (A)MVP mode is selected. Alternatively, the mode index may determine whether one of various inter prediction mode candidates is selected. The inter prediction mode candidates may include skip mode, merge mode, and / or (A)MVP mode, or may include various inter prediction modes described below.

[0143] The decoding apparatus derives motion information of the current block based on the determined inter prediction mode (S610). For example, when a skip mode or a merge mode is applied to the current block, the decoding apparatus may construct a merge candidate list (described below) and select one merge candidate from among the merge candidates included in the merge candidate list. The selection may be performed based on the selection information (merge index) described above. Motion information of the selected merge candidate may be derived for the current block using motion information of the selected merge candidate. The motion information of the selected merge candidate may be used as motion information of the current block.

[0144] As another example, when the (A)MVP mode is applied to the current block, the decoding apparatus may construct an (A)MVP candidate list (described below) and use a motion vector predictor (MVP) selected from among the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. The selection may be performed based on the selection information (MVP flag or MVP index). In this case, the MVD of the current block may be derived based on information related to the MVD, and the motion vector of the current block may be derived based on the MVP of the current block and the MVD. Furthermore, the decoding apparatus may derive a reference picture index of the current block based on the reference picture index information. A picture pointed to by the reference picture index in the reference picture list for the current block may be derived as a reference picture referenced for inter-prediction of the current block.

[0145] On the other hand, as will be described later, the motion information of the current block may be derived without constructing a candidate list, and in this case, the motion information of the current block may be derived according to a procedure disclosed in the prediction mode section, which will be described later. In this case, the candidate list construction as described above may be omitted.

[0146] The decoding apparatus may generate predictive samples for the current block based on the motion information of the current block (S620). In this case, the reference picture may be derived based on a reference picture index of the current block, and the predictive samples of the current block may be derived using samples of a reference block to which the motion vector of the current block points on the reference picture. In this case, as described below, a predictive sample filtering procedure may be further performed on all or some of the predictive samples of the current block, depending on the circumstances.

[0147] For example, the inter prediction unit 332 of the decoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit, and may determine a prediction mode for the current block based on prediction mode information received by the prediction mode determination unit, derive motion information (motion vector and / or reference picture index, etc.) of the current block based on information regarding the motion information received by the motion information derivation unit, and derive a prediction sample of the current block by the prediction sample derivation unit.

[0148] The decoding apparatus generates residual samples for the current block based on the received residual information (S630). The decoding apparatus generates reconstructed samples for the current block based on the predicted samples and the residual samples, and may derive a reconstructed block including the reconstructed samples (S640). A reconstructed picture for the current picture may be generated based on the reconstructed block.

[0149] Various inter-prediction modes can be used to predict a current block in a picture. For example, various modes can be used, such as merge mode, skip mode, motion vector prediction (MVP) mode, affine mode, sub-block merge mode, and merge with MVD (MMVD) mode. Decoder side motion vector refinement (DMVR) mode, adaptive motion vector resolution (AMVR) mode, bi-prediction with CU-level weight (BCW), bi-directional optical flow (BDOF), etc. can be used in addition to or instead of these modes as additional modes. Affine mode may also be referred to as affine motion prediction mode. MVP mode may also be referred to as advanced motion vector prediction mode. In this document, motion information candidates derived by some modes and / or some modes may be included as candidates for motion information in other modes. For example, an HMVP candidate may be added as a merge candidate in the merge / skip mode, or as an MVP candidate in the MVP mode.

[0150] Prediction mode information indicating the inter prediction mode of the current block may be signaled from the encoding apparatus to the decoding apparatus. The prediction mode information may be included in a bitstream and received by the decoding apparatus. The prediction mode information may include index information indicating one of multiple candidate modes. Alternatively, the inter prediction mode may be indicated through hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, a skip flag may be signaled to indicate whether the skip mode is applied, and if the skip mode is not applied, a merge flag may be signaled to indicate whether the merge mode is applied, and if the merge mode is not applied, an MVP mode may be applied, or a flag for additional distinction may be further signaled. The affine mode may be signaled as an independent mode or as a mode dependent on the merge mode, MVP mode, etc. For example, the affine mode may include affine merge mode and affine MVP mode.

[0151] Meanwhile, information indicating whether the above-mentioned list0 (L0) prediction, list1 (L1) prediction, or bi-prediction prediction is used for the current block (current coding unit) may be signaled for the current block. This information may be called motion prediction direction information, inter-prediction direction information, or inter-prediction indication information, and may be configured / encoded / signaled in the form of, for example, an inter_pred_idc syntax element. That is, the inter_pred_idc syntax element may indicate whether the above-mentioned list0 (L0) prediction, list1 (L1) prediction, or bi-prediction is used for the current block (current coding unit). In this document, for convenience of explanation, the inter-prediction type (L0 prediction, L1 prediction, or BI prediction) indicated by the inter_pred_idc syntax element may be referred to as a motion prediction direction. L0 prediction may be represented as pred_L0, L1 prediction as pred_L1, and bi-prediction as pred_BI. For example, the prediction type can be determined according to the value of the syntax element inter_pred_idc as shown in the following table.

[0152] [Table 1]

[0153] As described above, one picture may include one or more slices. A slice may have one of slice types, including an I slice (intra slice), a P slice (predictive slice), and a B slice (bi-predictive slice). The slice type may be indicated based on slice type information. For blocks in an I slice, inter prediction may not be used for prediction, and only intra prediction may be used. Of course, even in this case, original sample values ​​may be coded and signaled without prediction. For blocks in a P slice, intra prediction or inter prediction may be used, and when inter prediction is used, only uni prediction may be used. On the other hand, for blocks in a B slice, intra prediction or inter prediction may be used, and when inter prediction is used, up to bi prediction may be used.

[0154] L0 and L1 may include reference pictures encoded / decoded earlier than the current picture. For example, L0 may include reference pictures earlier and / or later than the current picture in POC order, and L1 may include reference pictures later and / or earlier than the current picture in POC order. In this case, L0 may be assigned a reference picture index lower than the reference picture earlier than the current picture in POC order, and L1 may be assigned a reference picture index lower than the reference picture later than the current picture in POC order. For B slices, bi-prediction may be applied, and in this case, unidirectional bi-prediction or bi-directional bi-prediction may be applied. Bi-directional bi-prediction may be referred to as true bi-prediction.

[0155] As described above, a residual block (residual sample) may be derived based on a predicted block (prediction sample) derived through prediction in the encoding stage, and residual information may be generated from the residual sample through transformation / quantization. The residual information may include information on quantized transform coefficients. The residual information may be included in video / image information, which may be encoded and transmitted to a decoding device in the form of a bitstream. A decoding device may obtain the residual information from the bitstream and derive residual samples based on the residual information. Specifically, the decoding device may derive quantized transform coefficients based on the residual information and derive residual blocks (residual samples) through inverse quantization / inverse transform procedures.

[0156] Meanwhile, at least one of the (inverse) transform and / or (inverse) quantization steps can be omitted.

[0157] An in-loop filtering procedure performed for a reconstructed picture will now be described. Modified reconstructed samples, blocks, or pictures (or modified filtered samples, blocks, or pictures) can be generated through the in-loop filtering procedure. The modified (modified filtered) reconstructed picture can be output as a decoded picture by a decoding device or stored in a decoded picture buffer or memory of an encoding / decoding device and subsequently used as a reference picture in an inter-prediction procedure during picture encoding / decoding. As described above, the in-loop filtering procedure can include a deblocking filtering procedure, a sample adaptive offset (SAO) procedure, and / or an adaptive loop filter (ALF) procedure. In this case, one or some of the deblocking filtering procedure, sample adaptive offset (SAO), adaptive loop filter (ALF), and bilateral filter procedures can be applied sequentially, or all of them can be applied sequentially. For example, the SAO procedure can be performed after the deblocking filtering procedure is applied to the reconstructed picture. Or, for example, the ALF procedure can be performed after a deblocking filtering procedure is applied to the reconstructed picture, which can also be performed in the encoding device.

[0158] Deblocking filtering is a filtering technique that removes distortions that occur at boundaries between blocks in a reconstructed picture. The deblocking filtering procedure may, for example, derive a target boundary in a reconstructed picture, determine a boundary strength (bS) for the target boundary, and perform deblocking filtering on the target boundary based on the bS. The bS may be determined based on the prediction modes of two blocks adjacent to the target boundary, the motion vector difference, whether the reference pictures are the same, whether there are significant non-zero coefficients, etc.

[0159] SAO is a method for compensating for an offset difference between a reconstructed picture and an original picture on a sample-by-sample basis, and may be applied based on a type such as a band offset or an edge offset. According to SAO, samples are classified into different categories according to each SAO type, and an offset value may be added to each sample based on the category. Filtering information for SAO may include information on whether SAO is applicable, SAO type information, SAO offset value information, etc. SAO may also be applied to a reconstructed picture after the deblocking filtering is applied.

[0160] An adaptive loop filter (ALF) is a technique for filtering a reconstructed picture on a sample-by-sample basis based on filter coefficients according to a filter shape. An encoding apparatus can determine whether to apply an ALF, the ALF shape, and / or ALF filtering coefficients by comparing a reconstructed picture with an original picture, and signal the determination to a decoding apparatus. That is, filtering information for the ALF can include information on whether to apply an ALF, ALF filter shape information, ALF filtering coefficient information, etc. The ALF can also be applied to a reconstructed picture after the deblocking filtering has been applied.

[0161] FIG. 7 shows an example of an ALF filter shape.

[0162] FIG. 7(a) shows a 7×7 diamond filter shape, and FIG. 7(b) shows a 5×5 diamond filter shape. In FIG. 7, Cn in the filter shape indicates a filter coefficient. When n is the same in Cn, this indicates that the same filter coefficient can be assigned. In this document, the position and / or unit to which filter coefficients are assigned according to the ALF filter shape may be referred to as a filter tab. In this case, one filter coefficient may be assigned to each filter tab, and the arrangement of the filter tabs may correspond to the filter shape. A filter tab located at the center of the filter shape may be referred to as a center filter tab. Two filter tabs with the same n value that are located at corresponding positions based on the center filter tab may be assigned the same filter coefficient. For example, a 7×7 diamond filter shape includes 25 filter tabs, and filter coefficients C0 to C11 are assigned in a centrally symmetrical manner, so that filter coefficients can be assigned to the 25 filter tabs with only 13 filter coefficients. For example, a 5x5 diamond filter shape includes 13 filter tabs, and filter coefficients C0 to C5 are assigned in a centrally symmetric manner, so that filter coefficients can be assigned to the 13 filter tabs using only seven filter coefficients. For example, to reduce the amount of data related to signaled filter coefficients, 12 of the 13 filter coefficients for a 7x7 diamond filter shape may be signaled (explicitly) and one filter coefficient may be derived (implicitly). For example, 6 of the 7 filter coefficients for a 5x5 diamond filter shape may be signaled (explicitly) and one filter coefficient may be derived (implicitly).

[0163] According to one embodiment of this document, ALF parameters used for the ALF procedure can be signaled via an adaptation parameter set (APS). The ALF parameters can be derived from filter information or ALF data for the ALF.

[0164] As mentioned above, ALF is a type of in-loop filtering technique that can be applied in video / picture coding. ALF can be implemented using a Wiener-based adaptive filter to minimize the mean square error (MSE) between original samples and decoded samples (or reconstructed samples). A high-level design for an ALF tool can incorporate syntax elements that can be accessed in the SPS and / or slice header (or tile group header).

[0165] In one example, before filtering each 4x4 luma block, a geometric transformation such as rotation or diagonal and vertical flipping can be applied to the filter coefficients f(k, l) and corresponding filter clipping values ​​c(k, l) depending on the gradient value calculated for that block. This is the same as applying these transformations to samples in the filter aided region. This is similar to generating other blocks to which the ALF is applied and aligning these blocks according to their orientation.

[0166] For example, three transformations, a diagonal, a vertical flip, and a rotation, can be performed based on the following equations:

[0167]

number

[0168]

number

[0169]

number

[0170] In Equations 1 to 3, K is the size of the filter. 0≦k, 1≦K−1 are coefficient coordinates. For example, (0,0) is the upper left corner coordinate, and / or (K−1,K−1) is the lower right corner coordinate. The relationship between the transformation and the four gradients in the four directions can be summarized as shown in the following table.

[0171] [Table 2]

[0172] ALF filter parameters can be signaled in the APS and slice header. Up to 25 luma filter coefficients and clipping value indices can be signaled in one APS. Up to 8 chroma filter coefficients and clipping value indices can be signaled in one APS. To reduce bit overhead, filter coefficients of different classifications for the luma component can be merged. The index of the APS used for the current slice (referenced by the current slice) can be signaled in the slice header.

[0173] The clipping value index decoded from the APS can be used to determine the clipping value using a luma table of clipping values ​​and a chroma table of clipping values. These clipping values ​​depend on the internal bit depth. More specifically, the luma table of clipping values ​​and the chroma table of clipping values ​​can be derived based on the following equations:

[0174]

number

[0175]

number

[0176] In the above formula, B is the internal bit depth, and N is the number of allowed clipping values ​​(a predetermined number), for example, N is 4.

[0177] In the slice header, up to seven APS indices can be signaled to indicate the luma filter set used for the current slice. The filtering procedure can be further controlled at the CTB level. For example, a flag indicating whether ALF is applied to the luma CTB can be signaled. The luma CTB can select one filter set from 16 fixed filter sets and a filter set from the APS. A filter set index can be signaled for the luma CTB to indicate which filter set to apply. The 16 fixed filter sets can be predefined and hard-coded in both the encoder and decoder.

[0178] For chroma components, an APS index can be signaled in the slice header to indicate the chroma filter set used for the current slice. At the CTB level, if there are more than one chroma filter sets in the APS, a filter index can be signaled for each chroma CTB.

[0179] The filter coefficients may be quantized to a norm of 128. To limit the multiplication complexity, bitstream conformance may be applied, such that coefficient values ​​of non-central positions are in the range of 0 to 28, and / or coefficient values ​​of other positions are in the range of -27 to 27-1. The central position coefficients may be predetermined to 128 without being signaled in the bitstream.

[0180] If the ALF is available for the current block, each sample R(i, j) can be filtered, and the filtered result R′(i, j) can be expressed as follows:

[0181]

number

[0182] In the above equation, f(k, l) is a decoded filter coefficient, K(x, y) is a clipping function, and c(k, l) is a decoded clipping parameter. For example, the variables k and / or l can vary from -L / 2 to L / 2, where L represents the filter length. The clipping function K(x, y) = min(y, max(-y, x)) can correspond to the function Clip3(-y, y, x).

[0183] In one example, to reduce the line buffer requirements of ALF, modified block classification and filtering can be applied for samples adjacent to horizontal CTU boundaries. For this purpose, virtual boundaries can be defined.

[0184] Figure 8 is a diagram illustrating a virtual boundary applied to a filtering procedure according to an embodiment of the present document. Figure 9 shows an example of an ALF procedure using a virtual boundary according to an embodiment of the present document. Figure 9 will be described together with Figure 8.

[0185] 8, the virtual boundary is a line defined by shifting the horizontal CTU boundary by N samples, where N is 4 for the luma component and / or N is 2 for the chroma component.

[0186] In Figure 8, modified block classification can be applied to the luma component. For the 1D Laplacian gradient calculation of a 4x4 block above the virtual boundary, only samples above the virtual boundary can be used. Similarly, for the 1D Laplacian gradient calculation of a 4x4 block below the virtual boundary, only samples below the virtual boundary can be used. The quantization of the activity value A can be scaled accordingly to account for the reduced number of samples used in the 1D Laplacian gradient calculation.

[0187] For the filtering procedure, a symmetric padding operation at the virtual boundary can be used for the luma and chroma components. Referring to Figure 8, if a filtered sample is located below the virtual boundary, adjacent samples located on the virtual boundary can be padded. Meanwhile, the corresponding samples on the other side can also be padded symmetrically.

[0188] The procedure described by Figure 9 can also be used for slice, brick, and / or tile boundaries when a filter is not available across the boundary. For ALF block classification, only samples contained in the same slice, brick, and / or tile can be used, and activity values ​​can be scaled accordingly. For ALF filtering, symmetric padding can be applied in the horizontal and / or vertical directions for horizontal and / or vertical boundaries, respectively.

[0189] 10 is a diagram illustrating a cross-component adaptive loop filtering (CCALF (CC-ALF)) procedure according to one embodiment of the present document. The CCALF procedure may also be referred to as a cross-component filtering procedure.

[0190] In one aspect, the ALF procedure may include a general ALF procedure and a CCALF procedure. That is, the CCALF procedure may refer to a sub-procedure of the ALF procedure. In another aspect, the filtering procedure may include a deblocking procedure, an SAO procedure, an ALF procedure, and / or a CCALF procedure.

[0191] CC-ALF can refine each chroma component using luma sample values. CC-ALF is controlled by (video) information in the bitstream, which can include (a) information about filter coefficients for each chroma component and (b) information about a mask that controls filter application to a block of samples. The filter coefficients can be signaled in the APS, and the block size and mask can be signaled at the slice level.

[0192] Referring to Figure 10, CC-ALF can operate by applying a linear diamond-shaped filter (Figure 10(b)) to the luma channel for each chroma component. The filter coefficients are sent to the APS, scaled by a factor of 210, and rounded to the nearest integer for fixed-point representation. The application of the filter is controlled by a variable block size, which can be signaled by a context coding flag received for each block of samples. The block size, along with the CC-ALF availability flag, can be received at the slice level for each chroma component. The block size (for chroma samples) can be 16x16, 32x32, 64x64, or 128x128.

[0193] In the following embodiment, a method for re-filtering or modifying ALF filtered reconstructed chroma samples based on reconstructed luma samples is proposed.

[0194] One embodiment of this document relates to filter on / off transmission and filter coefficient transmission in CC-ALF. As described above, information (syntax elements) in the syntax table disclosed in this document can be included in image / video information, can be configured / encoded by an encoding device, and can be transmitted to a decoding device in the form of a bitstream. The decoding device can parse / decode the information (syntax elements) in the corresponding syntax table. The decoding device can perform a picture / image / video decoding procedure (specifically, for example, the CC-ALF procedure) based on the decoded information. The same applies to other embodiments below.

[0195] The following table shows a partial syntax of slice header information according to an embodiment of this document.

[0196] [Table 3]

[0197] The following table shows example semantics for the syntax elements contained in the table.

[0198] [Table 4]

[0199] Referring to the above two tables, if sps_cross_component_alf_enabled_flag is 1 in the slice header, parsing of slice_cross_component_alf_cb_enabled_flag can be performed to determine whether Cb CC-ALF is applied within the corresponding slice. If slice_cross_component_alf_cb_enabled_flag is 1, CC-ALF is applied to the corresponding Cb slice, and if slice_cross_component_alf_cb_reuse_temporal_layer_filter is 1, the existing filter of the same temporal layer can be reused. If slice_cross_component_alf_cb_enabled_flag is 0, CC-ALF can be applied using the filter in the corresponding APS (adaptation parameter set) id via slice_cross_component_alf_cb_aps_idparsing. slice_cross_component_alf_cb_log2_control_size_minus4 can mean the CC-ALF application block unit in Cb slice.

[0200] For example, if the value of slice_cross_component_alf_cb_log2_control_size_minus4 is 0, whether or not CC-ALF is applied is determined in units of 16x16. If the value of slice_cross_component_alf_cb_log2_control_size_minus4 is 1, whether or not CC-ALF is applied is determined in units of 32x32. If the value of slice_cross_component_alf_cb_log2_control_size_minus4 is 2, whether or not CC-ALF is applied is determined in units of 64x64. If the value of slice_cross_component_alf_cb_log2_control_size_minus4 is 3, whether or not CC-ALF is applied is determined in units of 128x128. Also, for Cr CC-ALF, the same syntax structure as above is used.

[0201] The following table shows an example syntax for ALF data:

[0202] [Table 5]

[0203] The following table shows example semantics for the syntax elements contained in the table.

[0204] [Table 6-1]

[0205] [Table 6-2]

[0206] Referring to the two tables, CC-ALF syntax elements are configured to be transmitted and applied independently, regardless of the existing (general) ALF syntax structure. That is, CC-ALF can be applied even when the ALF tool on the SPS is turned off. Because CC-ALF must be able to operate independently of the existing ALF structure, a new hardware pipeline design is required. This increases the cost and delay of hardware implementation.

[0207] In addition, ALF determines whether to apply it to both luma and chroma images in CTU units, and the result of this determination is sent to the decoder via signaling. However, because the application of variable CC-ALF is determined in units of 16x16 to 128x128 and then applied, conflicts can occur between the existing ALF structure and CC-ALF. This causes problems in hardware implementation and also increases the line buffer size required to apply various variable CC-ALFs.

[0208] The present invention aims to solve the above-mentioned problems in hardware implementation of CC-ALF by integrating the CC-ALF syntax structure with the ALF syntax structure.

[0209] According to one embodiment of this document, a sequence parameter set (SPS) may include a CC-ALF enable flag (sps_ccalf_enable_flag) to determine whether CC-ALF is used (applied). The CC-ALF enable flag may be transmitted independently of an ALF enable flag (sps_alf_enabled_flag) for determining whether ALF is used (applied).

[0210] The following table shows a portion of an example syntax for an SPS according to this embodiment.

[0211] [Table 7]

[0212] Referring to the table, CC-ALF can be applied only when ALF is always active. That is, the CC-ALF enabled flag (sps_ccalf_enabled_flag) can be parsed only when the ALF enabled flag (sps_alf_enabled_flag) is 1. According to the table, CC-ALF and ALF can be combined. The CC-ALF enabled flag can indicate whether or not CC-ALF is enabled (or can be related to it).

[0213] The following table shows some example syntax for a slice header:

[0214] [Table 8]

[0215] Referring to the table, parsing of sps_ccalf_enabled_flag can be performed only if sps_alf_enabled_flag is 1. Syntax elements included in the table can be described based on Table 4. In one example, video information encoded by an encoding device or acquired (received) by a decoding device may include slice header information (slice_header()). Based on a determination that the value of the CC-ALF availability flag (sps_ccalf_flag) is 1, the slice header information may include a first flag (slice_cross_component_alf_cb_enabled_flag) associated with whether CC-ALF is available for the Cb color component of the filtered reconstructed chroma sample, and a second flag (slice_cross_component_alf_cr_enabled_flag) associated with whether CC-ALF is available for the Cr color component of the filtered reconstructed chroma sample.

[0216] In one example, the slice header information may include ID information (slice_cross_component_alf_cb_aps_id) of a first APS for deriving cross-component filter coefficients for the Cb color component based on a determination that the value of the first flag (slice_cross_component_alf_cb_enabled_flag) is 1. Based on a determination that the value of the second flag (slice_cross_component_alf_cr_enabled_flag) is 1, the slice header information may include ID information (slice_cross_component_alf_cr_aps_id) of a second APS for deriving cross-component filter coefficients for the Cr color component.

[0217] The following table shows a portion of the SPS syntax according to another example of this embodiment.

[0218] [Table 9]

[0219] The following table shows an example of a portion of the slice header syntax:

[0220] [Table 10]

[0221] Referring to Table 9, the SPS can include a CCALF enabled flag (sps_ccalf_enabled_flag) if ChromaArrayType is not 0 and the ALF enabled flag (sps_alf_enabled_flag) is 1. For example, if ChromaArrayType is not 0, the chroma format is not monochrome, and the CCALF enabled flag can be transmitted via the SPS based on the case where the chroma format is not monochrome.

[0222] Referring to Table 9, if ChromaArrayType is not 0, information about CCALF (slice_cross_component_alf_cb_enabled_flag, slice_cross_component_alf_cb_aps_id, slice_cross_component_alf_cr_enabled_flag, slice_cross_component_alf_cr_aps_id) can be included in the slice header information.

[0223] In one example, video information encoded by an encoding device or acquired by a decoding device may include the SPS. The SPS may include a first ALF enable flag (sps_alf_enabled_flag) associated with whether ALF is enabled. For example, based on a determination that the value of the first ALF enable flag is 1, the SPS may include a CCALF enable flag associated with whether the cross-component filtering is enabled. In another example, without using sps_ccalf_enabled_flag, if sps_alf_enabled_flag is 1, CCALF may always be applied (sps_ccalf_enabled_flag==1).

[0224] The following table shows a portion of the slice header syntax according to another example of this embodiment.

[0225] [Table 11]

[0226] Referring to the table above, CCALF enabled flag (sps_ccalf_enabled_flag) parsing can be performed only if the ALF enabled flag (sps_alf_enabled_flag) is 1.

[0227] The following table shows example semantics for the syntax elements contained in the table.

[0228] [Table 12]

[0229] The slice_ccalf_chroma_idc in the table above can also be explained by the semantics in the table below.

[0230] [Table 13]

[0231] The following table shows a portion of the slice header syntax according to another example of this embodiment.

[0232] [Table 14]

[0233] The syntax elements included in the above table may be described by Table 12 or Table 13. Also, if the chroma format is not monochrome, CCALF-related information may be included in the slice header.

[0234] The following table shows a portion of the slice header syntax according to another example of this embodiment.

[0235] [Table 15]

[0236] The following table shows example semantics for the syntax elements contained in the table.

[0237] [Table 16]

[0238] The following table shows a part of the slice header syntax according to another example of this embodiment. The syntax elements included in the following table can be explained by Table 12 or Table 13.

[0239] [Table 17]

[0240] Referring to the table, whether ALF and CC-ALF are applied per slice can be determined at once via slice_alf_enabled_flag. After parsing slice_alf_chroma_idc, if the first ALF enabled flag (sps_alf_enabled_flag) is 1, slice_ccalf_chroma_idc can be parsed.

[0241] Referring to the table, it can be determined whether sps_ccalf_enabled_flag is 1 in the slice header information only if slice_alf_enabled_flag is 1. The slice header information may include a second ALF availability flag (slice_alf_enabled_flag) related to whether ALF is available. Based on the determination that the value of the second ALF availability flag (slice_alf_enabled_flag) is 1, the CCALF can be available for the slice.

[0242] The following table exemplarily shows a portion of the APS syntax: The syntax element adaptation_parameter_set_id can indicate identifier information (ID information) of the APS.

[0243] [Table 18]

[0244] The following table shows an example syntax for ALF data:

[0245] [Table 19]

[0246] Referring to the two tables, an APS can include ALF data (alf_data()). An APS including ALF data can be called an ALF APS (ALF type APS). That is, the type of an APS including ALF data is ALF type. The type of APS can be determined by information about the APS type or a syntax element (aps_params_type). The ALF data can include a Cb filter signal flag (alf_cross_component_cb_filter_signal_flag or alf_cc_cb_filter_signal_flag) associated with whether a cross-component filter for the Cb color component is signaled. The ALF data can include a Cr filter signal flag (alf_cross_component_cr_filter_signal_flag or alf_cc_cr_filter_signal_flag) associated with whether a cross-component filter for the Cr color component is signaled.

[0247] In one example, based on the Cr filter signal flag, the ALF data may include information regarding the absolute value of the cross-component filter coefficient for the Cr color component (alf_cross_component_cr_coeff_abs) and information regarding the sign of the cross-component filter coefficient for the Cr color component (alf_cross_component_cr_coeff_sign). Based on the information regarding the absolute value of the cross-component filter coefficient for the Cr color component and the information regarding the sign of the cross-component filter coefficient for the Cr color component, the cross-component filter coefficient for the Cr color component may be derived.

[0248] In one example, the ALF data may include information on the absolute values ​​of cross-component filter coefficients for the Cb color component (alf_cross_component_cb_coeff_abs) and information on the signs of cross-component filter coefficients for the Cb color component (alf_cross_component_cb_coeff_sign). Based on the information on the absolute values ​​of cross-component filter coefficients for the Cb color component and the information on the signs of cross-component filter coefficients for the Cb color component, cross-component filter coefficients for the Cb color component may be derived.

[0249] The following table shows the syntax for another example of ALF data.

[0250] [Table 20]

[0251] Referring to the table, after alf_cross_component_filter_signal_flag is transmitted first, the Cb / Cr filter signal flag can be transmitted if the alf_cross_component_filter_signal_flag is set to 1. That is, the alf_cross_component_filter_signal_flag determines whether to transmit the CC-ALF filter coefficient by integrating Cb / Cr.

[0252] The following table shows the syntax for another example of ALF data.

[0253] [Table 21]

[0254] The following table shows example semantics for the syntax elements contained in the table.

[0255] [Table 22-1]

[0256] [Table 22-2]

[0257] The following table shows the syntax for another example of ALF data.

[0258] [Table 23]

[0259] The following table shows example semantics for the syntax elements contained in the table.

[0260] [Table 24-1]

[0261] [Table 24-2]

[0262] In the two tables, the degree of exp-Golombbinarization for parsing the alf_cross_component_cb_coeff_abs[j] and alf_cross_component_cr_coeff_abs[j] syntaxes can be defined as one of values ​​0 to 9.

[0263] Referring to the two tables, the ALF data may include a Cb filter signal flag (alf_cross_component_cb_filter_signal_flag or alf_cc_cb_filter_signal_flag) related to whether a cross-component filter for the Cb color component is signaled. Based on the Cb filter signal flag (alf_cross_component_cb_filter_signal_flag), the ALF data may include information related to the number of cross-component filters for the Cb color component (ccalf_cb_num_alt_filters_minus1). Based on the information related to the number of cross-component filters for the Cb color component, the ALF data may include information related to the absolute value of the cross-component filter coefficient for the Cb color component (alf_cross_component_cb_coeff_abs) and information related to the sign of the cross-component filter coefficient for the Cb color component (alf_cross_component_cr_coeff_sign). Based on information regarding the absolute values ​​of the cross-component filter coefficients for the Cb color component and information regarding the signs of the cross-component filter coefficients for the Cb color component, cross-component filter coefficients for the Cb color component can be derived.

[0264] In one example, the ALF data may include a Cr filter signal flag (alf_cross_component_cr_filter_signal_flag or alf_cc_cr_filter_signal_flag) related to whether a cross-component filter for the Cr color component is signaled. Based on the Cr filter signal flag (alf_cross_component_cr_filter_signal_flag), the ALF data may include information related to the number of cross-component filters for the Cr color component (ccalf_cr_num_alt_filters_minus1). Based on the information related to the number of cross-component filters for the Cr color component, the ALF data may include information related to the absolute value of the cross-component filter coefficient for the Cr color component (alf_cross_component_cr_coeff_abs) and information related to the sign of the cross-component filter coefficient for the Cr color component (alf_cross_component_cr_coeff_sign). Based on information regarding the absolute values ​​of the cross-component filter coefficients for the Cr color component and information regarding the signs of the cross-component filter coefficients for the Cr color component, cross-component filter coefficients for the Cr color component can be derived.

[0265] The following table shows the syntax for a coding tree unit according to one embodiment of this document.

[0266] [Table 25]

[0267] The following table shows example semantics for the syntax elements contained in the table.

[0268] [Table 26]

[0269] The following table shows the coding tree unit syntax according to another example of this embodiment.

[0270] [Table 27]

[0271] Referring to the table, CCALF can be applied in CTU units. In one example, the video information may include information about a coding tree unit (coding_tree_unit()). The information about the coding tree unit may include information about whether a cross-component filter is applied to the current block of a Cb color component (ccalf_ctb_flag[0]) and / or information about whether a cross-component filter is applied to the current block of a Cr color component (ccalf_ctb_flag[1]). In addition, the information about the coding tree unit may include information about a filter set index of a cross-component filter applied to the current block of a Cb color component (ccalf_ctb_filter_alt_idx[0]) and / or information about a filter set index of a cross-component filter applied to the current block of a Cr color component (ccalf_ctb_filter_alt_idx[1]). The syntax may be adaptively transmitted using slice_ccalf_enabled_flag and slice_ccalf_chroma_idc syntaxes.

[0272] The following table shows the coding tree unit syntax according to another example of this embodiment.

[0273] [Table 28]

[0274] The following table shows example semantics for the syntax elements contained in the table.

[0275] [Table 29]

[0276] The following table shows a coding tree unit syntax according to another example of this embodiment. The syntax elements included in the table below can be explained by Table 29.

[0277] [Table 30]

[0278] In one example, the video information may include information about a coding tree unit (coding_tree_unit()). The information about the coding tree unit may include information about whether a cross-component filter is applied to the current block of a Cb color component (ccalf_ctb_flag[0]) and / or information about whether a cross-component filter is applied to the current block of a Cr color component (ccalf_ctb_flag[1]). In addition, the information about the coding tree unit may include information about a filter set index of a cross-component filter applied to the current block of a Cb color component (ccalf_ctb_filter_alt_idx[0]) and / or information about a filter set index of a cross-component filter applied to the current block of a Cr color component (ccalf_ctb_filter_alt_idx[1]).

[0279] 11 and 12 schematically illustrate an example of a video / image encoding method and related components according to embodiment(s) of the present document. The method disclosed in FIG. 11 may be performed by the encoding device disclosed in FIG. 2. Specifically, for example, S1100 of FIG. 11 may be performed by the addition unit 250 of the encoding device, S1110 to S1140 may be performed by the filtering unit 260 of the encoding device, and S1150 may be performed by the entropy encoding unit 240 of the encoding device. The method disclosed in FIG. 11 may include embodiments detailed herein.

[0280] 11, an encoding apparatus may generate reconstructed luma samples and reconstructed chroma samples of a current block (S1100). The encoding apparatus may generate residual luma samples and / or residual chroma samples. The encoding apparatus may generate reconstructed luma samples based on the residual luma samples, and may generate reconstructed chroma samples based on the residual chroma samples.

[0281] In one example, a residual sample for the current block may be generated based on an original sample and a predicted sample of the current block. Specifically, the encoding apparatus may generate a predicted sample for the current block based on a prediction mode. In this case, various prediction methods disclosed herein, such as inter prediction or intra prediction, may be applied. A residual sample may be generated based on the predicted sample and the original sample.

[0282] In one example, the encoding device can generate residual luma samples. The residual luma samples can be generated based on the original luma samples and the predicted luma samples. In one example, the encoding device can generate residual chroma samples. The residual chroma samples can be generated based on the original chroma samples and the predicted chroma samples.

[0283] The encoding apparatus may derive transform coefficients. The encoding apparatus may derive transform coefficients based on a transform procedure for the residual samples. The encoding apparatus may derive transform coefficients for the residual luma samples (luma transform coefficients) and / or transform coefficients for the residual chroma samples (chroma transform coefficients). For example, the transform procedure may include at least one of DCT, DST, GBT, or CNT.

[0284] The encoding apparatus may derive quantized transform coefficients. The encoding apparatus may derive the quantized transform coefficients based on a quantization procedure for the transform coefficients. The quantized transform coefficients may have a one-dimensional vector form based on a coefficient scan order. The quantized transform coefficients may include quantized luma transform coefficients and / or quantized chroma transform coefficients.

[0285] The encoding apparatus may generate residual information indicating (including) the quantized transform coefficients, which may be generated through various encoding methods such as Exponential-Golomb, CAVLC, CABAC, etc.

[0286] The encoding device may generate prediction-related information based on prediction samples and / or modes applied thereto, and the prediction-related information may include information on various prediction modes (e.g., merge mode, MVP mode, etc.), MVD information, etc.

[0287] The encoding apparatus may derive ALF filter coefficients for the ALF procedure (S1110). The ALF filter coefficients may include ALF luma filter coefficients for reconstructed luma samples and ALF chroma filter coefficients for reconstructed chroma samples. Filtered reconstructed luma samples and / or filtered reconstructed chroma samples may be generated based on the ALF filter coefficients.

[0288] The encoding apparatus may generate ALF-related information (S1120). The encoding apparatus may generate the ALF-related information based on the ALF filter coefficients. The encoding apparatus may derive ALF-related parameters that can be applied for filtering the reconstructed samples, and generate the ALF-related information. For example, the ALF-related information may include the ALF-related information described in detail herein.

[0289] The encoding apparatus may derive cross-component filters (CCALF filters) and / or cross-component filter coefficients (CCALF filter coefficients) (S1130). The cross-component filters and / or cross-component filter coefficients may be used in the CCALF procedure. Modified filtered reconstructed chroma samples may be generated based on the cross-component filters and / or cross-component filter coefficients.

[0290] The encoding apparatus may generate loss-component filtering-related information (or CCALF-related information) (S1140). In one example, the cross-component filtering-related information may include information about the number of cross-component filters and information about the cross-component filter coefficients. The cross-component filters may include a cross-component filter for a Cb color component and a cross-component filter for a Cr color component.

[0291] In one example, the CCALF-related information may include a CCALF availability flag, a flag related to whether CCALF is available for the Cb (or Cr) color component, a Cb (or Cr) filter signal flag related to whether a cross-component filter for the Cb (or Cr) color component is signaled, information related to the number of cross-component filters for the Cb (or Cr) color component, information related to values ​​of cross-component filter coefficients for the Cb (or Cr) color component, information related to absolute values ​​of cross-component filter coefficients for the Cb (or Cr) color component, information related to signs of cross-component filter coefficients for the Cb (or Cr) color component, and / or information related to the coding tree unit (coding tree unit syntax) related to whether a cross-component filter is applied to the current block of the Cb (or Cr) color component.

[0292] The image / video information may include various information according to embodiments of this document, for example, the image / video information may include at least one of the information disclosed in Tables 1 to 30 above.

[0293] In one embodiment, the image information may include a sequence parameter set (SPS). The SPS may include a cross-component adaptive loop filter (CCALF) availability flag related to whether the cross-component filtering is available. Based on a determination that the CCALF availability flag is 1, ID information (identifier information) of adaptation parameter sets (APSs) including ALF data used to derive cross-component filter coefficients for CCALF may be derived. The image information may include slice header information.

[0294] According to one example of an embodiment, the slice header information may include ID information of an APS including ALF data used to derive the cross-component filter coefficients. In another example, based on a determination that the CCALF available flag is 1, the slice header information may include ID information of an APS including ALF data used to derive the cross-component filter coefficients.

[0295] In one embodiment, the SPS may include an ALF enabled flag (sps_ccalf_enabled_flag) associated with whether ALF is enabled. Based on a determination that the value of the first ALF enabled flag is 1, the SPS may include a CCALF enabled flag associated with whether the cross-component filtering is enabled.

[0296] In one embodiment, the video information may include slice header information and an adaptation parameter set (APS). The header information may include information related to an identifier of an APS including ALF data. For example, the cross-component filter coefficients may be derived based on the ALF data.

[0297] In one embodiment, the slice header information may include an ALF enabled flag (slice_alf_enabled_flag) associated with whether ALF is enabled. The sps_alf_enabled_flag and slice_alf_enabled_flag may be referred to as a first ALF enabled flag and a second ALF enabled flag, respectively. Whether the CCALF enabled flag is set to 1 may be determined based on whether the ALF enabled flag (slice_alf_enabled_flag) is set to 1. In one example, whether the CCALF (Cross-Component Filtering) is enabled may be determined based on whether the ALF enabled flag (slice_alf_enabled_flag) is set to 1.

[0298] In one embodiment, the header information (slice header information) may include a first flag associated with whether CCALF is enabled for the Cb color component of the filtered reconstructed chroma sample and a second flag associated with whether CCALF is enabled for the Cr color component of the filtered reconstructed chroma sample. In another example, based on a determination that the ALF enabled flag (slice_alf_enabled_flag) is 1, the header information (slice header information) may include a first flag associated with whether CCALF is enabled for the Cb color component of the filtered reconstructed chroma sample and a second flag associated with whether CCALF is enabled for the Cr color component of the filtered reconstructed chroma sample.

[0299] In one embodiment, based on a determination that the value of the first flag is 1, the slice header information may include ID information of a first APS (information associated with an identifier of a second APS) for deriving cross-component filter coefficients for the Cb color component. Based on a determination that the value of the second flag is 1, the slice header information may include ID information of a second APS (information associated with an identifier of a second APS) for deriving cross-component filter coefficients for the Cr color component.

[0300] In one embodiment, the first ALF data included in the first APS may include a Cb filter signal flag related to whether a cross-component filter for the Cb color component is signaled. Based on the Cb filter signal flag, the first ALF data may include information related to the number of cross-component filters for the Cb color component. Based on the information related to the number of cross-component filters for the Cb color component, the first ALF data may include information related to absolute values ​​of cross-component filter coefficients for the Cb color component and information related to signs of the cross-component filter coefficients for the Cb color component. Based on the information related to the absolute values ​​of the cross-component filter coefficients for the Cb color component and the information related to the signs of the cross-component filter coefficients for the Cb color component, the cross-component filter coefficients for the Cb color component may be derived.

[0301] In one embodiment, the information related to the number of cross-component filters for the Cb color component is a zeroth order exponential Golomb (0 th EG) can be coded.

[0302] In one embodiment, the second ALF data included in the second APS may include a Cr filter signal flag associated with whether a cross-component filter for the Cr color component is signaled. Based on the Cr filter signal flag, the second ALF data may include information associated with the number of cross-component filters for the Cr color component. Based on the information associated with the number of cross-component filters for the Cr color component, the second ALF data may include information associated with the absolute values ​​of cross-component filter coefficients for the Cr color component and information associated with the signs of the cross-component filter coefficients for the Cr color component. Based on the information associated with the absolute values ​​of the cross-component filter coefficients for the Cr color component and the information associated with the signs of the cross-component filter coefficients for the Cr color component, the cross-component filter coefficients for the Cr color component may be derived.

[0303] In one embodiment, the information related to the number of cross-component filters for the Cr color component is a zeroth order exponential Golomb (0 th EG) can be coded.

[0304] In one embodiment, the image information may include information about a coding tree unit, the information about the coding tree unit including information about whether a cross-component filter is applied to the current block of a Cb color component and / or information about whether a cross-component filter is applied to the current block of a Cr color component.

[0305] In one embodiment, the information about the coding tree unit may include information about a filter set index of a cross-component filter applied to the current block of a Cb color component, and / or information about a filter set index of a cross-component filter applied to the current block of a Cr color component.

[0306] Figures 13 and 14 schematically illustrate an example of a video / image decoding method and related components according to embodiment(s) of the present document. The method disclosed in Figure 13 may be performed by the decoding device disclosed in Figure 3 or Figure 14. Specifically, for example, S1300 in Figure 13 may be performed by the entropy decoding unit 310 of the decoding device, S1310 may be performed by the addition unit 340 of the decoding device, and S1320 and S1330 may be performed by the filtering unit 350 of the decoding device.

[0307] Referring to FIG. 13, a decoding apparatus may receive / acquire video / image information (S1300). The video / image information may include prediction-related information and / or residual information. The decoding apparatus may receive / acquire the video / image information through a bitstream. The residual information may be generated through various encoding methods such as Exponential-Golomb, CAVLC, CABAC, etc. In one example, the video / image information may further include CCAL-related information. For example, the CCALF-related information may include a CCALF availability flag, a flag related to whether CCALF is available for the Cb (or Cr) color component, a Cb (or Cr) filter signal flag related to whether a cross-component filter for the Cb (or Cr) color component has been signaled, information related to the number of cross-component filters for the Cb (or Cr) color component, information related to the absolute values ​​of the cross-component filter coefficients for the Cb (or Cr) color component, information related to the signs of the cross-component filter coefficients for the Cb (or Cr) color component, and / or information related to the coding tree unit (coding tree unit syntax) related to whether a cross-component filter is applied to the current block of the Cb (or Cr) color component.

[0308] The image / video information may include various information according to the embodiments of this document, for example, the image / video information may include at least one of the information disclosed in Tables 1 to 30 above.

[0309] A decoding device may derive transform coefficients. Specifically, the decoding device may derive quantized transform coefficients based on residual information. The transform coefficients may include luma transform coefficients and chroma transform coefficients. The quantized transform coefficients may have a one-dimensional vector form based on a coefficient scan order. The decoding device may derive the transform coefficients based on an inverse quantization procedure for the quantized transform coefficients.

[0310] The decoding device may derive residual samples. The decoding device may derive residual samples based on transform coefficients. The residual samples may include residual luma samples and residual chroma samples. For example, the residual luma samples may be derived based on luma transform coefficients, and the residual chroma samples may be derived based on chroma transform coefficients. Furthermore, the residual samples for the current block may be derived based on original samples and predicted samples of the current block.

[0311] The decoding apparatus may perform prediction based on the image / video information to derive predicted samples of the current block. The decoding apparatus may derive the predicted samples of the current block based on prediction-related information. The prediction-related information may include prediction mode information. The decoding apparatus may determine whether inter-prediction or intra-prediction is applied to the current block based on the prediction mode information, and perform prediction based on the determination. The predicted samples may include predicted luma samples and / or predicted chroma samples.

[0312] A decoding apparatus may generate / derive reconstructed luma samples and / or reconstructed chroma samples (S1310). The reconstructed samples may include reconstructed luma samples and / or reconstructed chroma samples. The decoding apparatus may generate reconstructed luma samples based on the residual luma samples. The decoding apparatus may generate reconstructed chroma samples based on the residual chroma samples. The luma components of the reconstructed samples may correspond to the reconstructed luma samples, and the chroma components of the reconstructed samples may correspond to the reconstructed chroma samples.

[0313] The decoding device may perform an adaptive loop filtering (ALF) procedure on the reconstructed chroma samples to generate filtered reconstructed chroma samples (S1320). In the ALF procedure, the decoding device may derive ALF filter coefficients for the ALF procedure on the reconstructed chroma samples. Additionally, the decoding device may derive ALF filter coefficients for the ALF procedure on the reconstructed luma samples. The ALF filter coefficients may be derived based on ALF parameters included in the ALF data in the APS.

[0314] The decoding device may generate filtered reconstructed chroma samples. The decoding device may generate filtered reconstructed samples based on the reconstructed chroma samples and the ALF filter coefficients.

[0315] The decoding apparatus may perform a cross-component filtering procedure on the filtered reconstructed chroma samples to generate modified filtered reconstructed chroma samples (S1330). In the cross-component filtering procedure, the decoding apparatus may derive cross-component filter coefficients for the cross-component filtering. The cross-component filter coefficients may be derived based on CCALF-related information in the ALF data included in the APS, and identifier (ID) information of the corresponding APS may be included in (or signaled via) the slice header.

[0316] A decoding device may generate modified filtered reconstructed chroma samples. The decoding device may generate the modified filtered reconstructed chroma samples based on the reconstructed luma samples, the filtered reconstructed chroma samples, and the cross-component filter coefficients. In one example, the decoding device may derive a difference between two of the reconstructed luma samples and multiply the difference by one of the cross-component filter coefficients. Based on the result of the multiplication and the filtered reconstructed chroma samples, the decoding device may generate the modified filtered reconstructed chroma samples. For example, the decoding device may generate the modified filtered reconstructed chroma samples based on the sum of the product and one of the filtered reconstructed chroma samples.

[0317] In one embodiment, the image information may include a sequence parameter set (SPS). The SPS may include a CCALF availability flag associated with whether the cross-component filtering is available. Based on a determination that the CCALF availability flag is 1, ID information (identifier information) of adaptation parameter sets (APSs) including ALF data used to derive cross-component filter coefficients for CCALF may be derived. The image information may include slice header information.

[0318] According to one example embodiment, the slice header information may include ID information of an APS including ALF data used to derive the cross-component filter coefficients. In another example, based on a determination that the CCALF available flag is 1, the slice header information may include ID information of an APS including ALF data used to derive the cross-component filter coefficients.

[0319] In one embodiment, the SPS may include an ALF enabled flag (sps_ccalf_enabled_flag) associated with whether ALF is enabled. Based on a determination that the value of the ALF enabled flag (sps_ccalf_enabled_flag) is 1, the SPS may include a CCALF enabled flag associated with whether the cross-component filtering is enabled.

[0320] In one embodiment, the video information may include slice header information and an adaptation parameter set (APS). The header information may include information related to an identifier of an APS including ALF data. For example, the cross-component filter coefficients may be derived based on the ALF data.

[0321] In one embodiment, the slice header information may include an ALF enabled flag (slice_alf_enabled_flag) associated with whether ALF is enabled. The sps_alf_enabled_flag and slice_alf_enabled_flag may be referred to as a first ALF enabled flag and a second ALF enabled flag, respectively. Whether the CCALF enabled flag is set to 1 may be determined based on whether the ALF enabled flag (slice_alf_enabled_flag) is set to 1. In one example, whether the CCALF (Cross-Component Filtering) is enabled may be determined based on whether the ALF enabled flag (slice_alf_enabled_flag) is set to 1.

[0322] In one embodiment, the header information (slice header information) may include a first flag associated with whether CCALF is enabled for the Cb color component of the filtered reconstructed chroma sample and a second flag associated with whether CCALF is enabled for the Cr color component of the filtered reconstructed chroma sample. In another example, based on a determination that the ALF enabled flag (slice_alf_enabled_flag) is 1, the header information (slice header information) may include a first flag associated with whether CCALF is enabled for the Cb color component of the filtered reconstructed chroma sample and a second flag associated with whether CCALF is enabled for the Cr color component of the filtered reconstructed chroma sample.

[0323] In one embodiment, based on a determination that the value of the first flag is 1, the slice header information may include ID information of a first APS (information associated with an identifier of a second APS) for deriving cross-component filter coefficients for the Cb color component. Based on a determination that the value of the second flag is 1, the slice header information may include ID information of a second APS (information associated with an identifier of a second APS) for deriving cross-component filter coefficients for the Cr color component.

[0324] In one embodiment, the first ALF data included in the first APS may include a Cb filter signal flag related to whether a cross-component filter for the Cb color component is signaled. Based on the Cb filter signal flag, the first ALF data may include information related to the number of cross-component filters for the Cb color component. Based on the information related to the number of cross-component filters for the Cb color component, the first ALF data may include information related to absolute values ​​of cross-component filter coefficients for the Cb color component and information related to signs of the cross-component filter coefficients for the Cb color component. Based on the information related to the absolute values ​​of the cross-component filter coefficients for the Cb color component and the information related to the signs of the cross-component filter coefficients for the Cb color component, the cross-component filter coefficients for the Cb color component may be derived.

[0325] In one embodiment, the information related to the number of cross-component filters for the Cb color component is a zeroth order exponential Golomb (0 th EG) can be coded.

[0326] In one embodiment, the second ALF data included in the second APS may include a Cr filter signal flag associated with whether a cross-component filter for the Cr color component is signaled. Based on the Cr filter signal flag, the second ALF data may include information associated with the number of cross-component filters for the Cr color component. Based on the information associated with the number of cross-component filters for the Cr color component, the second ALF data may include information associated with the absolute values ​​of cross-component filter coefficients for the Cr color component and information associated with the signs of the cross-component filter coefficients for the Cr color component. Based on the information associated with the absolute values ​​of the cross-component filter coefficients for the Cr color component and the information associated with the signs of the cross-component filter coefficients for the Cr color component, the cross-component filter coefficients for the Cr color component may be derived.

[0327] In one embodiment, the information related to the number of cross-component filters for the Cr color component is a zeroth order exponential Golomb (0 th EG) can be coded.

[0328] In one embodiment, the image information may include information about a coding tree unit, the information about the coding tree unit including information about whether a cross-component filter is applied to the current block of a Cb color component and / or information about whether a cross-component filter is applied to the current block of a Cr color component.

[0329] In one embodiment, the information about the coding tree unit may include information about a filter set index of a cross-component filter applied to the current block of a Cb color component, and / or information about a filter set index of a cross-component filter applied to the current block of a Cr color component.

[0330] If residual samples for the current block exist, the decoding apparatus may receive information about the residuals for the current block. The information about the residuals may include transform coefficients related to the residual samples. The decoding apparatus may derive residual samples (or residual sample arrays) for the current block based on the residual information. Specifically, the decoding apparatus may derive quantized transform coefficients based on the residual information. The quantized transform coefficients may have a one-dimensional vector form based on a coefficient scanning order. The decoding apparatus may derive transform coefficients based on a dequantization procedure for the quantized transform coefficients. The decoding apparatus may derive residual samples based on the transform coefficients.

[0331] The decoding apparatus may generate reconstructed samples based on (intra) predicted samples and residual samples, and derive reconstructed blocks or pictures based on the reconstructed samples. Specifically, the decoding apparatus may generate reconstructed samples based on the sum of the (intra) predicted samples and the residual samples. As described above, the decoding apparatus may then apply an in-loop filtering procedure, such as deblocking filtering and / or an SAO procedure, to the reconstructed pictures as needed to improve subjective / objective image quality.

[0332] For example, a decoding device can decode a bitstream or encoded information to obtain video information including all or part of the above-described information (or syntax elements). Also, the bitstream or encoded information can be stored in a computer-readable storage medium and can be used to cause the above-described decoding method to be performed.

[0333] In the above-described embodiments, the method is described based on a flowchart as a series of steps or blocks, but the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and that different steps may be included, or one or more steps in the flowcharts may be deleted without affecting the scope of the embodiments herein.

[0334] The methods according to the embodiments of the present document described above may be implemented in the form of software, and the encoding device and / or decoding device according to the present document may be included in devices that perform video processing, such as TVs, computers, smartphones, set-top boxes, display devices, etc.

[0335] When an embodiment of this document is implemented in software, the method described above may be implemented with modules (processes, functions, etc.) that perform the functions described above. The modules may be stored in memory and executed by a processor. The memory may be internal or external to the processor and may be coupled to the processor in various well-known ways. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described herein may be implemented on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in the figures may be implemented on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information on instructions) or algorithms for implementation may be stored on a digital storage medium.

[0336] In addition, the decoding device and encoding device to which the embodiments of this document are applied may be included in a multimedia broadcast transmitting / receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video interaction device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a customized video (VoD) service providing device, an over-the-top (OTT) video (over-the-top) device, an internet streaming service providing device, a three-dimensional (3D) video device, a virtual reality (VR) device, an augmented reality (AR) device, a video telephone video device, a vehicle terminal (e.g., a vehicle terminal (including an autonomous vehicle), an airplane terminal, a ship terminal, etc.), a medical video device, etc., and may be used to process video signals or data signals. For example, over-the-top (OTT) video (over-the-top) devices may include a game console, a Blu-ray player, an internet access TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.

[0337] In addition, a processing method to which the embodiments of this document are applied may be produced in the form of a computer-executable program and stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiments of this document may also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The computer-readable recording medium may also include media embodied in the form of a carrier wave (e.g., transmission via the Internet). The bitstream generated by the encoding method may be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.

[0338] Furthermore, the embodiments of the present document may be embodied in a computer program product with program code, which may be executed by a computer in accordance with the embodiments of the present document. The program code may be stored on a computer-readable carrier.

[0339] FIG. 15 illustrates an example of a content streaming system in which the embodiments disclosed herein can be applied.

[0340] Referring to FIG. 15, a content streaming system to which the embodiments of this document are applied may broadly include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0341] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted.

[0342] The bitstream can be generated by an encoding method or a bitstream generation method to which an embodiment of this document is applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0343] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which controls commands and responses between devices in the content streaming system.

[0344] The streaming server can receive content from a media repository and / or an encoding server. For example, if content is received from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.

[0345] Examples of the user device include a mobile phone, a smartphone, a laptop computer, a digital broadcasting terminal, a PDA (personal digital assistant), a PMP (portable multimedia player), a navigation system, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, a smart glass, a head mounted display (HMD)), a digital TV, a desktop computer, a digital signage, etc.

[0346] Each server in the content streaming system can be operated as a distributed server, in which case data received by each server can be processed in a distributed manner.

[0347] The claims described herein may be combined in various ways. For example, technical features of method claims herein may be combined to be embodied as an apparatus, and technical features of apparatus claims herein may be combined to be embodied as a method. Furthermore, technical features of method claims herein and technical features of apparatus claims herein may be combined to be embodied as an apparatus, and technical features of method claims herein and technical features of apparatus claims herein may be combined to be embodied as a method.

Claims

1. A video decoding method performed by a decoding device, comprising: receiving video information including prediction-related information via a bitstream; deriving an intra prediction mode of a current block in a current picture based on the prediction-related information; generating predicted luma samples and predicted chrominance samples based on the intra prediction mode; generating reconstructed luma samples and reconstructed chroma samples based on the predicted luma samples and the predicted chroma samples; performing an adaptive loop filtering (ALF) procedure on the reconstructed chroma samples to generate filtered reconstructed chroma samples; performing a cross-component filtering procedure on the filtered reconstructed chroma samples to generate modified filtered reconstructed chroma samples; the video information includes information regarding cross-component filtering; The step of performing a cross-component filtering procedure comprises: deriving a number of cross-component filters for the cross-component filtering based on the information regarding the cross-component filtering; deriving cross-component filter coefficients for the cross-component filtering procedure based on the number of cross-component filters; generating the modified filtered reconstructed chroma samples based on the filtered reconstructed chroma samples and the cross-component filter coefficients; The video information includes a sequence parameter set (SPS) and slice header information, The SPS includes an ALF availability flag related to whether the ALF procedure is available; Based on a determination that the value of the ALF enable flag is 1, the SPS includes a cross-component adaptive loop filter (CCALF) enable flag associated with whether the cross-component filtering is enabled; Based on a determination that the value of the ALF available flag included in the SPS is 1, the slice header information includes an ALF available flag related to whether the ALF is available; based on a determination that the value of the ALF available flag included in the slice header information is 1 and the value of the CCALF available flag included in the SPS is 1, the slice header information includes information on whether the CCALF is available for the filtered reconstructed chroma samples; If the value of the information regarding whether the CCALF is available for the filtered reconstructed chroma samples is 1, the slice header information includes ID (identification) information of an adaptation parameter set (APS) including ALF data used to derive the cross-component filter coefficients; the video information includes information about coding tree units; The information about the coding tree unit may include: Information about whether a cross-component filter is applied to the Cb color components of the current block; and Information about whether a cross-component filter is applied to the Cr color component of the current block; and Information regarding a filter set index of a cross-component filter to be applied to the Cb color component of the current block; and The method includes information about a filter set index of a cross-component filter to be applied to the Cr color component of the current block.

2. A video encoding method performed by an encoding device, comprising: determining an intra-prediction mode for a current block in a current picture; generating predicted luma samples and predicted chrominance samples based on the intra prediction mode; generating reconstructed luma samples and reconstructed chroma samples based on the predicted luma samples and the predicted chroma samples; deriving adaptive loop filtering (ALF) filter coefficients for an ALF procedure; generating ALF-related information based on the ALF filter coefficients; deriving cross-component filters and cross-component filter coefficients for a cross-component filtering procedure; generating cross-component filtering related information based on the cross-component filter and the cross-component filter coefficients; encoding video information including information for generating reconstruction samples, the ALF-related information, and the cross-component filtering-related information; the cross-component filtering related information includes information about the number of the cross-component filters and information about the cross-component filter coefficients; The video information includes an SPS (sequence parameter set) and slice header information, The SPS includes an ALF availability flag related to whether the ALF procedure is available; Based on a determination that the value of the ALF enable flag is 1, the SPS includes a cross-component adaptive loop filter (CCALF) enable flag associated with whether cross-component filtering is enabled; Based on a determination that the value of the ALF available flag included in the SPS is 1, the slice header information includes an ALF available flag related to whether the ALF is available; based on a determination that the value of the ALF available flag included in the slice header information is 1 and the value of the CCALF available flag included in the SPS is 1, the slice header information includes information on whether the CCALF is available for filtered reconstructed chroma samples; If the value of the information regarding whether the CCALF is available for the filtered reconstructed chroma samples is 1, the slice header information includes ID (identification) information of an adaptation parameter set (APS) including ALF data used to derive the cross-component filter coefficients; the video information includes information about coding tree units; The information about the coding tree unit may include: Information about whether a cross-component filter is applied to the Cb color components of the current block; and Information about whether a cross-component filter is applied to the Cr color component of the current block; and Information regarding a filter set index of a cross-component filter to be applied to the Cb color component of the current block; and The method includes information about a filter set index of a cross-component filter to be applied to the Cr color component of the current block.

3. A method for transmitting data for video, comprising: obtaining the video bitstream, The bitstream comprises: determining an intra prediction mode for a current block in a current picture; generating predicted luma samples and predicted chroma samples based on the intra prediction mode; generating reconstructed luma samples and reconstructed chroma samples based on the predicted luma samples and the predicted chroma samples; Deriving adaptive loop filtering (ALF) filter coefficients for an ALF procedure; generating ALF-related information based on the ALF filter coefficients; Deriving cross-component filters and cross-component filter coefficients for the cross-component filtering procedure; generating cross-component filtering related information based on the cross-component filter and the cross-component filter coefficients; based on encoding video information including information for generating reconstruction samples, the ALF-related information, and the cross-component filtering-related information; transmitting the data including the bitstream; the cross-component filtering related information includes information about the number of the cross-component filters and information about the cross-component filter coefficients; The video information includes an SPS (sequence parameter set) and slice header information, The SPS includes an ALF availability flag related to whether the ALF procedure is available; Based on a determination that the value of the ALF enable flag is 1, the SPS includes a cross-component adaptive loop filter (CCALF) enable flag associated with whether the cross-component filtering is enabled; Based on a determination that the value of the ALF available flag included in the SPS is 1, the slice header information includes an ALF available flag related to whether the ALF is available; based on a determination that the value of the ALF available flag included in the slice header information is 1 and the value of the CCALF available flag included in the SPS is 1, the slice header information includes information on whether the CCALF is available for filtered reconstructed chroma samples; If the value of the information regarding whether the CCALF is available for the filtered reconstructed chroma samples is 1, the slice header information includes ID (identification) information of an adaptation parameter set (APS) including ALF data used to derive the cross-component filter coefficients; the video information includes information about a coding tree unit; The information about the coding tree unit may include: Information about whether a cross-component filter is applied to the Cb color components of the current block; and Information about whether a cross-component filter is applied to the Cr color component of the current block; and Information regarding a filter set index of a cross-component filter to be applied to the Cb color component of the current block; and The method includes information about a filter set index of a cross-component filter to be applied to the Cr color component of the current block.

Citation Information

Patent Citations

  • Cross-component adaptive loop filtering based video coding apparatus and method

    JP7734259B2

  • Cross-component adaptive loop filter for chroma

    WO2021032751A1