Filtering-based video coding apparatus and method

The filtering-based video coding method addresses the inefficiencies in compressing high-resolution and immersive media by applying adaptive loop filtering and cross-component ALF, enhancing compression efficiency and image quality.

JP7869362B2Active Publication Date: 2026-06-02LG ELECTRONICS INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2025-03-21
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality video/images, such as 4K or UHD, and immersive media like VR and AR, has led to higher transmission and storage costs due to increased video data, necessitating a more efficient video/image compression technique.

Method used

A filtering-based video coding method that includes adaptive loop filtering (ALF) and cross-component ALF (CCALF) applied to reconstructed chroma samples based on luma samples, with signaling options for CCALF availability and filter coefficients, allowing adaptive application in units of pictures, slices, and coding blocks.

Benefits of technology

Enhances video compression efficiency, improves subjective and objective visual quality, and corrects chroma samples based on luma samples, increasing encoding efficiency and image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007869362000040
    Figure 0007869362000040
  • Figure 0007869362000041
    Figure 0007869362000041
  • Figure 0007869362000042
    Figure 0007869362000042
Patent Text Reader

Abstract

To provide a device and a method of enhancing efficiency of image / video coding.SOLUTION: According to an embodiment of the present disclosure, a cross-component filter coefficient for cross-component filtering is derived. Modified filtered reconstructed chroma samples are generated on the basis of the cross-component filter coefficient. This embodiment increases accuracy of in-loop filtering.SELECTED DRAWING: Figure 12
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to a filtering-based video coding apparatus and method.

Background Art

[0002] In recent years, the demand for high-resolution and high-quality video / images such as 4K or UHD (Ultra High Definition) video / images of 8K or higher has been increasing in various fields. As the video / image data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to existing video / image data. Therefore, when transmitting video data using a medium such as an existing wired or wireless broadband line, or storing video / image data using an existing storage medium, the transmission cost and storage cost increase.

[0003] Also, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcast of video / images having video characteristics different from real-world video / images, such as game video, has been increasing.

[0004] Thus, there is a need for a highly efficient video / image compression technique to effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality video / images having various characteristics as described above.

Summary of the Invention

Means for Solving the Problems

[0005] According to one embodiment of this document, a method and apparatus for enhancing the efficiency of video / image coding are provided.

[0006] According to one embodiment of this document, an efficient filtering application method and apparatus are provided.

[0007] According to one embodiment of this document, an efficient ALF application method and apparatus are provided.

[0008] According to one embodiment of this document, a filtering procedure can be performed on a reconstructed chroma sample based on a reconstructed luma sample.

[0009] According to one embodiment of this document, the reconstructed chroma sample filtered based on the reconstructed chroma sample can be modified.

[0010] According to one embodiment of this document, information regarding whether or not CCALF is available in SPS can be signaled.

[0011] According to one embodiment of this document, information regarding the values ​​of cross-component filter coefficients can be derived from ALF data (general ALF data or CCALF data).

[0012] According to one embodiment of this document, APS identifier (ID) information, including ALF data for deriving cross-component filter coefficients, can be signaled in the slice.

[0013] According to one embodiment of this document, information regarding the filter set index for CCALF can be signaled on a CTU (block) basis.

[0014] According to one embodiment of this document, a video / image decoding method performed by a decoding device is provided.

[0015] According to one embodiment of this document, a decoding device for video / image decoding is provided.

[0016] According to one embodiment of this document, a video / image encoding method performed by an encoding device is provided.

[0017] According to one embodiment of this document, an encoding device for performing video / video encoding is provided.

[0018] According to one embodiment of this document, a computer-readable digital storage medium storing encoded video / video information generated by a video / video encoding method disclosed in at least one of the embodiments of this document is provided.

[0019] According to one embodiment of this document, a computer-readable digital storage medium storing encoded information or encoded video / video information that causes a decoding device to perform a video / video decoding method disclosed in at least one of the embodiments of this document is provided.

Advantages of the Invention

[0020] According to one embodiment of this document, the overall video / video compression efficiency can be increased.

[0021] According to one embodiment of this document, the subjective / objective visual quality can be improved through efficient filtering.

[0022] According to one embodiment of this document, the ALF procedure can be efficiently executed and the filtering performance can be improved.

[0023] According to one embodiment of this document, the restored chroma samples filtered based on the restored luma samples can be corrected, and the image quality and coding accuracy for the chroma component of the decoded picture can be improved.

[0024] According to one embodiment of this document, the CCALF procedure can be efficiently executed.

[0025] According to one embodiment of this document, ALF-related information can be efficiently signaled.

[0026] According to one embodiment of this document, CCALF-related information can be efficiently signaled.

[0027] According to one embodiment of this document, ALF and / or CCALF can be adaptively applied in units of pictures, slices, and / or coding blocks.

[0028] According to one embodiment of this document, when CCALF is used in an encoding and decoding method and apparatus for still images or moving images, the filter coefficients for CCALF and the on / off transmission method in units of blocks or CTUs are improved, and the encoding efficiency can be increased.

Brief Description of Drawings

[0029] [Figure 1] An example of a video / image coding system applicable to embodiments of this document is schematically shown. [Figure 2] This is a drawing schematically explaining the configuration of a video / image encoding apparatus applicable to embodiments of this document. [Figure 3] This is a drawing schematically explaining the configuration of a video / image decoding apparatus applicable to embodiments of this document. [Figure 4] An exemplary hierarchical structure for coded video / images is shown. [Figure 5] This is a flowchart for explaining an intra-prediction-based block restoration method in an encoding apparatus. [Figure 6] This is a flowchart for explaining an intra-prediction-based block restoration method in a decoding apparatus. [Figure 7] This is a flowchart for explaining an inter-prediction-based block restoration method in an encoding apparatus. [Figure 8] This is a flowchart for explaining an inter-prediction-based block restoration method in a decoding apparatus. [Figure 9] An example of the ALF filter shape is shown. [Figure 10] This diagram illustrates the virtual boundary applied to the filtering procedure according to one embodiment of this document. [Figure 11] One embodiment of this document demonstrates an example of an ALF procedure that utilizes a virtual boundary. [Figure 12] This diagram illustrates a cross-component adaptive loop filtering (CC-ALF (CCALF)) procedure according to one embodiment of this document. [Figure 13] This document provides a schematic overview of a video / image encoding method and related components as described in the examples. [Figure 14] This document provides a schematic overview of a video / image encoding method and related components as described in the examples. [Figure 15] This document provides a schematic overview of an example of a video decoding method and related components according to the embodiments described herein. [Figure 16] This document provides a schematic overview of an example of a video decoding method and related components according to the embodiments described herein. [Figure 17] Examples of content streaming systems to which the embodiments disclosed in this document can be applied are shown. [Modes for carrying out the invention]

[0030] This document may be modified in various ways and may have various embodiments, but it attempts to illustrate and describe in detail specific embodiments with the drawings. However, this does not mean that this document is intended to limit itself to specific embodiments. The terms used herein are used solely to describe specific embodiments and are not intended to limit the technical ideas of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as "includes" or "has" are intended to indicate the existence of features, figures, stages, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preemptively exclude the possibility of the existence or addition of one or more different features, figures, stages, operations, components, parts, or combinations thereof.

[0031] On the other hand, each configuration shown in the diagrams described in this document is shown independently for the convenience of explaining its distinct characteristic functions, and does not mean that each configuration is embodied in separate hardware or separate software. For example, two or more configurations may be combined to form a single configuration, and a single configuration may be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included within the scope of the rights of this document, as long as they do not deviate from the essence of this document.

[0032] The following describes a preferred embodiment of this document in more detail with reference to the attached diagrams. The same reference numerals may be used for the same components in the drawings, and redundant descriptions of the same components may be omitted.

[0033] This document relates to video / image coding. For example, the methods / executions disclosed in this document are related to the VVC (Versatile Video Coding) standard (ITU-T Rec.H.266), next-generation video / image coding standards after VVC, or other video coding-related standards (e.g., HEVC (High Efficiency Video Coding) standard (ITU-T Rec.H.265), EVC (essential video coding) standard, AVS2 standard, etc.).

[0034] This document presents various examples of video / image coding, and unless otherwise noted, these examples may be combined with each other.

[0035] In this document, "video" can mean a collection of images over time. "Picture" generally refers to a single image representing a specific time period, while "slice" or "tile" is a unit that constitutes part of a picture in coding. A slice or tile can contain one or more CTUs (coding tree units). A single picture can consist of one or more slices or tiles. A single picture can consist of one or more tile groups. A single tile group can consist of one or more tiles.

[0036] A pixel or pel can refer to the smallest unit that makes up a picture (or image). Alternatively, the term "sample" may be used as a counterpart to pixel. A sample can generally represent a pixel or a pixel value, or it can represent only the luma component pixel / pixel value, or only the chroma component pixel / pixel value. Alternatively, a sample may refer to a pixel value in the spatial domain, and when these pixel values ​​are converted to the frequency domain, it can also refer to the conversion coefficient in the frequency domain.

[0037] A unit can represent a basic unit of image processing. A unit can contain at least one of a specific region of a picture and information associated with that region. A unit can contain one luma block and two chroma (e.g., cb, cr) blocks. The term unit may be used interchangeably with terms such as block or area. In general, an M×N block can contain a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.

[0038] In this document, the terms " / " and "," are interpreted as "and / or." For example, "A / B" is interpreted as "A and / or B," and "A, B" is interpreted as "A and / or B." Additionally, "A / B / C" means "at least one of A, B, and / or C." Similarly, "A, B, C" also means "at least one of A, B, and / or C."

[0039] In addition, in this document, “or” is interpreted as “and / or.” For example, “A or B” can mean 1) only “A,” 2) only “B,” or 3) both “A and B.” Alternatively, “or” in this document can mean “additionally or alternatively.”

[0040] In this specification, "at least one of A and B" may mean "just A," "just B," or "both A and B." Furthermore, in this specification, the expressions "at least one of A or B" and "at least one of A and / or B" may be interpreted similarly to "at least one of A and B."

[0041] Furthermore, in this specification, "at least one of A, B and C" may mean "just A," "just B," "just C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" may mean "at least one of A, B and C."

[0042] Furthermore, parentheses used herein may mean "for example." Specifically, when "prediction (intra-prediction)" is indicated, "intra-prediction" may be proposed as an example of "prediction." In other words, "prediction" as used herein is not limited to "intra-prediction," and "intra-prediction" may be proposed as an example of "prediction." Similarly, when "prediction (i.e., intra-prediction)" is indicated, "intra-prediction" may be proposed as an example of "prediction."

[0043] Technical features described individually in each drawing in this specification may be embodied individually or simultaneously.

[0044] Figure 1 schematically shows an example of a video / image coding system to which this document can be applied.

[0045] As shown in Figure 1, a video / image coding system may comprise a source device and a receiving device. The source device can transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or network.

[0046] The source device may comprise a video source, an encoding device, and a transmitter. The receiving device may comprise a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be provided in the encoding device. The receiver may be provided in the decoding device. The renderer may comprise a display unit, which may consist of a separate device or external component.

[0047] A video source can acquire video / images through processes such as video / image capture, synthesis, or generation. A video source may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, or video / image archives containing previously captured video / images. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images may be generated via a computer, in which case the video / image capture process may be replaced by the process of generating the associated data.

[0048] An encoding device can encode input video / image data. For compression and coding efficiency, the encoding device can perform a series of steps, including prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in bitstream format.

[0049] The transmitting unit can transmit encoded video / image information or data output in bitstream format to the receiving unit of a receiving device via a digital storage medium or network in file or streaming format. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit may include elements for generating media files via a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.

[0050] A decoding device can decode video / images by performing a series of steps, such as inverse quantization, inverse transformation, and prediction, corresponding to the operation of an encoding device.

[0051] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0052] Figure 2 is a schematic diagram illustrating the configuration of a video / image encoding device to which this document applies. Hereinafter, the term "video encoding device" may include image encoding devices.

[0053] As shown in Figure 2, the encoding device 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be called a reconstructor or a reconstructed block generator. The aforementioned video splitting unit 210, prediction unit 220, residual processing unit 230, entropy encoding unit 240, addition unit 250, and filtering unit 260 can be configured by one or more hardware components (e.g., an encoder chipset or processor) depending on the embodiment. The memory 270 may also include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.

[0054] The video splitting unit 210 can split the input video (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units can be recursively split from a coding tree unit (CTU) or the largest coding unit (LCU) using a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary-tree structure and / or the ternary structure. Alternatively, the binary-tree structure may be applied first. The coding procedure according to this disclosure may be performed based on the final coding unit that is not further split. In this case, based on coding efficiency due to video characteristics, the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into lower-depth coding units so that the optimally sized coding unit is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further comprise a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit can each be separated or partitioned from the final coding unit described above.The prediction unit may be a unit of sample prediction, and the conversion unit may be a unit for deriving conversion coefficients and / or a unit for deriving a residual signal from conversion coefficients.

[0055] The term "unit" can sometimes be used interchangeably with terms such as "block" or "area." Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the luminance (luma) component pixel / pixel value, or only the chroma component pixel / pixel value. A sample can be used as the term corresponding to a single picture (or image) pixel or pel.

[0056] The subtraction unit 231 can generate a residual signal (residual block, residual sample, or residual sample array) by subtracting the predicted signal (predicted block, predicted sample, or predicted sample array) output from the prediction unit 220 using the input video signal (original block, original sample, or original sample array), and the generated residual signal is transmitted to the conversion unit 232. The prediction unit 220 can make predictions for the block to be processed (hereinafter referred to as the current block) and generate a predicted block that includes predicted samples for the current block. The prediction unit 220 can determine whether intra-prediction or inter-prediction is applied on a current block or CU basis. As will be described later in the explanation of each prediction mode, the prediction unit can generate various information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240. The information related to prediction can be encoded by the entropy encoding unit 240 and output in bitstream format.

[0057] The intra-prediction unit 222 can predict the current block by referring to a sample in the current picture. The referenced sample may be located adjacent to the current block or at a distance, depending on the prediction mode. The prediction mode in intra-prediction may include multiple non-directional modes and multiple directional modes. Non-directional modes may include, for example, DC mode and planar mode. Directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is illustrative, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 222 may also determine the prediction mode to apply to the current block using the prediction modes applied to adjacent blocks.

[0058] The interprediction unit 221 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, adjacent blocks may include spatial neighboring blocks that exist in the current picture and temporal neighboring blocks that exist in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, col CU, etc., and the reference picture containing the temporal neighboring block may be called a collocated picture (colPic). For example, the inter-prediction unit 221 can construct a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter-prediction can be performed based on various prediction modes; for example, in skip mode and merge mode, the inter-prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted.In motion vector prediction (MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.

[0059] The prediction unit 220 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra-prediction (CIIP). The prediction unit can also perform intra-block copy (IBC) for predictions on blocks. The intra-block copy can be used for content video / video coding such as in games, for example, in SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed similarly to inter-prediction in that it derives a reference block within the current picture. That is, IBC can use at least one of the inter-prediction techniques described in this document.

[0060] The prediction signals generated via the interpretation unit 221 and / or intrapretation unit 222 can be used to generate a reconstructed signal or a residual signal. The transformation unit 232 can apply transformation techniques to the residual signal to generate transformation coefficients. For example, transformation techniques may include DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to a transformation obtained from a graph when relational information between pixels is represented by this graph. CNT refers to a transformation obtained by generating a prediction signal using all previously reconstructed pixels and based on that. The transformation process may also be applied to pixel blocks of the same size and square shape, or to non-square blocks of variable size.

[0061] The quantization unit 233 quantizes the conversion coefficients and transmits them to the entropy encoding unit 240, which can encode the quantized signal (information about the quantized conversion coefficients) and output it as a bitstream. The information about the quantized conversion coefficients can be called residual information. The quantization unit 233 can rearrange the block-form quantized conversion coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector form of the quantized conversion coefficients. The entropy encoding unit 240 can perform various encoding methods, such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). In addition to the quantized conversion coefficients, the entropy encoding unit 240 can also encode information necessary for video / image restoration (e.g., the values ​​of syntax elements) together with or separately from the quantized conversion coefficients. Encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form in units of NAL (network abstraction layer) units. The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. In this document, the signaling / transmitted information and / or syntax elements described later may be encoded via the encoding procedure described above and included in the bitstream.The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be transmitted by a transmitting unit (not shown) and / or stored by a storage unit (not shown) which are configured as internal / external elements of the encoding device 200, or the transmitting unit may be included in the entropy encoding unit 240.

[0062] The quantized conversion coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transformation to the quantized conversion coefficients via the inverse quantization unit 234 and the inverse transformation unit 235. The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample, or reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the prediction unit 220. If there is no residual for the block to be processed, as in the case of skip mode, the predicted block can be used as the reconstructed block. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, and can also be used for inter-prediction of the next picture after filtering, as described later.

[0063] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture encoding and / or restoration process.

[0064] The filtering unit 260 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 270, specifically in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter, and bilateral filter. The filtering unit 260 can generate various filtering-related information and transmit it to the entropy encoding unit 240, as will be described later in the description of each filtering method. The filtering-related information can be encoded by the entropy encoding unit 240 and output in bitstream format.

[0065] The corrected restored picture sent to memory 270 can be used as a reference picture in the interpretation unit 221. When interpretation is applied via this, the encoding device can avoid prediction mismatches between the encoding device 100 and the decoding device, and can also improve encoding efficiency.

[0066] The DPB in memory 270 can store the corrected restored picture for use as a reference picture in the inter-prediction unit 221. Memory 270 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 221 for use as motion information of spatially adjacent blocks or motion information of temporally adjacent blocks. Memory 270 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 222.

[0067] Figure 3 is a schematic diagram illustrating the configuration of a video / image decoding device to which this document can be applied.

[0068] As shown in Figure 3, the decoding device 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-prediction unit 331 and an intra-prediction unit 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. The entropy decoder 310, residual processor 320, predictor 330, adder 340, and filtering unit 350 described above can be configured by a single hardware component (e.g., a decoder chipset or processor) depending on the embodiment. The memory 360 may include a decoded picture buffer (DPB) and may also be configured by a digital storage medium. The aforementioned hardware component may also further include memory 360 as an internal / external component.

[0069] When a bitstream containing video / image information is input, the decoding device 300 can reconstruct the image in accordance with the process by which the video / image information was processed in the encoding device shown in Figure 3. For example, the decoding device 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the decoding processing unit can be, for example, a coding unit, which can be divided from a coding tree unit or a maximum coding unit according to a quad-tree structure, a binary tree structure, and / or a terminally tree structure. One or more conversion units can be derived from the coding unit. The reconstructed video signal decoded and output via the decoding device 300 can then be played back via a playback device.

[0070] The decoding device 300 can receive the signal output from the encoding device shown in Figure 3 in bitstream form, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information necessary for video restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The decoding device can further decode the picture based on the parameter set information and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values ​​of syntax elements necessary for image restoration and the quantized values ​​of conversion coefficients related to the residual. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded and the decoding information of adjacent and decoded blocks or symbol / bin information decoded in a previous step, predicts the probability of bin occurrence based on the determined context model, performs arithmetic decoding of the bins, and generates symbols corresponding to the values ​​of each syntax element.In this case, the CABAC entropy decoding method can update the context model after determining the context model by utilizing the decoded symbol / bin information for the context model of the next symbol / bin. Of the information decoded by the entropy decoding unit 310, information related to prediction is provided to the prediction unit 330, and residual information that has been entropy decoded by the entropy decoding unit 310, i.e., quantized conversion coefficients and related parameter information, can be input to the inverse quantization unit 321. In addition, of the information decoded by the entropy decoding unit 310, information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit can be a component of the entropy decoding unit 310. On the other hand, the decoding device relating to this document may be called a video / image / picture decoding device, and the decoding device may also be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, inverse transformation unit 322, prediction unit 330, addition unit 340, filtering unit 350, and memory 360.

[0071] The inverse quantization unit 321 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 321 can rearrange the quantized transformation coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain the transformation coefficients.

[0072] In the inverse conversion unit 322, the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).

[0073] The prediction unit can make predictions for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 310, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block and can determine a specific intra / inter-prediction mode.

[0074] The prediction unit can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra-prediction (CIIP). The prediction unit can also perform intra-block copying (IBC) for predictions on blocks. This intra-block copying can be used for content video / video coding such as in games, for example, as in SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed similarly to inter-prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction.

[0075] The intra-prediction unit 331 can predict the current block by referring to a sample in the current picture. The referenced sample can be located adjacent to or far from the current block depending on the prediction mode. In intra-prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra-prediction unit 331 can also determine the prediction mode to be applied to the current block using the prediction modes applied to adjacent blocks.

[0076] The interprediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in blocks, subblocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, adjacent blocks may include spatially adjacent blocks that exist in the current picture and temporally adjacent blocks that exist in the reference picture. For example, the interprediction unit 332 can construct a motion information candidate list based on adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction can be performed based on various prediction modes, and the prediction information may include information indicating the mode of interprediction for the current block.

[0077] The summing unit 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit. If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the restored block.

[0078] The addition unit 340 may be called the restoration unit or restoration block generation unit. The generated restoration signal can be used for intra-prediction of the next block to be processed in the current picture, and can be output after filtering as described later, or it can be used for intra-prediction of the next picture.

[0079] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.

[0080] The filtering unit 350 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.

[0081] The (modified) restored picture stored in the DPB of memory 360 can be used as a reference picture by the inter-prediction unit 332. Memory 360 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 332 for use as motion information of spatially adjacent blocks or motion information of temporally adjacent blocks. Memory 360 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 331.

[0082] In this specification, the embodiments described for the prediction unit 330, inverse quantization unit 321, inverse conversion unit 322, and filtering unit 350 of the decoding device 300 can be applied identically or in a corresponding manner to the prediction unit 220, inverse quantization unit 234, inverse conversion unit 235, and filtering unit 260 of the encoding device 200, respectively.

[0083] As mentioned above, prediction is performed to improve compression efficiency when performing video coding. Through this, a predicted block containing predicted samples for the current block, which is the block to be coded, can be generated. Here, the predicted block contains predicted samples in the spatial domain (or pixel domain). The predicted block is derived in both the encoding and decoding devices, and the encoding device can improve video coding efficiency by signaling the decoding device with information (residual information) about the residual between the original block and the predicted block, which is not the original sample value of the original block itself. The decoding device can derive a residual block containing residual samples based on the residual information, and can generate a restored block containing restored samples by combining the residual block and the predicted block, and can generate a restored picture containing the restored block.

[0084] The residual information can be generated through transformation and quantization procedures. For example, an encoding device can signal the relevant residual information (via a bitstream) to a decoding device by deriving a residual block between the original block and the predicted block, performing a transformation procedure on the residual samples (residual sample array) contained in the residual block to derive transformation coefficients, and performing a quantization procedure on the transformation coefficients to derive quantized transformation coefficients. Here, the residual information may include information such as the value information, position information, transformation technique, transformation kernel, and quantization parameters of the quantized transformation coefficients. The decoding device can derive a residual sample (or residual block) by performing an inverse quantization / inverse transformation procedure based on the residual information. The decoding device can generate a reconstructed picture based on the predicted block and the residual block. The encoding device can also derive a residual block by inverse quantization / inverse transformation of the quantized transformation coefficients for reference for subsequent interpretation of the picture, and generate a reconstructed picture based on this.

[0085] In this document, at least one of quantization / inverse quantization and / or transformation / inverse transformation may be omitted. If quantization / inverse quantization is omitted, the quantized transformation coefficients may be called transformation coefficients. If transformation / inverse transformation is omitted, the transformation coefficients may also be called coefficients or residual coefficients, or for consistency of expression, they may still be called transformation coefficients.

[0086] In this document, quantized transformation coefficients and transformation coefficients may be referred to as transformation coefficients and scaled transformation coefficients, respectively. In this case, residual information may include information about the transformation coefficients, and such information about the transformation coefficients may be signaled via residual coding syntax. Based on the residual information (or information about the transformation coefficients), transformation coefficients may be derived, and scaled transformation coefficients may be derived via inverse transformation (scaling) of the transformation coefficients. Based on inverse transformation (transformation) of the scaled transformation coefficients, residual samples may be derived. This may be applied / expressed similarly in other parts of this document.

[0087] The prediction unit of the encoding / decoding device can perform interpretation on a block-by-block basis to derive predicted samples. Interpretation can indicate predictions derived in a manner dependent on data elements (e.g., sample values ​​or motion information) of pictures other than the current picture. When interpretation is applied to the current block, a predicted block (predicted sample array) for the current block can be derived based on the reference block (reference sample array) identified by the motion vector on the reference picture pointed to by the index of the reference picture. In this case, in order to reduce the amount of motion information transmitted in interpretation mode, the motion information of the current block can be predicted on a block, subblock, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information may include the motion vector and the index of the reference picture. The motion information may further include information on the interpretation type (L0 prediction, L1 prediction, Bi prediction, etc.). When interpretation is applied, neighboring blocks may include spatial neighboring blocks that exist in the current picture and temporal neighboring blocks that exist in the reference picture. The reference picture containing the aforementioned reference block and the reference picture containing the aforementioned temporally adjacent block may be the same or different. The temporally adjacent block may be referred to by names such as collocated reference block, colCU, etc., and the reference picture containing the aforementioned temporally adjacent block may be referred to as collocated picture (colPic). For example, a candidate list of motion information can be constructed based on the adjacent blocks of the current block, and flags or index information can be signaled to indicate which candidate is selected (used) in order to derive the motion vector and / or index of the reference picture of the current block.Interpretation is performed based on various prediction modes. For example, in skip mode and merge mode, the motion information of the current block may be identical to the motion information of the selected adjacent block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected adjacent block can be used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.

[0088] The motion information may include L0 motion information and / or L1 motion information depending on the interpretation type (L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction may be called an L0 motion vector or MVL0, and a motion vector in the L1 direction may be called an L1 motion vector or MVL1. A prediction based on an L0 motion vector may be called an L0 prediction, a prediction based on an L1 motion vector may be called an L1 prediction, and a prediction based on both the L0 motion vector and the L1 motion vector may be called a bi (Bi) prediction. Here, an L0 motion vector may represent a motion vector associated with a reference picture list L0 (L0), and an L1 motion vector may represent a motion vector associated with a reference picture list L1 (L1). The reference picture list L0 may include pictures earlier in the output order than the current picture, and the reference picture list L1 may include pictures later in the output order than the current picture. The aforementioned earlier picture may be called a forward (reference) picture, and the aforementioned later picture may be called a reverse (reference) picture. The reference picture list L0 may include further reference pictures that are later in the output order than the current picture. In this case, the earlier picture may be indexed first in the reference picture list L0, and the later picture may be indexed afterward. The reference picture list L1 may include further reference pictures that are earlier in the output order than the current picture. In this case, the later picture may be indexed first in the reference picture list L1, and the earlier picture may be indexed afterward. Here, the output order may correspond to the POC (picture order count) order.

[0089] Figure 4 illustrates the hierarchical structure for coded video.

[0090] Referring to Figure 4, coded video is divided into the VCL (video coding layer), which handles the video decoding process and the video itself; a lower-level system that transmits and stores the coded information; and the NAL (network abstraction layer), which exists between the VCL and the lower-level system and is responsible for network adaptation functions.

[0091] VCL can generate VCL data containing compressed video data (slice data), or generate parameter sets containing information such as Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS), or SEI (Supplemental Enhancement Information) messages that are additionally necessary during the video decoding process.

[0092] In NAL, a NAL unit can be generated by adding header information (NAL unit header) to the RBSP (Raw Byte Sequence Payload) generated by VCL. In this case, the RBSP refers to the slice data, parameter set, SEI message, etc., generated by VCL. The NAL unit header can include NAL unit type information, which is identified by the RBSP data contained in the NAL unit.

[0093] As shown in the above diagram, NAL units can be divided into VCL NAL units and Non-VCL NAL units by the RBSP generated in VCL. A VCL NAL unit can mean a NAL unit that contains information about the video (slice data), and a Non-VCL NAL unit can mean a NAL unit that contains information necessary for decoding the video (parameter set or SEI message).

[0094] The aforementioned VCL NAL units and Non-VCL NAL units can be transmitted over a network with header information added according to the data standards of the lower-level system. For example, NAL units can be transformed into data formats of predetermined standards such as H.266 / VVC file format, RTP (Real-time Transport Protocol), and TS (Transport Stream) and transmitted over various networks.

[0095] As mentioned above, the NAL unit type can be identified by the RBSP data structure contained within the NAL unit, and information about such NAL unit types can be stored in the NAL unit header and signaled.

[0096] For example, NAL units can be broadly classified into VCL NAL unit types and Non-VCL NAL unit types depending on whether or not they contain information (slice data) related to the video. VCL NAL unit types can be further classified by the nature and type of picture they contain, while Non-VCL NAL unit types can be further classified by the type of parameter set.

[0097] The following is an example of a NAL unit type identified by the type of parameter set included in the Non-VCL NAL unit type.

[0098] -APS (Adaptation Parameter Set) NAL unit: Type for NAL units that include APS

[0099] -DPS (Decoding Parameter Set) NAL unit: Type for NAL units including DPS

[0100] -VPS (Video Parameter Set) NAL unit: Type for NAL unit including VPS

[0101] -SPS (Sequence Parameter Set) NAL unit: Type for NAL units that include SPS

[0102] -PPS (Picture Parameter Set) NAL unit: Type for NAL units that include PPS

[0103] -PH (Picture header) NAL unit: Type for NAL units that include PH

[0104] The aforementioned NAL unit type has syntax information for the NAL unit type, and this syntax information can be stored in the NAL unit header and signaled. For example, the syntax information is nal_unit_type, and the NAL unit type can be identified by the nal_unit_type value.

[0105] On the other hand, as mentioned above, a single picture can contain multiple slices, and a single slice can contain a slice header and slice data. In this case, a single picture header can be added to each of the multiple slices (slice headers and slice data sets) within a single picture. The picture header (picture header syntax) can contain information / parameters that are commonly applicable to the picture. In this document, slices can be mixed with or replaced by tile groups. Also in this document, slice headers can be mixed with or replaced by type group headers.

[0106] The slice header (slice header syntax, slice header information) may include information / parameters that can be commonly applied to the slice. The APS (APS syntax) or PPS (PPS syntax) may include information / parameters that can be commonly applied to one or more slices or pictures. The SPS (SPS syntax) may include information / parameters that can be commonly applied to one or more sequences. The VPS (VPS syntax) may include information / parameters that can be commonly applied to multiple layers. The DPS (DPS syntax) may include information / parameters that can be commonly applied to video in general. The DPS may include information / parameters related to the concatenation of CVS (coded video sequence). In this document, High-level syntax (HLS) may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, and slice header syntax.

[0107] In this document, the video information encoded from the encoding device to the decoding device and signaled in bitstream form includes not only partitioning-related information, intra / inter prediction information, residual information, and in-loop filtering information within the picture, but may also include information contained in the slice header, the picture header, the APS, the PPS, the SPS, the VPS, and / or the DPS. Furthermore, the video information may further include information from the NAL unit header.

[0108] On the other hand, to compensate for differences between the original and restored video due to errors that occur during the compression encoding process, such as quantization, an in-loop filtering procedure can be performed on the restored sample or restored picture, as described above. As described above, in-loop filtering can be performed in the filter section of the encoding device and the filter section of the decoding device, and a deblocking filter, SAO, and / or adaptive loop filter (ALF) can be applied. For example, the ALF procedure can be performed after the deblocking filtering procedure and / or SAO procedure are completed. However, even in this case, the deblocking filtering procedure and / or SAO procedure may be omitted.

[0109] The following provides a detailed explanation of picture restoration and filtering. In video coding, a restored block can be generated based on intra-prediction / inter-prediction for each block, and a restored picture containing the restored block can be generated. If the current picture / slice is an I-picture / slice, the blocks contained in the current picture / slice can be restored based solely on intra-prediction. On the other hand, if the current picture / slice is a P or B-picture / slice, the blocks contained in the current picture / slice can be restored based on either intra-prediction or inter-prediction. In this case, intra-prediction may be applied to some blocks within the current picture / slice, while inter-prediction may be applied to the remaining blocks.

[0110] Intra prediction can represent a prediction that generates prediction samples for the current block based on reference samples within the picture to which the current block belongs (hereinafter referred to as the current picture). When intra prediction is applied to the current block, adjacent reference samples to be used for intra prediction of the current block can be derived. The adjacent reference samples of the current block may include samples adjacent to the left boundary of the current block of size nW × nH and a total of 2 × nH samples adjacent to the bottom left, samples adjacent to the top boundary of the current block and a total of 2 × nW samples adjacent to the top right, and one sample adjacent to the top left of the current block. Alternatively, the adjacent reference samples of the current block may include multiple columns of upper adjacent samples and multiple rows of left adjacent samples. Furthermore, the adjacent reference samples of the current block may also include a total of nH samples adjacent to the right boundary of the current block, which is of size nW × nH, a total of nW samples adjacent to the bottom boundary of the current block, and one sample adjacent to the bottom-right side of the current block.

[0111] However, some of the adjacent reference samples in the current block may not yet be decoded or available. In this case, the decoder can construct adjacent reference samples to be used for prediction by substituting the unavailable samples as available samples, or by constructing adjacent reference samples to be used for prediction through interpolation of available samples.

[0112] If neighboring reference samples are derived, (i) predicted samples can be derived based on the average or interpolation of neighboring reference samples in the current block, or (ii) predicted samples can be derived based on reference samples in the current block that are located in a specific (predicted) direction relative to the predicted sample. Case (i) is called the non-directional mode or non-angular mode, and case (ii) is called the directional mode or angular mode. Alternatively, the predicted sample can be generated by interpolation between the first neighboring sample and a second neighboring sample located in the opposite direction to the prediction direction of the current block's intra-prediction mode, based on the predicted sample of the current block. In this case, it can be called linear interpolation intra-prediction (LIP). Alternatively, chroma predicted samples can be generated based on chroma samples using a linear model. In this case, it can be called the LM mode. Alternatively, a temporary predicted sample for the current block can be derived based on filtered adjacent reference samples, and the predicted sample for the current block can be derived by performing a weighted sum of the temporary predicted sample and at least one reference sample derived by the intra-prediction mode from the existing adjacent reference samples, i.e., unfiltered adjacent reference samples. In the above case, it can be called PDPC (Position dependent intra prediction). In addition, intra-predictive coding can be performed by selecting the reference sample line with the highest prediction accuracy from the adjacent multiple reference sample lines of the current block, deriving the predicted sample using the reference sample located in the prediction direction on that line, and instructing (signaling) the decoding device with the reference sample line used at this time.In the aforementioned cases, this can be called multi-reference line (MRL) intra prediction or MRL-based intra prediction. Furthermore, the current block can be divided into vertical or horizontal subpartitions, and intra prediction can be performed based on the same intra prediction mode, with adjacent reference samples derived and available for use on a subpartition basis. That is, in this case, the intra prediction mode for the current block is also applied to the subpartition, and by deriving and using adjacent reference samples on a subpartition basis, intra prediction performance can be improved in some cases. Such prediction methods can be called intra subpartitions (ISP) or ISP-based intra prediction. The aforementioned intra prediction methods can be distinguished from the intra prediction modes in Table of Contents 1 and 2 and referred to as intra prediction types. These intra prediction types can be referred to by various terms, such as intra prediction techniques or additional intra prediction modes. For example, the intra prediction type (or additional intra prediction mode, etc.) may include at least one of the aforementioned LIP, PDPC, MRL, and ISP. A general intra-prediction method that excludes specific intra-prediction types such as LIP, PDPC, MRL, and ISP can be called a normal intra-prediction type. The normal intra-prediction type can be generally applied when the aforementioned specific intra-prediction types are not applicable, and predictions can be performed based on the intra-prediction modes described above. Meanwhile, post-processing filtering can be performed on the derived prediction samples as needed.

[0113] Specifically, the intra-prediction procedure may include an intra-prediction mode / type determination step, an adjacent reference sample derivation step, and an intra-prediction mode / type-based predictive sample derivation step. Additionally, a post-filtering step may be performed on the derived predictive samples as needed.

[0114] Figure 5 is a flowchart illustrating an intra-prediction-based block reconstruction method in an encoding device. The method in Figure 5 may include steps S500, S510, S520, S530, and S540.

[0115] S500 can be performed by the intra-prediction unit 222 of the encoding device, and S510 to S530 can be performed by the residual processing unit 230 of the encoding device. Specifically, S510 can be performed by the subtraction unit 231 of the encoding device, S520 can be performed by the conversion unit 232 and quantization unit 233 of the encoding device, and S530 can be performed by the inverse quantization unit 234 and inverse conversion unit 235 of the encoding device. In S500, prediction information is derived by the intra-prediction unit 222 and can be encoded by the entropy encoding unit 240. Residual information is derived via S510 and S520 and can be encoded by the entropy encoding unit 240. The residual information is information about the residual sample. The residual information may include information about the quantized conversion coefficients for the residual sample. As described above, the residual sample can be derived as a conversion coefficient via the conversion unit 232 of the encoding device, and the conversion coefficient can be derived as a quantized conversion coefficient via the quantization unit 233. Information regarding the quantized conversion coefficient can be encoded in the entropy encoding unit 240 via the residual coding procedure.

[0116] The encoding device performs intra-prediction for the current block (S500). The encoding device can derive an intra-prediction mode for the current block, derive adjacent reference samples for the current block, and generate predicted samples within the current block based on the intra-prediction mode and the adjacent reference samples. Here, the intra-prediction mode determination, adjacent reference sample derivation, and predicted sample generation procedures can be performed simultaneously, or one procedure can be performed before another. For example, the intra-prediction unit 222 of the encoding device may include a prediction mode / type determination unit, a reference sample derivation unit, and a predicted sample derivation unit, where the prediction mode / type determination unit determines the intra-prediction mode / type for the current block, the reference sample derivation unit derives adjacent reference samples for the current block, and the predicted sample derivation unit derives motion samples for the current block. On the other hand, although not shown, if a predicted sample filtering procedure described later is performed, the intra-prediction unit 222 may further include a predicted sample filtering unit (not shown). The encoding device can determine which of a plurality of intra-prediction modes is to be applied to the current block. The encoding device can determine the optimal intra-prediction mode for the current block by comparing the RD costs for the intra-prediction modes.

[0117] On the other hand, the encoding device can also perform a predictive sample filtering procedure. This predictive sample filtering may be called post-filtering. The predictive sample filtering procedure may filter out some or all of the predictive samples. In some cases, the predictive sample filtering procedure may be omitted.

[0118] The encoding device derives a residual sample for the current block based on the predicted sample (S510). The encoding device can derive the residual sample by comparing the predicted sample with the original sample of the current block based on phase.

[0119] The encoding device converts / quantizes the residual sample and derives the quantized conversion coefficients (S520). Subsequently, the quantized conversion coefficients can be inversely quantized / inversely transformed again to derive the (corrected) residual sample (S530). The reason for performing inverse quantization / inverse transformation again after conversion / quantization is, as mentioned above, to derive the same residual sample as the residual sample derived by the decoding device.

[0120] The encoding device can generate a restored block containing restored samples for the current block based on the predicted samples and the (modified) residual samples (S540). Based on the restored block, a restored picture can be generated for the current picture.

[0121] As described above, the encoding device can encode video information including prediction information related to the intra prediction (for example, prediction mode information indicating the prediction mode), and residual information related to the intra and the residual sample, and output the encoded video information in bitstream format. The residual information may include residual coding syntax. The encoding device can convert / quantize the residual sample and derive quantized conversion coefficients. The residual information may include information relating to the quantized conversion coefficients.

[0122] Figure 6 is a flowchart illustrating an intra-prediction-based block reconstruction method in a decoding device. The method in Figure 6 may include steps S600, S610, S620, S630, and S640. The decoding device can perform operations corresponding to the operations performed in the encoding device.

[0123] S600 to S620 can be performed by the intra-prediction unit 331 of the decoding device, and the prediction information in S600 and the residual information in S630 can be obtained from the bitstream by the entropy decoding unit 310 of the decoding device. The residual processing unit 320 of the decoding device can derive a residual sample for the current block based on the residual information. Specifically, the inverse quantization unit 321 of the residual processing unit 320 can derive conversion coefficients by performing inverse quantization based on the quantized conversion coefficients derived from the residual information, and the inverse transformation unit 322 of the residual processing unit can derive a residual sample for the current block by performing an inverse transformation on the conversion coefficients. S640 can be performed by the addition unit 340 or the restoration unit of the decoding device.

[0124] Specifically, the decoding device can derive an intra-prediction mode for the current block based on the received prediction mode information (S600). The decoding device can derive adjacent reference samples for the current block (S610). The decoding device generates prediction samples within the current block based on the intra-prediction mode and the adjacent reference samples (S620). In this case, the decoding device can perform a prediction sample filtering procedure. Prediction sample filtering may be called post-filtering. The prediction sample filtering procedure may filter some or all of the prediction samples. Depending on the circumstances, the prediction sample filtering procedure may be omitted.

[0125] The decoding device generates a residual sample for the current block based on the received residual information (S630). The decoding device can generate a restored sample for the current block based on the predicted sample and the residual sample, and derive a restored block containing the restored sample (S640). A restored picture for the current picture can be generated based on the restored block.

[0126] Here, the intra-prediction unit 331 of the decoding device may include a prediction mode / type determination unit, a reference sample derivation unit, and a prediction sample derivation unit. The prediction mode / type determination unit determines the intra-prediction mode for the current block based on the prediction mode information acquired by the entropy decoding unit 310 of the decoding device. The reference sample derivation unit derives adjacent reference samples for the current block, and the prediction sample derivation unit derives prediction samples for the current block. On the other hand, although not shown, if the prediction sample filtering procedure described above is performed, the intra-prediction unit 331 may further include a prediction sample filter unit (not shown).

[0127] The prediction information may include intra-prediction mode information and / or intra-prediction type information. The intra-prediction mode information may include, for example, flag information (e.g., intra_luma_mpm_flag) indicating whether MPM (most probable mode) is applied to the current block or whether remaining mode is applied. If MPM is applied to the current block, the prediction mode information may further include index information (e.g., intra_luma_mpm_idx) pointing to one of the intra-prediction mode candidates (MPM candidates). The intra-prediction mode candidates (MPM candidates) may consist of an MPM candidate list or an MPM list. If MPM is not applied to the current block, the intra-prediction mode information may further include remaining mode information (e.g., intra_luma_mpm_remainder) pointing to one of the remaining intra-prediction modes excluding the intra-prediction mode candidates (MPM candidates). The decoding device can determine the intra-prediction mode of the current block based on the intra-prediction mode information. A separate MPM list can be configured for the aforementioned MIP.

[0128] Furthermore, the intra-prediction type information can be embodied in various forms. For example, the intra-prediction type information may include intra-prediction type index information indicating one of the intra-prediction types. As another example, the intra-prediction type information may include reference sample line information (e.g., intra_luma_ref_idx) indicating whether the MRL is applied to the current block and, if so, which reference sample line is used; ISP flag information (e.g., intra_subpartitions_mode_flag) indicating whether the ISP is applied to the current block; ISP type information (e.g., intra_subpartitions_split_flag) indicating the split type of the subpartition if the ISP is applied; and at least one of the following flag information: flag information indicating whether PDCP is applicable or flag information indicating whether LIP is applicable. The intra-prediction type information may also include an MIP flag indicating whether MIP is applied to the current block.

[0129] The intra-prediction mode information and / or the intra-prediction type information can be encoded / decoded via the coding methods described in this document. For example, the intra-prediction mode information and / or the intra-prediction type information can be encoded / decoded via entropy coding (e.g., CABAC, CAVLC coding) based on truncated (rice) binary code.

[0130] The prediction unit of the encoding / decoding device can perform inter-prediction on a block-by-block basis to derive predicted samples. Inter-prediction can be a prediction derived in a manner that is dependent on data elements (e.g., sample values ​​or motion information) of picture(s) other than the current picture. When inter-prediction is applied to the current block, a predicted block (predicted sample array) for the current block can be derived based on the reference block (reference sample array) identified by the motion vector on the reference picture pointed to by the reference picture index. At this time, in order to reduce the amount of motion information transmitted in inter-prediction mode, the motion information of the current block can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indexes. The motion information may further include inter-prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. When inter-prediction is applied, an adjacent block can include a spatial neighboring block currently present in the picture and a temporal neighboring block present in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to as a collocated reference block or colCU, and the reference picture containing the temporal neighboring block may be referred to as a collocated picture (colPic).For example, a list of motion information candidates can be constructed based on the adjacent blocks of the current block, and flags or index information can be signaled indicating which candidates are selected (used) to derive the motion vector and / or reference picture index of the current block. Interpretation can be performed based on various prediction modes; for example, in skip mode and merge mode, the motion information of the current block is the same as the motion information of the selected adjacent block. In skip mode, unlike merge mode, no residual signal is transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected adjacent block is used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.

[0131] Figure 7 is a flowchart illustrating an interpretation-based block reconstruction method in an encoding device. The method in Figure 7 may include steps S700, S710, S720, S730, and S740.

[0132] S700 can be performed by the interprediction unit 221 of the encoding device, and S710 to S730 can be performed by the residual processing unit 230 of the encoding device. Specifically, S710 can be performed by the subtraction unit 231 of the encoding device, S720 can be performed by the conversion unit 232 and quantization unit 233 of the encoding device, and S730 can be performed by the inverse quantization unit 234 and inverse conversion unit 235 of the encoding device. In S700, prediction information is derived by the interprediction unit 221 and can be encoded by the entropy encoding unit 240. Residual information is derived via S710 and S720 and can be encoded by the entropy encoding unit 240. The residual information is information about the residual sample. The residual information may include information about the quantized conversion coefficients for the residual sample. As described above, the residual sample can be derived as a conversion coefficient via the conversion unit 232 of the encoding device, and the conversion coefficient can be derived as a quantized conversion coefficient via the quantization unit 233. Information regarding the quantized conversion coefficient can be encoded in the entropy encoding unit 240 via the residual coding procedure.

[0133] The encoding device performs interpretation for the current block (S700). The encoding device can derive the interpretation mode and motion information of the current block and generate a prediction sample of the current block. Here, the interpretation mode determination, motion information derivation, and prediction sample generation procedures can be performed simultaneously, or one procedure can be performed before another. For example, the interpretation unit 221 of the encoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit, where the prediction mode determination unit determines the prediction mode for the current block, the motion information derivation unit derives the motion information of the current block, and the prediction sample derivation unit derives a motion sample of the current block. For example, the interpretation unit 221 of the encoding device can search for blocks similar to the current block within a certain area (search area) of the reference picture via motion estimation and derive a reference block whose difference from the current block is minimal or below a certain standard. Based on this, a reference picture index pointing to the reference picture in which the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding device can determine which mode to apply to the current block from among a variety of prediction modes. The encoding device can determine the optimal prediction mode for the current block by comparing the RD costs for the various prediction modes.

[0134] For example, when skip mode or merge mode is applied to the current block, the encoding device can configure a merge candidate list, as described later, and derive a reference block from among the reference blocks pointed to by the merge candidates included in the merge candidate list whose difference between the current block and the previous current block is minimal or below a certain standard. In this case, a merge candidate associated with the derived reference block is selected, and merge index information pointing to the selected merge candidate is generated and signaled to the decoding device. The movement information of the current block can be derived using the movement information of the selected merge candidate.

[0135] As another example, when the (A)MVP mode is applied to the current block, the encoding device can configure the (A)MVP candidate list described later, and use the motion vector of the mvp (motion vector predictor) candidate selected from the mvp candidates included in the (A)MVP candidate list as the mvp of the current block. In this case, for example, the motion vector pointing to the reference block derived by the motion estimation described above can be used as the motion vector of the current block, and the mvp candidate having the motion vector with the smallest difference from the motion vector of the current block can become the selected mvp candidate. The MVD (motion vector difference), which is the difference obtained by subtracting the mvp from the motion vector of the current block, can be derived. In this case, information regarding the MVD can be signaled to the decoding device. Also, when the (A)MVP mode is applied, the value of the reference picture index can be configured with reference picture index information and separately signaled to the decoding device.

[0136] The encoding device can derive a residual sample based on the predicted sample (S710). The encoding device can derive the residual sample by comparing the original sample of the current block with the predicted sample.

[0137] The encoding device converts / quantizes the residual sample and derives the quantized conversion coefficients (S720). Subsequently, the quantized conversion coefficients can be inversely quantized / inversely transformed again to derive the (corrected) residual sample (S730). The reason for performing inverse quantization / inverse transformation again after conversion / quantization is, as mentioned above, to derive the same residual sample as the residual sample derived by the decoding device.

[0138] The encoding device can generate a restored block containing restored samples for the current block based on the predicted samples and the (modified) residual samples (S740). Based on the restored block, a restored picture can be generated for the current picture.

[0139] Although not shown in the diagram, as described above, the encoding device can encode video information including prediction information and residual information. The encoding device can output the encoded video information in bitstream format. The prediction information is information related to the prediction procedure and may include prediction mode information (e.g., skip flag, merge flag, or mode index, etc.) and motion information. The motion information may include candidate selection information (e.g., merge index, mvp flag, or mvp index) which is information for deriving a motion vector. The motion information may also include the aforementioned MVD information and / or reference picture index information. Furthermore, the motion information may include information indicating whether L0 prediction, L1 prediction, or paired (bi) prediction is applied. The residual information is information about the residual sample. The residual information may include information about the quantized conversion coefficients for the residual sample.

[0140] The output bitstream can be stored in a (digital) storage medium and transmitted to a decoding device, or it can be transmitted to a decoding device via a network.

[0141] Figure 8 is a flowchart illustrating an interpretation-based block reconstruction method in a decoding device. The method in Figure 8 may include steps S800, S810, S820, S830, and S840. The decoding device can perform operations corresponding to the operations performed in the encoding device.

[0142] S800 to S820 can be performed by the interpretation unit 332 of the decoding device, and the prediction information of S800 and the residual information of S830 can be obtained from the bitstream by the entropy decoding unit 310 of the decoding device. The residual processing unit 320 of the decoding device can derive a residual sample for the current block based on the residual information. Specifically, the inverse quantization unit 321 of the residual processing unit 320 can perform inverse quantization to derive conversion coefficients based on the quantized conversion coefficients derived based on the residual information, and the inverse transformation unit 322 of the residual processing unit can perform an inverse transformation on the conversion coefficients to derive a residual sample for the current block. S840 can be performed by the addition unit 340 or the restoration unit of the decoding device.

[0143] Specifically, the decoding device can determine the prediction mode for the current block based on the received prediction information (S800). The decoding device can determine which inter-prediction mode is applied to the current block based on the prediction mode information in the prediction information.

[0144] For example, based on the merge flag, it can be determined whether the merge mode is applied to the current block, or whether the (A)MVP mode is determined. Alternatively, one can be selected from a variety of inter-prediction mode candidates based on the mode index. The inter-prediction mode candidates may include skip mode, merge mode, and / or (A)MVP mode, or may include a variety of inter-prediction modes as described later.

[0145] The decoding device derives motion information for the current block based on the determined inter prediction mode (S810). For example, if a skip mode or merge mode is applied to the current block, the decoding device can configure a merge candidate list, as described later, and select one merge candidate from among the merge candidates included in the merge candidate list. This selection can be performed based on the selection information (merge index) described above. The motion information for the current block can be derived using the motion information for the selected merge candidate. The motion information for the selected merge candidate can be used as the motion information for the current block.

[0146] As another example, when the (A)MVP mode is applied to the current block, the decoding device can configure the (A)MVP candidate list described later, and use the motion vector of the mvp candidate selected from the mvp (motion vector predictor) candidates included in the (A)MVP candidate list as the mvp of the current block. The selection can be performed based on the selection information (mvp flag or mvp index) described above. In this case, the MVD of the current block can be derived based on the information regarding the MVD, and the motion vector of the current block can be derived based on the mvp of the current block and the MVD. Furthermore, the reference picture index of the current block can be derived based on the reference picture index information. The picture pointed to by the reference picture index in the reference picture list for the current block can be derived as the reference picture referenced for interpretation of the current block.

[0147] On the other hand, as will be described later, the movement information of the current block can be derived without constructing a candidate list, in which case the movement information of the current block can be derived by the procedure disclosed in the prediction mode described later. In this case, the candidate list configuration described above can be omitted.

[0148] The decoding device can generate predicted samples for the current block based on the motion information of the current block (S820). In this case, the reference picture can be derived based on the reference picture index of the current block, and the predicted samples for the current block can be derived using the sample of the reference block pointed to by the motion vector of the current block on the reference picture. In this case, as will be described later, a prediction sample filtering procedure may be further performed on all or some of the predicted samples for the current block.

[0149] For example, the interpretation unit 332 of the decoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit determines the prediction mode for the current block based on the prediction mode information received, the motion information derivation unit derives motion information (motion vector and / or reference picture index, etc.) for the current block based on the motion information received, and the prediction sample derivation unit derives the prediction sample for the current block.

[0150] The decoding device generates a residual sample for the current block based on the received residual information (S830). The decoding device can generate a restored sample for the current block based on the predicted sample and the residual sample, and derive a restored block containing the restored sample (S840). A restored picture for the current picture can be generated based on the restored block.

[0151] A variety of interpretation modes can be used to predict the current block within a picture. For example, various modes such as merge mode, skip mode, MVP (motion vector prediction) mode, affine mode, subblock merge mode, and MMVD (merge with MVD) mode can be used. DMVR (Decoder side motion vector refinement) mode, AMVR (adaptive motion vector resolution) mode, Bi-prediction with CU-level weight (BCW), and Bi-directional optical flow (BDOF) can be used as additional or alternative modes. The affine mode is sometimes called affine motion prediction mode. The MVP mode is sometimes called AMVP (advanced motion vector prediction) mode. In this document, candidate motion information derived from some modes and / or some modes may be included as one of the candidate motion information for other modes. For example, an HMVP candidate may be added as a merge candidate in the merge / skip modes, or as an mvp candidate in the MVP mode.

[0152] Prediction mode information indicating the inter-prediction mode of the current block can be signaled from the encoding device to the decoding device. The prediction mode information can be included in the bitstream and received by the decoding device. The prediction mode information may include index information indicating one of a number of candidate modes. Alternatively, the inter-prediction mode may be indicated through hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, a skip flag may be signaled to indicate whether the skip mode is applicable, and if the skip mode is not applicable, a merge flag may be signaled to indicate whether the merge mode is applicable, and if the merge mode is not applicable, it may indicate that the MVP mode is applicable, or further flags for additional distinctions may be signaled. Affine modes may be signaled as independent modes, or as modes dependent on the merge mode or MVP mode, etc. For example, affine modes may include affine merge mode and affine MVP mode.

[0153] On the other hand, the current block may be signaled with information indicating whether the aforementioned list0 (L0) prediction, list1 (L1) prediction, or bi-prediction prediction is used in the current block (current coding unit). This information may be called motion prediction direction information, inter-prediction direction information, or inter-prediction instruction information, and can be composed / encoded / signaled, for example, in the form of an inter_pred_idc syntax element. That is, the inter_pred_idc syntax element can indicate whether the aforementioned list0 (L0) prediction, list1 (L1) prediction, or bi-prediction prediction is used in the current block (current coding unit). For the sake of explanation, in this document, the inter-prediction type (L0 prediction, L1 prediction, or BI prediction) pointed to by the inter_pred_idc syntax element may be expressed as motion prediction direction. L0 prediction may also be expressed as pred_L0, L1 prediction as pred_L1, and bi-prediction as pred_BI. For example, depending on the value of the syntax element in inter_pred_idc, the prediction type can be determined as shown in the following table.

[0154] [Table 1]

[0155] As mentioned above, a single picture can contain one or more slices. A slice can have one of the following slice types: I-slice (intra slice), P-slice (predictive slice), and B-slice (bi-predictive slice). The slice type can be indicated based on the slice type information. For blocks in an I-slice, only intra-prediction can be used for prediction, and inter-prediction is not used. Of course, even in this case, original sample values ​​may be coded and signaled without prediction. For blocks in a P-slice, intra-prediction or inter-prediction can be used, and if inter-prediction is used, only uni-prediction can be used. On the other hand, for blocks in a B-slice, intra-prediction or inter-prediction can be used, and if inter-prediction is used, up to the maximum bi-prediction can be used.

[0156] L0 and L1 may contain reference pictures that were encoded / decoded before the current picture. For example, L0 may contain reference pictures that are earlier and / or later than the current picture in the POC order, and L1 may contain reference pictures that are later and / or earlier than the current picture in the POC order. In this case, L0 may be assigned an index of a reference picture that is even lower relative to the reference picture that is earlier than the current picture in the POC order, and L1 may be assigned an index of a reference picture that is even lower relative to the reference picture that is later than the current picture in the POC order. In the case of B slices, biprediction can be applied, and in this case as well, unidirectional biprediction or bidirectional biprediction can be applied. Bidirectional biprediction may be called true biprediction.

[0157] As described above, a residual block (residual sample) can be derived based on a predicted block (predicted sample) derived through prediction in the encoding stage, and residual information can be generated through conversion / quantization of the residual sample. The residual information may include information on the quantized conversion coefficients. The residual information may be included in video / image information, which can be encoded and transmitted to a decoding device in bitstream form. The decoding device can obtain the residual information from the bitstream and derive a residual sample based on the residual information. Specifically, the decoding device can derive quantized conversion coefficients based on the residual information and derive a residual block (residual sample) through an inverse quantization / inverse conversion procedure.

[0158] On the other hand, at least one of the (inverse) transformation and / or (inverse) quantization steps can be omitted.

[0159] The following describes an in-loop filtering procedure performed for a restored picture. Through the in-loop filtering procedure, modified restored samples, blocks, and pictures (or modified filtered samples, blocks, and pictures) can be generated, and the decoder can output the modified restored picture as a decoded picture, or it can be stored in the decoded picture buffer or memory of the encoding / decoding device and used as a reference picture in the interpretation procedure during picture encoding / decoding. The in-loop filtering procedure may include, as previously stated, a deblocking filtering procedure, a sample adaptive offset (SAO) procedure, and / or an adaptive loop filter (ALF) procedure. In this case, one or part of the deblocking filtering procedure, the SAO procedure, the ALF procedure, and the bi-lateral filter procedure may be applied sequentially, or all of them may be applied sequentially. For example, after the deblocking filtering procedure is applied to the restored picture, the SAO procedure may be executed. Alternatively, for example, the ALF procedure can be performed after a deblocking filtering procedure has been applied to the restored picture. This can also be done on the encoding device.

[0160] Deblocking filtering is a filtering technique that removes distortion occurring at the boundaries between blocks in a reconstructed picture. The deblocking filtering procedure can, for example, involve deriving a target boundary in the reconstructed picture, determining a boundary strength (bS) for the target boundary, and performing deblocking filtering on the target boundary based on the bS. The bS can be determined based on the prediction modes of two adjacent blocks, the difference in motion vectors, whether the reference pictures are identical, and whether a non-zero effectiveness coefficient exists.

[0161] SAO is a method for compensating for the offset difference between a restored picture and the original picture on a sample-by-sample basis, and can be applied based on types such as Band Offset and Edge Offset. According to SAO, each SAO type can classify samples into different categories, and an offset value can be added to each sample based on the category. Filtering information for SAO can include information on whether SAO is applicable, SAO type information, SAO offset value information, etc. SAO can also be applied to the restored picture after the deblocking filtering has been applied.

[0162] ALF (Adaptive Loop Filter) is a technique that filters a restored picture on a sample-by-sample basis based on a filter coefficient determined by the filter shape. The encoding device can determine whether ALF is applicable, the ALF shape, and / or the ALF filtering coefficient by comparing the restored picture with the original picture, and can signal this to the decoding device. That is, filtering information for ALF can include information on whether ALF is applicable, ALF filter shape information, ALF filtering coefficient information, etc. ALF can also be applied to the restored picture after the deblocking filtering has been applied.

[0163] Figure 9 shows an example of an ALF filter shape.

[0164] Figure 9(a) shows a 7x7 diamond filter shape, and (b) shows a 5x5 diamond filter shape. In Figure 9, Cn within the filter shape represents the filter coefficient. In Cn, if n is the same, this indicates that the same filter coefficient can be assigned. In this document, the position and / or unit to which the filter coefficient is assigned by the ALF filter shape can be called a filter tab. In this case, one filter coefficient can be assigned to each filter tab, and the arrangement of the filter tabs can be called a filter shape. The filter tab located in the center of the filter shape can be called a center filter tab. Two filter tabs with the same n value located at positions corresponding to each other with respect to the center filter tab can be assigned the same filter coefficient. For example, in the case of a 7x7 diamond filter shape, it contains 25 filter tabs, and the filter coefficients C0 to C11 are assigned in a centrally symmetrical manner, so only 13 filter coefficients are needed to assign filter coefficients to the 25 filter tabs. Furthermore, for example, in the case of a 5x5 diamond filter shape, since it includes 13 filter tabs and the filter coefficients C0 to C5 are assigned in a centrally symmetrical manner, only 7 filter coefficients are needed to assign the filter coefficients to the 13 filter tabs. For example, to reduce the amount of data related to the signaled filter coefficients, of the 13 filter coefficients for a 7x7 diamond filter shape, 12 filter coefficients can be (explicitly) signaled and 1 filter coefficient can be (implicitly) derived. Also, for example, of the 7 filter coefficients for a 5x5 diamond filter shape, 6 filter coefficients can be (explicitly) signaled and 1 filter coefficient can be (implicitly) derived.

[0165] According to one embodiment of this document, the ALF parameters used for the ALF procedure can be signaled via an APS (adaptation parameter set). The ALF parameters can be derived from filter information or ALF data for the ALF.

[0166] As mentioned earlier, ALF is a type of in-loop filtering technique that can be applied in video / image coding. ALF can be performed using Wiener-based adaptive filters to minimize the mean square error (MSE) between the original sample and the decoded sample (or reconstructed sample). The high-level design for the ALF tool can incorporate syntactic elements that can be accessed in the SPS and / or slice header (or tile group header).

[0167] In one example, before filtering each of the 4x4 Luma blocks, geometric transformations such as rotation or diagonal and vertical flipping can be applied to the filter coefficients f(k, l) and corresponding filter clipping values ​​c(k, l), which depend on the slope values ​​calculated for the block. This is equivalent to applying these transformations to the samples within the filter assistance region. This is similar to generating other blocks to which the ALF is applied and aligning these blocks according to their orientation.

[0168] For example, the three transformations—diagonal, vertical flip, and rotation—can be performed based on the following formulas.

[0169]

number

[0170]

number

[0171]

number

[0172] In equations 1 to 3 above, K is the size of the filter. 0 ≤ k, 1 ≤ K-1 are coefficient coordinates. For example, (0, 0) is the upper-left corner coordinate, and / or (K-1, K-1) is the lower-right corner coordinate. The relationship between the transformation and the four slopes in the four directions can be summarized in the table below.

[0173] [Table 2]

[0174] ALF filter parameters can be signaled in the APS and slice header. Up to 25 lumina filter coefficients and clipping value indices can be signaled in a single APS. Up to 8 chroma filter coefficients and clipping value indices can be signaled in a single APS. To reduce bit overhead, filter coefficients of different classifications for lumina components can be merged. The slice header can signal the index of the APS currently used for the slice (the one the slice currently references).

[0175] The clipping value index decoded from the APS can be made to allow the clipping value to be determined using the clipping value luma table and clipping value chroma table. These clipping values ​​are dependent on the internal bit depth. More specifically, the clipping value luma table and clipping value chroma table can be derived based on the following formulas.

[0176]

number

[0177]

number

[0178] In the above formula, B is the internal bit depth, and N is the number of allowed clipping values ​​(a predetermined number). For example, N is 4.

[0179] In the slice header, up to seven APS indices can be signaled to indicate the Luma filter set currently used for the slice. The filtering procedure can be further controlled at the CTB level. For example, a flag indicating whether ALF is applied to the Luma CTB can be signaled. The Luma CTB can select one filter set from 16 fixed filter sets and filter sets from APS. Filter set indices can be signaled for the Luma CTB to indicate which filter set is applied. The 16 fixed filter sets can be predefined and hardcoded for both the encoder and decoder.

[0180] For chroma components, the APS index can be signaled in the slice header to indicate the chroma filter set currently used for the slice. At the CTB level, if the APS has more than one chroma filter set, the filter index can be signaled for each chroma CTB.

[0181] The filter coefficients can be quantized relative to (norm) 128. To limit the multiplicative complexity, bitstream conformance can be applied, so that coefficient values ​​of non-central positions are in the range of 0 to 28, and / or coefficient values ​​of other positions are in the range of -27 to 27-1. The central position coefficient can be pre-determined (considered) as 128 without signaling in the bitstream.

[0182] If ALF is currently available for a block, each sample R(i, j) can be filtered, and the filtered result R'(i, j) can be expressed as follows:

[0183]

number

[0184] In the above formula, f(k, l) is the decoded filter coefficient, K(x, y) is the clipping function, and c(k, l) is the decoded clipping parameter. For example, the variables k and / or l can vary from -L / 2 to L / 2, where L can represent the filter length. The clipping function K(x, y) = min(y, max(-y, x)) corresponds to the function Clip3(-y, y, x).

[0185] In one example, to reduce the ALF line buffer requirements, modified block classification and filtering can be applied to samples adjacent to a horizontal CTU boundary. For this purpose, a virtual boundary can be defined.

[0186] Figure 10 is a diagram illustrating the virtual boundary applied to the filtering procedure according to one embodiment of this document. Figure 11 shows an example of an ALF procedure that utilizes the virtual boundary according to one embodiment of this document. Figure 11 is described together with Figure 10.

[0187] Referring to Figure 10, the virtual boundary is a line defined by shifting the horizontal CTU boundary by N samples. In one example, N is 4 for the luma component and / or N is 2 for the chroma component.

[0188] In Figure 10, the modified block classification can be applied to the Luma component. For the calculation of the 1D Laplacian slope of a 4x4 block on the virtual boundary, only samples on the virtual boundary can be used. Similarly, for the calculation of the 1D Laplacian slope of a 4x4 block below the virtual boundary, only samples below the virtual boundary can be used. The quantization of the activity value A can be scaled by considering the reduced number of samples used in the 1D Laplacian slope calculation.

[0189] For the filtering procedure, symmetric padding operations at the virtual boundary can be used for the luma and chroma components. Referring to Figure 10, if a filtered sample is located below the virtual boundary, adjacent samples located on the virtual boundary can be padded. Conversely, the corresponding samples on the other side can also be padded symmetrically.

[0190] The procedure illustrated in Figure 11 can also be used for slice, brick, and / or tile boundaries when filtering is not available across the boundary. For ALF block classification, only samples contained within the same slice, brick, and / or tile can be used, and the activity value can be scaled accordingly. For ALF filtering, symmetrical padding can be applied to the horizontal and / or vertical boundaries, respectively.

[0191] Figure 12 is a diagram illustrating a cross-component adaptive loop filtering (CCALF (CC-ALF)) procedure according to one embodiment of this document. The CCALF procedure can also be called a cross-component filtering procedure.

[0192] From one perspective, the ALF procedure can include the general ALF procedure and the CCALF procedure. That is, the CCALF procedure can refer to a subset of the ALF procedure. From another perspective, the filtering procedure can include the deblocking procedure, the SAO procedure, the ALF procedure, and / or the CCALF procedure.

[0193] CC-ALF can refine each chroma component using chroma sample values. CC-ALF is controlled by (image) information in the bitstream, which may include (a) information about filter coefficients for each chroma component and (b) information about a mask that controls the application of the filter to a block of sample. The filter coefficients can be signaled at the APS level, and the block size and mask can be signaled at the slice level.

[0194] Referring to Figure 12, CC-ALF can operate by applying a linear diamond-shaped filter (Figure 12(b)) to the chroma channel for each chroma component. The filter coefficients are sent to the APS, scaled by a factor of 210, and rounded for fixed-point representation. The application of the filter can be controlled by a variable block size and signaled by a context coding flag received for each sample block. The block size, along with the CC-ALF enabled flag, can be received at the slice level for each chroma component. The block size (for a chroma sample) is 16×16, 32×32, 64×64, or 128×128.

[0195] In the following examples, a method is proposed for re-filtering or modifying a reconstructed chromatic sample that has been filtered by ALF, based on the reconstructed chromatic sample.

[0196] One embodiment described herein relates to the filter on / off transmission and filter coefficient transmission within CC-ALF. As previously stated, the information (syntax elements) in the syntax table disclosed herein can be included in image / video information and can be configured / encoded by an encoding device and transmitted to a decoding device in bitstream form. The decoding device can parse / decode the information (syntax elements) in the syntax table. Based on the decoded information, the decoding device can execute a picture / image / video decoding procedure (specifically, for example, the CC-ALF procedure described above). The same applies to other embodiments below.

[0197] The table below shows some of the syntax of slice header information according to the examples in this document.

[0198] [Table 3]

[0199] The following table shows exemplary semantics for the syntax elements included in the table above.

[0200] [Table 4]

[0201] Referring to the two tables above, if sps_cross_component_alf_enabled_flag is 1 in the slice header, parsing of slice_cross_component_alf_cb_enabled_flag can be performed to determine whether Cb CC-ALF should be applied to the slice. If slice_cross_component_alf_cb_enabled_flag is 1, CC-ALF can be applied to the Cb slice, and if slice_cross_component_alf_cb_reuse_temporal_layer_filter is 1, the filter of the existing same temporal layer can be reused. If slice_cross_component_alf_cb_enabled_flag is 0, CC-ALF can be applied using the filter in the corresponding APS (adaptation parameter set) id via slice_cross_component_alf_cb_aps_idparsing. `slice_cross_component_alf_cb_log2_control_size_minus4` can mean a CC-ALF applied block unit in a Cb slice.

[0202] For example, if the value of slice_cross_component_alf_cb_log2_control_size_minus4 is 0, the application of CC-ALF is determined in 16x16 units. If the value of slice_cross_component_alf_cb_log2_control_size_minus4 is 1, the application of CC-ALF is determined in 32x32 units. If the value of slice_cross_component_alf_cb_log2_control_size_minus4 is 2, the application of CC-ALF is determined in 64x64 units. If the value of slice_cross_component_alf_cb_log2_control_size_minus4 is 3, the application of CC-ALF is determined in 128x128 units. In addition, the same syntax structure as above is used for Cr CC-ALF.

[0203] The table below shows exemplary syntax for ALF data.

[0204] [Table 5]

[0205] The following table shows exemplary semantics for the syntax elements included in the table above.

[0206] [Table 6-1]

[0207] [Table 6-2]

[0208] Referring to the two tables above, CC-ALF syntax elements are configured to be transmitted and applied independently of the existing (general) ALF syntax structure. That is, CC-ALF can be applied even when the ALF tool on SPS is turned off. Because CC-ALF must be able to operate independently of the existing ALF structure, a new hardware pipeline design is required. This will result in increased hardware implementation costs and hardware delays.

[0209] Furthermore, ALF determines whether to apply both luminous and chroma images on a CTU basis, and transmits the result of this determination to the decoder via signaling. However, because the application of variable CC-ALF from 16x16 to 128x128 units is determined and applied, a conflict can occur between the existing ALF structure and the CC-ALF. This creates problems during hardware implementation and also leads to an increase in the line buffer required for various variable CC-ALF applications.

[0210] This invention aims to solve the aforementioned hardware implementation problems of CC-ALF by integrally applying the CC-ALF syntax structure to the ALF syntax structure.

[0211] According to one embodiment of this document, a sequence parameter set (SPS) may include a CC-ALF enable flag (sps_ccalf_enable_flag) to determine whether CC-ALF is used (applied). The CC-ALF enable flag can be transmitted independently of an ALF enable flag (sps_alf_enabled_flag) for determining whether ALF is used (applied).

[0212] The following table shows some of the exemplary syntax for SPS according to this embodiment.

[0213] [Table 7]

[0214] Referring to the table above, CC-ALF can only be applied when ALF is always active. That is, the CC-ALF enabled flag (sps_ccalf_enabled_flag) can only be parsed when the ALF enabled flag (sps_alf_enabled_flag) is 1. According to the table above, CC-ALF and ALF can be combined. The CC-ALF enabled flag can indicate (or be associated with) whether CC-ALF is enabled.

[0215] The table below shows some exemplary syntax for slice headers.

[0216] [Table 8]

[0217] Referring to the table above, parsing of sps_ccalf_enabled_flag can only be performed if sps_alf_enabled_flag is 1. The syntax elements included in the table above can be explained based on Table 4. In one example, video information encoded by an encoding device or acquired (received) by a decoding device may include slice header information (slice_header()). Based on the determination that the value of the CCALF enabled flag (sps_ccalf_flag) is 1, the slice header information may include a first flag (slice_cross_component_alf_cb_enabeld_flag) related to whether CC-ALF is enabled for the Cb color component of the filtered reconstructed chroma sample, and a second flag (slice_cross_component_alf_cr_enabeld_flag) related to whether CC-ALF is enabled for the Cr color component of the filtered reconstructed chroma sample.

[0218] In one example, based on the determination that the value of the first flag (slice_cross_component_alf_cb_enabeld_flag) is 1, the slice header information may include the ID information of the first APS (slice_cross_component_alf_cb_aps_id) for deriving the cross-component filter coefficients for the Cb color component. Based on the determination that the value of the second flag (slice_cross_component_alf_cr_enabeld_flag) is 1, the slice header information may include the ID information of the second APS (slice_cross_component_alf_cr_aps_id) for deriving the cross-component filter coefficients for the Cr color component.

[0219] The following table shows some of the SPS syntax from other examples in this embodiment.

[0220] [Table 9]

[0221] The following table illustrates some of the slice header syntax.

[0222] [Table 10]

[0223] Referring to Table 9 above, if ChromaArrayType is not 0 and the ALF enabled flag (sps_alf_enabled_flag) is 1, the SPS may include the CCALF enabled flag (sps_ccalf_enabled_flag). For example, if ChromaArrayType is not 0, the CCALF enabled flag may be sent via the SPS based on whether the chroma format is not monochrome.

[0224] Referring to Table 9 above, if ChromaArrayType is not 0, information about CCALF (slice_cross_component_alf_cb_enabled_flag, slice_cross_component_alf_cb_aps_id, slice_cross_component_alf_cr_enabled_flag, slice_cross_component_alf_cr_aps_id) can be included in the slice header information.

[0225] In one example, video information encoded by an encoding device or acquired by a decoding device may include the SPS. The SPS may include a first ALF enabled flag (sps_alf_enabled_flag) related to whether ALF is available. For example, based on the determination that the value of the first ALF enabled flag is 1, the SPS may include a CCALF enabled flag related to whether cross-component filtering is available. In another example, instead of using sps_ccalf_enabled_flag, CCALF may always be applied if sps_alf_enabled_flag is 1 (sps_ccalf_enabled_flag==1).

[0226] The following table shows some of the slice header syntax from other examples in this embodiment.

[0227] [Table 11]

[0228] Referring to the table above, CCALF enable flag (sps_ccalf_enabled_flag) parsing can only be performed if the ALF enable flag (sps_alf_enabled_flag) is 1.

[0229] The following table shows exemplary semantics for the syntax elements included in the table above.

[0230] [Table 12]

[0231] The value `slice_ccalf_chroma_idc` in the table above can also be explained by the semantics in the table below.

[0232] [Table 13]

[0233] The following table shows some of the slice header syntax from other examples in this embodiment.

[0234] [Table 14]

[0235] The syntax elements included in the aforementioned table can be explained by Table 12 or Table 13. Furthermore, if the chroma format is not monochrome, CCALF-related information may be included in the slice header.

[0236] The following table shows some of the slice header syntax from other examples in this embodiment.

[0237] [Table 15]

[0238] The following table shows exemplary semantics for the syntax elements included in the table above.

[0239] [Table 16]

[0240] The following table shows some of the slice header syntax from other examples of this embodiment. The syntax elements included in the following table can be explained by Table 12 or Table 13.

[0241] [Table 17]

[0242] Referring to the table above, the application of slice-unit ALF and CC-ALF can be determined at once via slice_alf_enabled_flag. After parsing slice_alf_chroma_idc, if the first ALF enabled flag (sps_alf_enabled_flag) is 1, slice_ccalf_chroma_idc can be parsed.

[0243] Referring to the table above, it can be determined whether sps_ccalf_enabled_flag is 1 in the slice header information only if slice_alf_enabled_flag is 1. The slice header information may include a second ALF enabled flag (slice_alf_enabled_flag) associated with whether ALF is enabled. Based on the determination that the value of the second ALF enabled flag (slice_alf_enabled_flag) is 1, CCALF is enabled for the slice.

[0244] The following table illustrates some of the APS syntax. The syntax element `adaptation_parameter_set_id` can represent the identifier information (ID information) of the APS.

[0245] [Table 18]

[0246] The table below shows exemplary syntax for ALF data.

[0247] [Table 19]

[0248] Referring to the two tables, the APS can include ALF data (alf_data()). The APS including ALF data can be called an ALF APS (ALF type APS). That is, the type of the APS including ALF data is the ALF type. The type of the APS can be determined by information regarding the APS type or a syntax element (aps_params_type). The ALF data can include a Cb filter signal flag (alf_cross_component_cb_filter_signal_flag or alf_cc_cb_filter_signal_flag) related to whether a cross-component filter for the Cb color component is signaled. The ALF data can include a Cr filter signal flag (alf_cross_component_cr_filter_signal_flag or alf_cc_cr_filter_signal_flag) related to whether a cross-component filter for the Cr color component is signaled.

[0249] In one example, based on the Cr filter signal flag, the ALF data can include information regarding the absolute value of the cross-component filter coefficient for the Cr color component (alf_cross_component_cr_coeff_abs) and information regarding the sign of the cross-component filter coefficient for the Cr color component (alf_cross_component_cr_coeff_sign). Based on the information regarding the absolute value of the cross-component filter coefficient for the Cr color component and the information regarding the sign of the cross-component filter coefficient for the Cr color component, the cross-component filter coefficient for the Cr color component can be derived.

[0250] In one example, the ALF data can include information (alf_cross_component_cb_coeff_abs) regarding the absolute value of the cross-component filter coefficient for the Cb color component and information (alf_cross_component_cb_coeff_sign) regarding the sign of the cross-component filter coefficient for the Cb color component. Based on the information regarding the absolute value of the cross-component filter coefficient for the Cb color component and the information regarding the sign of the cross-component filter coefficient for the Cb color component, the cross-component filter coefficient for the Cb color component can be derived.

[0251] The following table shows the syntax for ALF data according to another example.

[0252]

Table 20

[0253] Referring to the table above, after transmitting alf_cross_component_filter_signal_flag first, if alf_cross_component_filter_signal_flag is 1, the Cb / Cr filter signal flag can be transmitted. That is, alf_cross_component_filter_signal_flag determines whether to integrate Cb / Cr and transmit the CC-ALF filter coefficient.

[0254] The following table shows the syntax for ALF data according to another example.

[0255]

Table 21

[0256] The following table shows exemplary semantics regarding the syntax elements included in the table above.

[0257]

Table 22-1

[0258]

Table 22-2

[0259] The following table shows the syntax for ALF data according to other examples.

[0260]

Table 23

[0261] The following table shows exemplary semantics regarding the syntax elements included in the said table.

[0262]

Table 24-1

[0263]

Table 24-2

[0264] In the said two tables, the order of exp-Golomb binarization for parsing the alf_cross_component_cb_coeff_abs[j] and alf_cross_component_cr_coeff_abs[j] syntax can be defined as one of the values from 0 to 9.

[0265] Referring to the two tables above, the ALF data may include a Cb filter signal flag (alf_cross_component_cb_filter_signal_flag or alf_cc_cb_filter_signal_flag) related to whether a cross-component filter for the Cb color component was signaled. Based on the Cb filter signal flag (alf_cross_component_cb_filter_signal_flag), the ALF data may include information related to the number of cross-component filters for the Cb color component (ccalf_cb_num_alt_filters_minus1). Based on the information related to the number of cross-component filters for the Cb color component, the ALF data may include information regarding the absolute value of the cross-component filter coefficients for the Cb color component (alf_cross_component_cb_coeff_abs) and information regarding the sign of the cross-component filter coefficients for the Cb color component (alf_cross_component_cr_coeff_sign). Based on information regarding the absolute value of the cross-component filter coefficient for the Cb color component and information regarding the sign of the cross-component filter coefficient for the Cb color component, the cross-component filter coefficient for the Cb color component can be derived.

[0266] In one example, the ALF data may include a Cr filter signal flag (alf_cross_component_cr_filter_signal_flag or alf_cc_cr_filter_signal_flag) related to whether a cross-component filter for the Cr color component was signaled. Based on the Cr filter signal flag (alf_cross_component_cr_filter_signal_flag), the ALF data may include information related to the number of cross-component filters for the Cr color component (ccalf_cr_num_alt_filters_minus1). Based on the information related to the number of cross-component filters for the Cr color component, the ALF data may include information regarding the absolute value of the cross-component filter coefficients for the Cr color component (alf_cross_component_cr_coeff_abs) and information regarding the sign of the cross-component filter coefficients for the Cr color component (alf_cross_component_cr_coeff_sign). Based on information regarding the absolute value of the cross-component filter coefficient for the Cr color component and information regarding the sign of the cross-component filter coefficient for the Cr color component, the cross-component filter coefficient for the Cr color component can be derived.

[0267] The following table shows the syntax for a coding tree unit according to one embodiment of this document.

[0268] [Table 25]

[0269] The following table shows exemplary semantics for the syntax elements included in the table above.

[0270] [Table 26]

[0271] The following table shows coding tree unit syntax for other examples in this embodiment.

[0272] [Table 27]

[0273] Referring to the table above, CCALF can be applied on a CTU basis. In one example, the video information may include information about the coding tree unit (coding_tree_unit()). The coding tree unit information may include information about whether a cross-component filter is applied to the current block of the Cb color component (ccalf_ctb_flag[0]), and / or information about whether a cross-component filter is applied to the current block of the Cr color component (ccalf_ctb_flag[1]). The coding tree unit information may also include information about the filter set index of the cross-component filter applied to the current block of the Cb color component (ccalf_ctb_filter_alt_idx[0]), and / or information about the filter set index of the cross-component filter applied to the current block of the Cr color component (ccalf_ctb_filter_alt_idx[1]). The syntax can be transmitted adaptively by the slice_ccalf_enabled_flag and slice_ccalf_chroma_idc syntax.

[0274] The following table shows coding tree unit syntax for other examples in this embodiment.

[0275] [Table 28]

[0276] The following table shows exemplary semantics for the syntax elements included in the table above.

[0277] [Table 29]

[0278] The following table shows the coding tree unit syntax for other examples of this embodiment. The syntax elements included in the following table can be explained by Table 29.

[0279] [Table 30]

[0280] In one example, the video information may include information about a coding tree unit (coding_tree_unit()). The coding tree unit information may include information about whether a cross-component filter is applied to the current block of the Cb color component (ccalf_ctb_flag[0]), and / or information about whether a cross-component filter is applied to the current block of the Cr color component (ccalf_ctb_flag[1]). The coding tree unit information may also include information about the filter set index of the cross-component filter applied to the current block of the Cb color component (ccalf_ctb_filter_alt_idx[0]), and / or information about the filter set index of the cross-component filter applied to the current block of the Cr color component (ccalf_ctb_filter_alt_idx[1]).

[0281] Figures 13 and 14 schematically illustrate an example of a video / image encoding method and related components according to the embodiments described in this document.

[0282] The method disclosed in Figure 13 can be performed by the encoding apparatus disclosed in Figure 2 or Figure 14. Specifically, for example, steps S1300 to S1330 in Figure 13 can be performed by the residual processing unit 230 of the encoding apparatus in Figure 14, step S1340 in Figure 13 can be performed by the adder 250 of the encoding apparatus in Figure 14, step S1350 in Figure 13 can be performed by the filtering unit 260 of the encoding apparatus in Figure 14, and step S1360 in Figure 13 can be performed by the entropy encoding unit 240 of the encoding apparatus in Figure 14. Although not shown in Figure 13, in Figure 13, the prediction unit 220 of the encoding apparatus can derive prediction samples or prediction-related information, and the entropy encoding unit 240 of the encoding apparatus can generate a bitstream from the residual information or prediction-related information. The method disclosed in Figure 13 may include embodiments detailed in this document.

[0283] Referring to Figure 13, the encoding device can derive a residual sample (S1300). The encoding device can derive a residual sample for the current block, which can be derived based on the original sample and predicted sample of the current block. Specifically, the encoding device can derive a predicted sample of the current block based on a prediction mode. In this case, various prediction methods disclosed in this document, such as interpretation or intrapretation, can be applied. A residual sample can be derived based on the predicted sample and the original sample.

[0284] The encoding device can derive conversion coefficients (S1310). The encoding device can derive conversion coefficients based on the conversion procedure for the residual sample. For example, the conversion procedure may include at least one of DCT, DST, GBT, or CNT.

[0285] The encoding device can derive quantized transformation coefficients (S1320). The encoding device can derive quantized transformation coefficients based on a quantization procedure for the transformation coefficients. The quantized transformation coefficients may have a one-dimensional vector form based on the coefficient scan order.

[0286] The encoding device can generate residual information (S1330). The encoding device can generate residual information indicating the quantized conversion coefficients. Residual information can be generated through various encoding methods such as exponential golome, CAVLC, CABAC, etc.

[0287] The encoding device can generate a reconstructed sample (S1340). The encoding device can generate a reconstructed sample based on the residual information. The reconstructed sample can be generated by adding the residual sample based on the residual information to the predicted sample. Specifically, the encoding device can perform a prediction (intra or inter prediction) for the current block and generate a reconstructed sample based on the original sample and the predicted sample generated from the prediction.

[0288] A restored sample may include a restored luminal sample and a restored chroma sample. Specifically, a residual sample may include a residual luminal sample and a residual chroma sample. A residual luminal sample can be generated based on an original luminal sample and a predicted luminal sample. A residual chroma sample can be generated based on an original chroma sample and a predicted chroma sample. An encoding device can derive conversion coefficients (luminal conversion coefficients) for the residual luminal sample and / or conversion coefficients (chroma conversion coefficients) for the residual chroma sample. Quantized conversion coefficients may include quantized luminal conversion coefficients and / or quantized chroma conversion coefficients.

[0289] The encoding device can generate ALF-related information and / or CCALF (CC-ALF)-related information for the reconstructed sample (S1350). The encoding device can generate ALF-related information for the reconstructed sample. The encoding device derives ALF-related parameters that can be applied for filtering of the reconstructed sample and generates ALF-related information. For example, the ALF-related information may include the ALF-related information detailed in this document.

[0290] The encoding device can encode video / image information (S1360). The image information may include residual information, ALF-related information, and / or CCALF-related information. The encoded video / image information can be output in bitstream format. The bitstream can be transmitted to the decoding device via a network or storage medium.

[0291] For example, CCALF-related information may include a CCALF enabled flag, a flag associated with whether CCALF is enabled for a Cb (or Cr) color component, a Cb (or Cr) filter signal flag associated with whether a cross-component filter for a Cb (or Cr) color component has been signaled, information associated with the number of cross-component filters for a Cb (or Cr) color component, information regarding the values ​​of the cross-component filter coefficients for a Cb (or Cr) color component, information regarding the absolute values ​​of the cross-component filter coefficients for a Cb (or Cr) color component, information regarding the signs of the cross-component filter coefficients for a Cb (or Cr) color component, and / or information regarding the coding tree unit (coding tree unit syntax) regarding whether a cross-component filter is applied to the current block of a Cb (or Cr) color component.

[0292] The aforementioned video information may include a variety of information as described in the embodiments of this document. For example, the video information may include information disclosed in at least one of the Tables 1 to 30 mentioned above.

[0293] In one embodiment, the video information may include header information and an adaptation parameter set (APS). The header information may include information related to an identifier of the APS, which includes ALF data. For example, the cross-component filter coefficients can be derived based on the ALF data.

[0294] In one embodiment, the video information may include a sequence parameter set (SPS). The SPS may include a CCALF enabled flag related to whether the cross-component filtering is available.

[0295] In one embodiment, the SPS may include an ALF enabled flag (sps_alf_enabled_flag) related to whether ALF is available. Based on the determination that the value of the ALF enabled flag is 1, the SPS may include a CCALF enabled flag related to whether cross-component filtering is available.

[0296] In one embodiment, the video information may include slice header information. The slice header information may include an ALF enabled flag (slice_alf_enabled_flag) associated with whether ALF is available. Based on the determination that the value of the ALF enabled flag is 1, it can be determined whether the value of the CCALF enabled flag is 1. In one example, based on the determination that the value of the ALF enabled flag is 1, the CCALF is available for the slice.

[0297] In one embodiment, the header information (slice header information) may include a first flag related to whether CCALF is available for the Cb color component of the filtered reconstructed chroma sample, and a second flag related to whether CCALF is available for the Cr color component of the filtered reconstructed chroma sample. In another example, based on the determination that the value of the ALF enabled flag (slice_alf_enabled_flag) is 1, the header information (slice header information) may include a first flag related to whether CCALF is available for the Cb color component of the filtered reconstructed chroma sample, and a second flag related to whether CCALF is available for the Cr color component of the filtered reconstructed chroma sample.

[0298] In one embodiment, the image information may include adaptation parameter sets (APSs). In one example, the slice header information may include ID information of a first APS (information associated with the identifier of a second APS) for deriving cross-component filter coefficients for the Cb color component of the filtered reconstructed chroma sample. The slice header information may include ID information of a second APS (information associated with the identifier of a second APS) for deriving cross-component filter coefficients for the Cr color component of the filtered reconstructed chroma sample. In another example, based on the determination that the value of the first flag is 1, the slice header information may include ID information of a first APS (information associated with the identifier of a second APS) for deriving cross-component filter coefficients for the Cb color component. Based on the determination that the value of the second flag is 1, the slice header information may include ID information of a second APS (information associated with the identifier of a second APS) for deriving cross-component filter coefficients for the Cr color component.

[0299] In one embodiment, the first ALF data included in the first APS may include a Cb filter signal flag related to whether the cross-component filter for the Cb color component was signaled. Based on the Cb filter signal flag, the first ALF data may include information related to the number of cross-component filters for the Cb color component. Based on the information related to the number of cross-component filters for the Cb color component, the first ALF data may include information regarding the absolute value of the cross-component filter coefficients for the Cb color component and information regarding the sign of the cross-component filter coefficients for the Cb color component. Based on the information regarding the absolute value of the cross-component filter coefficients for the Cb color component and information regarding the sign of the cross-component filter coefficients for the Cb color component, the cross-component filter coefficients for the Cb color component can be derived.

[0300] In one embodiment, the information related to the number of cross-component filters for the Cb color component is the 0th exponential Golomb (0 th EG) Can be coded.

[0301] In one embodiment, the second ALF data included in the second APS may include a Cr filter signal flag related to whether the cross-component filter for the Cr color component was signaled. Based on the Cr filter signal flag, the second ALF data may include information related to the number of cross-component filters for the Cr color component. Based on the information related to the number of cross-component filters for the Cr color component, the second ALF data may include information regarding the absolute value of the cross-component filter coefficients for the Cr color component and information regarding the sign of the cross-component filter coefficients for the Cr color component. Based on the information regarding the absolute value of the cross-component filter coefficients for the Cr color component and information regarding the sign of the cross-component filter coefficients for the Cr color component, the cross-component filter coefficients for the Cr color component can be derived.

[0302] In one embodiment, the information related to the number of cross-component filters for the Cr color component is expressed as a zero-order exponential Golomb (0 th EG) Can be coded.

[0303] In one embodiment, the video information may include information about a coding tree unit. The information about the coding tree unit may include information about whether a cross-component filter is applied to the current block of the Cb color component, and / or information about whether a cross-component filter is applied to the current block of the Cr color component.

[0304] In one embodiment, the information relating to the coding tree unit may include information relating to the filter set index of the cross-component filter applied to the current block of the Cb color component, and / or information relating to the filter set index of the cross-component filter applied to the current block of the Cr color component.

[0305] Figures 15 and 16 schematically illustrate an example of a video / image decoding method and related components according to the embodiments described in this document.

[0306] The method disclosed in Figure 15 can be performed by the decoding apparatus disclosed in Figure 3 or Figure 16. Specifically, for example, S1500 in Figure 15 can be performed by the entropy decoding unit 310 of the decoding apparatus, S1510 can be performed by the addition unit 340 of the decoding apparatus, and S1520 to S1550 can be performed by the filtering unit 350 of the decoding apparatus. The method disclosed in Figure 15 may include embodiments detailed in this document.

[0307] Referring to Figure 15, the decoding device can receive / acquire video / image information (S1500). The video / image information may include residual information. The decoding device can receive / acquire the video / image information via a bitstream. In one example, the video / image information may further include CCAL-related information. For example, CCALF-related information may include a CCALF-enabled flag, a flag associated with whether CCALF is enabled for a Cb (or Cr) color component, a Cb (or Cr) filter signal flag associated with whether a cross-component filter for a Cb (or Cr) color component has been signaled, information associated with the number of cross-component filters for a Cb (or Cr) color component, information regarding the absolute value of the cross-component filter coefficients for a Cb (or Cr) color component, information regarding the sign of the cross-component filter coefficients for a Cb (or Cr) color component, and / or information regarding whether a cross-component filter is applied to the current block of a Cb (or Cr) color component in the coding tree unit information (coding tree unit syntax).

[0308] The aforementioned video information may include a variety of information as described in the embodiments of this document. For example, the video information may include information disclosed in at least one of the Tables 1 to 30 mentioned above.

[0309] The decoding device can derive quantized transformation coefficients. The decoding device can derive quantized transformation coefficients based on the residual information. The quantized transformation coefficients may have a one-dimensional vector form based on the coefficient scan order. The quantized transformation coefficients may include quantized luma transformation coefficients and / or quantized chroma transformation coefficients.

[0310] The decoding device can derive conversion coefficients. The decoding device can derive conversion coefficients based on an inverse quantization procedure for the quantized conversion coefficients. The decoding device can derive Luma conversion coefficients via inverse quantization based on quantized Luma conversion coefficients. The decoding device can derive Chroma conversion coefficients via inverse quantization based on quantized Chroma conversion coefficients.

[0311] The decoding device can generate / derive a resistive sample. The decoding device can derive a resistive sample based on an inverse transformation procedure for the transformation coefficients. The decoding device can derive a resistive luma sample via an inverse transformation procedure based on luma transformation coefficients. The decoding device can derive a resistive chroma sample via an inverse transformation procedure based on chroma transformation coefficients.

[0312] The decoding device can generate / derive a restored luminal sample and / or a restored chroma sample (S1510). The decoding device can generate a restored luminal sample and / or a restored chroma sample based on the residual information. The decoding device can generate a restored sample based on the residual information. The restored sample may include a restored luminal sample and / or a restored chroma sample. The luminal component of the restored sample may correspond to the restored luminal sample, and the chroma component of the restored sample may correspond to the restored chroma sample. The decoding device can generate a predicted luminal sample and / or a predicted chroma sample via a prediction procedure. The decoding device can generate a restored luminal sample based on the predicted luminal sample and the residual luminal sample. The decoding device can generate a restored chroma sample based on the predicted chroma sample and the residual chroma sample.

[0313] The decoding device can derive ALF filter coefficients for the ALF procedure of the reconstructed chroma sample (S1520). In addition, the decoding device can derive ALF filter coefficients for the ALF procedure of the reconstructed chroma sample. The ALF filter coefficients can be derived based on the ALF parameters included in the ALF data in the APS.

[0314] The decoding device can generate a filtered reconstructed chromatic sample (S1530). The decoding device can generate a reconstructed sample filtered based on the reconstructed chromatic sample and the ALF filter coefficient.

[0315] The decoding device can derive the cross-component filter coefficients for the cross-component filtering (S1540). The cross-component filter coefficients can be derived based on the CCALF-related information in the ALF data included in the aforementioned APS, and the identifier (ID) information of the relevant APS can be included in the slice header (and signaled through it).

[0316] The decoding device can generate modified filtered reconstructed chroma samples (S1550). The decoding device can generate modified filtered reconstructed chroma samples based on the reconstructed chroma samples, the filtered reconstructed chroma samples, and the cross-component filter coefficients. In one example, the decoding device can derive the difference between two samples from the reconstructed chroma samples and multiply the difference by one of the filter coefficients from the cross-component filter coefficients. Based on the result of the multiplication and the filtered reconstructed chroma samples, the decoding device can generate the modified filtered reconstructed chroma samples. For example, the decoding device can generate the modified filtered reconstructed chroma samples based on the sum between the product and one of the filtered reconstructed chroma samples.

[0317] In one embodiment, the video information may include header information and an adaptation parameter set (APS). The header information may include information related to an identifier of the APS, which includes ALF data. For example, the cross-component filter coefficients can be derived based on the ALF data.

[0318] In one embodiment, the video information may include a sequence parameter set (SPS). The SPS may include a CCALF enabled flag related to whether the cross-component filtering is available.

[0319] In one embodiment, the SPS may include an ALF enabled flag (sps_alf_enabled_flag) related to whether ALF is available. Based on the determination that the value of the ALF enabled flag is 1, the SPS may include a CCALF enabled flag related to whether cross-component filtering is available.

[0320] In one embodiment, the video information may include slice header information. The slice header information may include an ALF enabled flag (slice_alf_enabled_flag) associated with whether ALF is available. Based on the determination that the value of the ALF enabled flag is 1, it can be determined whether the value of the CCALF enabled flag is 1. In one example, based on the determination that the value of the ALF enabled flag is 1, the CCALF is available for the slice.

[0321] In one embodiment, the header information (slice header information) may include a first flag related to whether CCALF is available for the Cb color component of the filtered reconstructed chroma sample, and a second flag related to whether CCALF is available for the Cr color component of the filtered reconstructed chroma sample. In another example, based on the determination that the value of the ALF enabled flag (slice_alf_enabled_flag) is 1, the header information (slice header information) may include a first flag related to whether CCALF is available for the Cb color component of the filtered reconstructed chroma sample, and a second flag related to whether CCALF is available for the Cr color component of the filtered reconstructed chroma sample.

[0322] In one embodiment, the image information may include adaptation parameter sets (APSs). In one example, the slice header information may include ID information of a first APS (information associated with the identifier of a second APS) for deriving cross-component filter coefficients for the Cb color component of the filtered reconstructed chroma sample. The slice header information may include ID information of a second APS (information associated with the identifier of a second APS) for deriving cross-component filter coefficients for the Cr color component of the filtered reconstructed chroma sample. In another example, based on the determination that the value of the first flag is 1, the slice header information may include ID information of a first APS (information associated with the identifier of a second APS) for deriving cross-component filter coefficients for the Cb color component. Based on the determination that the value of the second flag is 1, the slice header information may include ID information of a second APS (information associated with the identifier of a second APS) for deriving cross-component filter coefficients for the Cr color component.

[0323] In one embodiment, the first ALF data included in the first APS may include a Cb filter signal flag related to whether the cross-component filter for the Cb color component was signaled. Based on the Cb filter signal flag, the first ALF data may include information related to the number of cross-component filters for the Cb color component. Based on the information related to the number of cross-component filters for the Cb color component, the first ALF data may include information regarding the absolute value of the cross-component filter coefficients for the Cb color component and information regarding the sign of the cross-component filter coefficients for the Cb color component. Based on the information regarding the absolute value of the cross-component filter coefficients for the Cb color component and information regarding the sign of the cross-component filter coefficients for the Cb color component, the cross-component filter coefficients for the Cb color component can be derived.

[0324] In one embodiment, the information related to the number of cross-component filters for the Cb color component is the 0th exponential Golomb (0 th EG) Can be coded.

[0325] In one embodiment, the second ALF data included in the second APS may include a Cr filter signal flag related to whether the cross-component filter for the Cr color component was signaled. Based on the Cr filter signal flag, the second ALF data may include information related to the number of cross-component filters for the Cr color component. Based on the information related to the number of cross-component filters for the Cr color component, the second ALF data may include information regarding the absolute value of the cross-component filter coefficients for the Cr color component and information regarding the sign of the cross-component filter coefficients for the Cr color component. Based on the information regarding the absolute value of the cross-component filter coefficients for the Cr color component and information regarding the sign of the cross-component filter coefficients for the Cr color component, the cross-component filter coefficients for the Cr color component can be derived.

[0326] In one embodiment, the information related to the number of cross-component filters for the Cr color component is expressed as a zero-order exponential Golomb (0 th EG) Can be coded.

[0327] In one embodiment, the video information may include information about a coding tree unit. The information about the coding tree unit may include information about whether a cross-component filter is applied to the current block of the Cb color component, and / or information about whether a cross-component filter is applied to the current block of the Cr color component.

[0328] In one embodiment, the information relating to the coding tree unit may include information relating to the filter set index of the cross-component filter applied to the current block of the Cb color component, and / or information relating to the filter set index of the cross-component filter applied to the current block of the Cr color component.

[0329] The decoding device can receive information about the residual for the current block if a residual sample for the current block exists. The information about the residual may include transformation coefficients for the residual sample. Based on the residual information, the decoding device can derive a residual sample (or a residual sample array) for the current block. Specifically, the decoding device can derive quantized transformation coefficients based on the residual information. The quantized transformation coefficients may have a one-dimensional vector form based on the coefficient scan order. The decoding device can derive transformation coefficients based on an inverse quantization procedure for the quantized transformation coefficients. Based on the transformation coefficients, the decoding device can derive a residual sample.

[0330] The decoding device can generate a reconstructed sample based on an (intra) predicted sample and a residual sample, and can derive a reconstructed block or reconstructed picture based on the reconstructed sample. Specifically, the decoding device can generate a reconstructed sample based on the sum of an (intra) predicted sample and a residual sample. Thereafter, as described above, the decoding device can apply deblocking filtering and / or in-loop filtering procedures such as the SAO procedure to the reconstructed picture as needed to improve subjective / objective image quality.

[0331] For example, a decoding device can decode a bitstream or encoded information to obtain video information that includes all or part of the aforementioned information (or syntax elements). Furthermore, the bitstream or encoded information can be stored on a computer-readable storage medium, which can trigger the aforementioned decoding method.

[0332] In the embodiments described above, the method is explained based on a flowchart as a series of steps or blocks, but the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and that different steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments described herein.

[0333] The methods relating to the embodiments described in this document above can be implemented in software form, and the encoding and / or decoding devices relating to this document may be included in, for example, video processing devices such as TVs, computers, smartphones, set-top boxes, and display devices.

[0334] In this document, when embodiments are implemented in software, the methods described above can be implemented by modules (processes, functions, etc.) that perform the functions described above. These modules are stored in memory and can be executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by various well-known means. The processor may include an ASIC (application-specific integrated circuit), other chipsets, logic circuits, and / or data processing devices. The memory may include ROM (read-only memory), RAM (random access memory), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this document may be implemented on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing may be implemented on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored on a digital storage medium.

[0335] Furthermore, the decoding and encoding devices to which the embodiments described in this document apply may include multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video interaction devices, real-time communication devices such as video communications, mobile streaming devices, storage media, camcorders, customized video (VoD) service providers, OTT video (Over the Top Video) devices, internet streaming service providers, 3D video devices, VR (virtual reality) devices, AR (argumente reality) devices, video telephone video devices, transportation terminals (e.g., vehicle terminals (including autonomous vehicles), airplane terminals, ship terminals, etc.), and medical video equipment, and may be used to process video signals or data signals. For example, OTT video (Over the Top Video) devices may include game consoles, Blu-ray players, internet access TVs, home theater systems, smartphones, tablet PCs, DVRs (Digital Video Recorders), etc.

[0336] Furthermore, the processing methods to which the embodiments of this document apply can be produced in the form of programs executed by a computer and stored on a computer-readable recording medium. Multimedia data having the data structure relating to the embodiments of this document can also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices that store data to be read by a computer. The computer-readable recording medium may include, for example, Blu-ray discs (BDs), general-purpose serial buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media embodied in the form of carrier waves (e.g., transmission over the Internet). Furthermore, a bitstream generated by an encoding method can be stored on a computer-readable recording medium or transmitted over a wired wireless network.

[0337] Furthermore, the embodiments described in this document can be embodied in a computer program product using program code, and the program code can be executed on a computer according to the embodiments described in this document. The program code can be stored on a computer-readable carrier.

[0338] Figure 17 shows an example of a content streaming system to which the embodiments disclosed in this document can be applied.

[0339] Referring to Figure 17, the content streaming system to which the embodiments described in this document apply can broadly include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.

[0340] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then transmitting this bitstream to the streaming server. As an alternative, if a multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server may be omitted.

[0341] The bitstream can be generated by an encoding method or bitstream generation method to which the embodiments of this document apply, and the streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0342] The streaming server transmits multimedia data to user devices based on user requests via a web server, and the web server acts as an intermediary to inform users about available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server controls the commands and responses between the devices within the content streaming system.

[0343] The streaming server can receive content from a media storage and / or encoding server. For example, if it begins receiving content from the encoding server, it can receive the content in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0344] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (such as smartwatches, smart glasses, and HMDs), digital TVs, desktop computers, and digital signage.

[0345] Each server within the aforementioned content streaming system can be operated as a distributed server, in which case the data received by each server can be processed in a distributed manner.

[0346] The claims described herein can be combined in various ways. For example, the technical features of the method claims herein can be combined to embody an apparatus, and the technical features of the apparatus claims herein can be combined to embody a method. Furthermore, the technical features of the method claims and the technical features of the apparatus claims herein can be combined to embody an apparatus, and the technical features of the method claims and the technical features of the apparatus claims herein can be combined to embody a method.

Claims

1. In a method for video decoding performed by a decoding device, A step of acquiring video information including predictive mode information and residual information via a bitstream, The steps include generating prediction samples based on the prediction mode information, The steps include generating a residual sample based on the residual information, A step of generating a reconstructed sample based on the predicted sample and the residual sample, wherein the reconstructed sample includes a reconstructed luma sample and a reconstructed chromatic sample. The steps include: deriving the ALF (adaptive loop filter) coefficients for the ALF (adaptive loop filter) procedure of the restored chromatic sample; A step of generating a filtered reconstructed chromatic sample based on the reconstructed chromatic sample and the ALF filter coefficients, The steps include: deriving cross-component filter coefficients for cross-component filtering, The step of generating a modified filtered restored chroma sample based on the restored chroma sample, the filtered restored chroma sample, and the cross-component filter coefficients, The aforementioned video information includes SPS (sequence parameter set), APS (adaptation parameter set), and slice header information. The ALF data included in the APS includes a first flag information relating to whether a cross-component filter for the Cb color component is signaled, and a second flag information relating to whether a cross-component filter for the Cr color component is signaled. Based on the first flag information and the second flag information, the ALF data includes information regarding the absolute value of the cross-component filter and information regarding the sign of the cross-component filter. The SPS includes an ALF enabled flag related to whether the ALF procedure is available, Whether the SPS includes a CCALF (cross-component adaptive loop filter) enabled flag, which is related to whether the aforementioned cross-component filtering is available, is determined based on the value of the ALF enabled flag. The SPS includes the CCALF enable flag based on the determination that the value of the ALF enable flag is 1. Based on the determination that the value of the ALF enabled flag in the SPS is 1, the slice header information includes an ALF enabled flag related to whether the ALF is enabled or not. Based on the determination that the value of the ALF usable flag included in the slice header information is 1 and the value of the CCALF usable flag included in the SPS is 1, the slice header information includes information on whether the CCALF is usable for the filtered restored chroma sample. A method wherein, based on the value of the information regarding whether the CCALF is available for the filtered reconstructed chroma sample, the slice header information includes the ID (identification) information of the APS associated with the CCALF for the filtered reconstructed chroma sample.

2. In a video encoding method performed by an encoding device, The current step is to determine the prediction mode of the block, The steps include generating prediction samples based on the prediction mode, The steps include: deriving a residual sample for the current block; The steps include: deriving conversion coefficients based on the conversion procedure for the residual sample; The steps include: deriving the quantized transformation coefficients based on the quantization procedure for the transformation coefficients; A step of generating residual information showing the quantized conversion coefficients, A step of generating a reconstructed sample based on the predicted sample and the residual sample, The steps include generating information related to the ALF (adaptive loop filter) and information related to the CCALF (cross-component ALF) for the reconstructed sample, The step includes encoding video information that includes the residual information, the information related to ALF, and the information related to CCALF, One reconstructed sample includes a reconstructed luma sample and a reconstructed chromatic sample. The aforementioned method for video encoding is The steps include: deriving the ALF filter coefficients for the ALF procedure of the restored chromatic sample; A step of generating a filtered reconstructed chromatic sample based on the reconstructed chromatic sample and the ALF filter coefficients, The steps include: deriving cross-component filter coefficients for cross-component filtering, The method further includes the step of generating a modified filtered restored chromatic sample based on the restored chromatic sample, the filtered restored chromatic sample, and the cross-component filter coefficients, The aforementioned video information includes SPS (sequence parameter set), APS (adaptation parameter set), and slice header information. The ALF data included in the APS includes a first flag information relating to whether a cross-component filter for the Cb color component is signaled, and a second flag information relating to whether a cross-component filter for the Cr color component is signaled. Based on the first flag information and the second flag information, the ALF data includes information regarding the absolute value of the cross-component filter and information regarding the sign of the cross-component filter. The SPS includes an ALF enabled flag related to whether the ALF procedure is available, Whether the SPS includes a CCALF (cross-component adaptive loop filter) enabled flag, which is related to whether the aforementioned cross-component filtering is available, is determined based on the value of the ALF enabled flag. The SPS includes the CCALF enable flag based on the determination that the value of the ALF enable flag is 1. Based on the determination that the value of the ALF enabled flag in the SPS is 1, the slice header information includes an ALF enabled flag related to whether the ALF is enabled or not. Based on the determination that the value of the ALF usable flag included in the slice header information is 1 and the value of the CCALF usable flag included in the SPS is 1, the slice header information includes information on whether the CCALF is usable for the filtered restored chroma sample. A method wherein, based on the value of the information regarding whether the CCALF is available for the filtered reconstructed chroma sample, the slice header information includes identification information of an APS (adaptation parameter set) associated with the CCALF for the filtered reconstructed chroma sample.

3. A method relating to data for video, A step of generating a bitstream for the aforementioned video, wherein the bitstream is: The current step is to determine the prediction mode of the block, The steps include generating prediction samples based on the prediction mode, The steps include: deriving a residual sample for the current block; The steps include: deriving conversion coefficients based on the conversion procedure for the residual sample; The steps include: deriving the quantized transformation coefficients based on the quantization procedure for the transformation coefficients; A step of generating residual information showing the quantized conversion coefficients, A step of generating a reconstructed sample based on the predicted sample and the residual sample, The steps include generating information related to the ALF (adaptive loop filter) and information related to the CCALF (cross-component ALF) for the reconstructed sample, A step of encoding video information including the residual information, the information related to ALF, and the information related to CCALF, which is generated based on the step of: The step of transmitting the data, which includes the bitstream, One reconstructed sample includes a reconstructed luma sample and a reconstructed chromatic sample. The aforementioned bitstream is The steps include: deriving the ALF filter coefficients for the ALF procedure of the restored chromatic sample; A step of generating a filtered reconstructed chromatic sample based on the reconstructed chromatic sample and the ALF filter coefficients, The steps include: deriving cross-component filter coefficients for cross-component filtering, A step of generating a modified filtered restored chroma sample based on the restored chroma sample, the filtered restored chroma sample, and the cross-component filter coefficients, is generated based on the following: The aforementioned video information includes SPS (sequence parameter set), APS (adaptation parameter set), and slice header information. The ALF data included in the APS includes a first flag information relating to whether a cross-component filter for the Cb color component is signaled, and a second flag information relating to whether a cross-component filter for the Cr color component is signaled. Based on the first flag information and the second flag information, the ALF data includes information regarding the absolute value of the cross-component filter and information regarding the sign of the cross-component filter. The SPS includes an ALF enabled flag related to whether the ALF procedure is available, Whether the SPS includes a CCALF (cross-component adaptive loop filter) enabled flag, which is related to whether the aforementioned cross-component filtering is available, is determined based on the value of the ALF enabled flag. The SPS includes the CCALF enable flag based on the determination that the value of the ALF enable flag is 1. Based on the determination that the value of the ALF enabled flag in the SPS is 1, the slice header information includes an ALF enabled flag related to whether the ALF is enabled or not. Based on the determination that the value of the ALF usable flag included in the slice header information is 1 and the value of the CCALF usable flag included in the SPS is 1, the slice header information includes information on whether the CCALF is usable for the filtered restored chroma sample. A method wherein, based on the value of the information regarding whether the CCALF is available for the filtered reconstructed chroma sample, the slice header information includes identification information of an APS (adaptation parameter set) associated with the CCALF for the filtered reconstructed chroma sample.