Image coding device and method of filtering base

By employing filtering-based video coding techniques, such as adaptive loop filtering and cross-component adaptive loop filtering, the challenges of compressing high-resolution video/images are addressed, resulting in improved compression efficiency and visual quality.

JP2025089401AActive Publication Date: 2025-06-12LG ELECTRONICS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025046995
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-08-29
Filing Date
2025-03-21
Publication Date
2025-06-12
Estimated Expiration
2040-08-31

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality video/images, such as 4K or UHD, leads to higher bitrates, resulting in increased transmission and storage costs. Additionally, the growing interest in immersive media like VR and AR requires efficient video/image compression techniques.

Method used

The implementation of an efficient filtering-based video coding method, including the application of adaptive loop filtering (ALF) and cross-component adaptive loop filtering (CCALF), to enhance video/image coding efficiency. This method involves filtering procedures for restored chroma samples based on restored luma samples, signaling information for CCALF in the Sequence Parameter Set (SPS), and deriving cross-component filter coefficients from ALF data.

Benefits of technology

The proposed solution increases overall video compression efficiency, improves subjective and objective visual quality, and enhances the filtering performance, leading to better image quality and coding accuracy for the chroma component of decoded pictures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025089401000001_ABST
    Figure 2025089401000001_ABST
Patent Text Reader

Abstract

To provide a device and a method of enhancing efficiency of image / video coding.SOLUTION: According to an embodiment of the present disclosure, a cross-component filter coefficient for cross-component filtering is derived. Modified filtered reconstructed chroma samples are generated on the basis of the cross-component filter coefficient. This embodiment increases accuracy of in-loop filtering.SELECTED DRAWING: Figure 12
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to a filtering-based video coding apparatus and method.

Background Art

[0002] In recent years, the demand for high-resolution and high-quality video / images such as 4K or UHD (Ultra High Definition) video / images of 8K or higher has been increasing in various fields. As the video / image data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to existing video / image data. Therefore, when transmitting video data using a medium such as an existing wired or wireless broadband line, or storing video / image data using an existing storage medium, the transmission cost and storage cost increase.

[0003] Also, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcast of video / images having video characteristics different from real-world video / images, such as game video, has been increasing.

[0004] Therefore, a highly efficient video / image compression technique is required to effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality video / images having various characteristics as described above.

Summary of the Invention

Means for Solving the Problems

[0005] According to an embodiment of this document, a method and apparatus for enhancing the efficiency of video / image coding are provided.

[0006] According to an embodiment of this document, an efficient filtering application method and apparatus are provided.

[0007] According to an embodiment of this document, an efficient ALF application method and apparatus are provided.

[0008] According to an embodiment of this document, a filtering procedure for a restored chroma sample based on a restored luma sample can be executed.

[0009] According to an embodiment of this document, a restored chroma sample filtered based on a restored luma sample can be modified.

[0010] According to an embodiment of this document, information regarding whether CCALF can be used in the SPS can be signaled.

[0011] According to an embodiment of this document, information regarding the value of a cross-component filter coefficient can be derived from the ALF data (general ALF data or CCALF data).

[0012] According to an embodiment of this document, the identifier (ID) information of the APS including the ALF data for deriving the cross-component filter coefficient in a slice can be signaled.

[0013] According to an embodiment of this document, information regarding a filter set index for CCALF can be signaled in CTU (block) units.

[0014] According to an embodiment of this document, a video / video decoding method performed by a decoding apparatus is provided.

[0015] According to an embodiment of this document, a decoding apparatus for performing video / video decoding is provided.

[0016] According to an embodiment of this document, a video / video encoding method performed by an encoding apparatus is provided.

[0017] According to one embodiment of the present document, an encoding device for performing video / video encoding is provided.

[0018] According to one embodiment of the present document, a computer-readable digital storage medium storing encoded video / video information generated by the video / video encoding method disclosed in at least one of the embodiments of the present document is provided.

[0019] According to one embodiment of the present document, a computer-readable digital storage medium storing encoded information or encoded video / video information that causes a decoding device to perform the video / video decoding method disclosed in at least one of the embodiments of the present document is provided.

Advantages of the Invention

[0020] According to one embodiment of the present document, the overall video / video compression efficiency can be increased.

[0021] According to one embodiment of the present document, the subjective / objective visual quality can be improved through efficient filtering.

[0022] According to one embodiment of the present document, the ALF procedure can be efficiently executed and the filtering performance can be improved.

[0023] According to one embodiment of the present document, the restored chroma samples filtered based on the restored luma samples can be corrected, and the image quality and coding accuracy for the chroma component of the decoded picture can be improved.

[0024] According to one embodiment of the present document, the CCALF procedure can be efficiently executed.

[0025] According to one embodiment of the present document, the ALF-related information can be efficiently signaled.

[0026] According to one embodiment of this document, CCALF-related information can be efficiently signaled.

[0027] According to one embodiment of this document, ALF and / or CCALF can be adaptively applied in units of pictures, slices, and / or coding blocks.

[0028] According to one embodiment of this document, when CCALF is used in an encoding and decoding method and apparatus for still images or moving images, the filter coefficients for CCALF and the on / off signaling method in units of blocks or CTUs are improved, and the coding efficiency can be increased.

Brief Description of the Drawings

[0029]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Mode for Carrying Out the Invention

[0030] This document can be modified in various ways and can have various embodiments, but specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to specific embodiments. The terms used in this specification are merely used to describe specific embodiments and are not used with the intention of limiting the technical idea of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as "including" or "having" are intended to specify the existence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the possibility of the existence or addition of one or more different features, numbers, steps, operations, components, parts, or combinations thereof.

[0031] On the other hand, each configuration on the drawings described in this document is shown independently for the convenience of explaining different characteristic functions, and does not mean that each configuration is implemented by separate hardware or separate software. For example, among the configurations, two or more configurations may be combined to form one configuration, or one configuration may be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included in the scope of rights of this document as long as they do not deviate from the essence of this document.

[0032] Hereinafter, with reference to the accompanying drawings, preferred embodiments of this document will be described in more detail. Hereinafter, the same reference numerals may be used for the same components on the drawings, and duplicate descriptions of the same components may be omitted.

[0033] This document relates to video / video coding. For example, the methods / examples disclosed in this document may be associated with the VVC (Versatile Video Coding) standard (ITU-T Rec.H.266), the next-generation video / image coding standard after VVC, or other video coding-related standards (e.g., the HEVC (High Efficiency Video Coding) standard (ITU-T Rec.H.265), the EVC (essential video coding) standard, the AVS2 standard, etc.).

[0034] This document presents various examples related to video / video coding, and unless otherwise mentioned, the examples may be combined with each other.

[0035] In this document, video can mean a collection of a series of images over time. A picture generally means a unit representing one image in a specific time period, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile can contain one or more CTUs (coding tree units). One picture can be composed of one or more slices / tiles. One picture can be composed of one or more tile groups. One tile group can be composed of one or more tiles.

[0036] Pixel or pel can mean the smallest unit that constitutes one picture (or video). Also, as a term corresponding to a pixel, "sample" can be used. A sample can generally indicate a pixel or a pixel value, can indicate only the pixel / pixel value of the luma component, or can also indicate only the pixel / pixel value of the chroma component. Or a sample may mean a pixel value in the spatial domain, and when these pixel values are converted to the frequency domain, it can also mean the conversion coefficients in the frequency domain.

[0037] A unit can represent the basic unit of video processing. A unit can include at least one of a specific area of a picture and information related to the area. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. A unit can, in some cases, be used interchangeably with terms such as "block" or "area". In general, an M×N block can include a set (or, array) of samples (or, sample array) or transform coefficients consisting of M columns and N rows.

[0038] In this document, the symbols " / " and "," are interpreted to mean "and / or". For example, "A / B" is interpreted as "A and / or B", and "A, B" is interpreted as "A and / or B". Additionally, "A / B / C" means "at least one of A, B, and / or C". Also, "A, B, C" also means "at least one of A, B, and / or C".

[0039] Further, in this document, the term “or” should be interpreted to indicate “and / or.” For instance, the expression “A or B” may comprise 1) only A, 2) only B, and / or 3) both A and B. In other words, the term “or” in this document should be interpreted to indicate “additionally or alternatively.”

[0040] In this specification, the expression “at least one of A and B” may mean “only A,” “only B,” or “both A and B.” Also, in this specification, expressions such as “at least one of A or B” and “at least one of A and / or B” may be interpreted in the same manner as “at least one of A and B.”

[0041] Also, in this specification, "at least one of A, B and C" may mean "only A", "only B", "only C", or "any combination of A, B and C". Also, "at least one of A, B or C" and "at least one of A, B and / or C" may mean "at least one of A, B and C".

[0042] Also, the parentheses used in this specification may mean "for example". Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction". In other words, "prediction" in this specification is not limited to "intra prediction", and "intra prediction" may be proposed as an example of "prediction". Also, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction".

[0043] The technical features separately described within one drawing in this specification may be embodied separately or simultaneously.

[0044] FIG. 1 schematically shows an example of a video / image coding system to which this document can be applied.

[0045] As shown in FIG. 1, the video / image coding system can include a source device and a receiving device. The source device can transmit encoded video / image information or data in file or streaming form to the receiving device via a digital storage medium or a network.

[0046] The source device can include a video source, an encoding device, and a transmission unit. The receiving device can include a receiving unit, a decoding device, and a renderer. The encoding device can be called a video / video encoding device, and the decoding device can be called a video / video decoding device. A transmitter can be provided in the encoding device. A receiver can be provided in the decoding device. The renderer can include a display unit, and the display unit can also be composed of a separate device or an external component.

[0047] The video source can obtain video / video through processes such as video / video capture, synthesis, or generation. The video source can include a video / video capture device and / or a video / video generation device. The video / video capture device can include, for example, one or more cameras, a video / video archive containing previously captured video / video, etc. The video / video generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / video. For example, virtual video / video can be generated through a computer or the like, and in this case, the video / video capture process can be replaced by the process of generating related data.

[0048] The encoding device can encode the input video / video. The encoding device can perform a series of procedures such as prediction, conversion, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / video information) can be output in the form of a bitstream.

[0049] The transmitting unit can transmit the encoded video / video information or data output in bitstream form to the receiving unit of the receiving device via a digital storage medium or network in file or streaming form. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating media files via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.

[0050] The decoding device can decode the video / video by performing a series of procedures such as inverse quantization, inverse transformation, prediction, etc., corresponding to the operation of the encoding device.

[0051] The renderer can render the decoded video / video. The rendered video / video can be displayed via the display unit.

[0052] Figure 2 is a drawing schematically explaining the configuration of a video / video encoding device to which this document can be applied. Hereinafter, the video encoding device can include the video encoding device.

[0053] As shown in FIG. 2, the encoding apparatus 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor 231. The adder 250 can be called a reconstructor or a reconstructed block generator. The above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. Also, the memory 270 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 270 as an internal / external component.

[0054] The video segmentation unit 210 can divide the input video (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be divided into multiple coding units with a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure can be applied first, and then the binary-tree structure and / or the ternary structure can be applied. Or, the binary-tree structure can also be applied first. The coding procedure according to the present disclosure can be performed based on the final coding unit that cannot be further divided. In this case, based on the coding efficiency according to the video characteristics, etc., the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units with a deeper depth, and the coding unit with the optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transformation unit (TU: Transform Unit). In this case, the prediction unit and the transformation unit can be divided or partitioned from the above-described final coding unit, respectively.The prediction unit can be a unit of sample prediction, and the conversion unit can be a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.

[0055] The unit can, in some cases, be used interchangeably with terms such as block or area. In general, an M×N block can represent a set such as samples or transform coefficients consisting of M columns and N rows. Samples can generally represent pixels or pixel values, and can represent only the pixels / pixel values of the luma component, or only the pixels / pixel values of the chroma component. Samples can be used as terms corresponding to pixels or pels in one picture (or video).

[0056] The subtraction unit 231 can subtract the prediction signal (predicted block, predicted sample, or predicted sample array) output from the prediction unit 220 from the input video signal (original block, original sample, or original sample array) to generate a residual signal (residual block, residual sample, or residual sample array), and the generated residual signal is transmitted to the conversion unit 232. The prediction unit 220 can perform prediction on the processing target block (hereinafter referred to as the current block) and generate a predicted block including predicted samples for the current block. The prediction unit 220 can determine whether intra prediction or inter prediction is applied in units of the current block or CU. The prediction unit can generate various pieces of information related to prediction, such as prediction mode information, and transmit them to the entropy encoding unit 240, as will be described later in the description of each prediction mode. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0057] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located adjacent to the current block or at a distance from it, depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the level of detail of the prediction direction. However, this is an example, and more or fewer directional prediction modes can be used depending on the settings. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the adjacent blocks.

[0058] The inter prediction unit 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring blocks can be called by names such as collocated reference blocks and collocated CUs (col CUs), and the reference picture including the temporal neighboring blocks can also be called a collocated picture (colPic). For example, the inter prediction unit 221 can construct a motion information candidate list based on adjacent blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and the motion vector difference is signaled, whereby the motion vector of the current block can be indicated.

[0059] The prediction unit 220 can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction, or both intra prediction and inter prediction simultaneously, for the prediction of a single block. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can perform intra block copy (IBC) for the prediction of a block. The intra block copy can be used for content video / movie coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed in the same way as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document.

[0060] The prediction signal generated via the inter prediction unit 221 and / or the intra prediction unit 222 can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 232 can apply a conversion technique to the residual signal to generate transform coefficients. For example, the conversion technique can include DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when expressing the relationship information between pixels in a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process may be applied to a pixel block having the same size of a square or may be applied to a block of a variable size that is not a square.

[0061] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients can be referred to as residual information. The quantization unit 233 can reorder the block-form quantized transform coefficients in a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 240 can also encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / video information) can be transmitted or stored in units of NAL (network abstraction layer) units in bitstream form. The video / video information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS), etc. Also, the video / video information can further include general constraint information. In this document, the signaling / transmitted information and / or syntax elements described later can be encoded via the above-described encoding procedure and included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be configured as an internal / external element of the encoding device 200 by a transmission unit (not shown) for transmission and / or a storage unit (not shown) for storage, or the transmission unit can also be included in the entropy encoding unit 240.

[0062] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transformation unit 235, a residual signal (residual block or residual sample) can be restored. The addition unit 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample, or reconstructed sample array) by adding the restored residual signal to the prediction signal output from the prediction unit 220. When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture and, as will be described later, can also be used for inter prediction of the next picture after passing through filtering.

[0063] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied in the picture encoding and / or restoration process.

[0064] The filtering unit 260 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and store the modified restored picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter, bilateral filter, and the like. The filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 240 as described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0065] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. When inter prediction is applied through this, the encoding device can avoid prediction mismatches in the encoding device 100 and the decoding device, and can also improve the encoding efficiency.

[0066] The DPB of memory 270 can store the corrected reconstructed picture for use as a reference picture in the inter prediction unit 221. Memory 270 can store the motion information of the blocks for which the motion information in the current picture has been derived (or encoded) and / or the motion information of the blocks in the already reconstructed picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. Memory 270 can store the reconstructed samples of the reconstructed blocks in the current picture and can transmit them to the intra prediction unit 222.

[0067] FIG. 3 is a drawing schematically explaining the configuration of a video / video decoding apparatus to which this document can be applied.

[0068] As shown in FIG. 3, the decoding apparatus 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filtering unit 350, and a memory 360. The predictor 330 can include an inter prediction unit 331 and an intra prediction unit 332. The residual processor 320 can include a dequantizer 321 and an inverse transformer 321. The above-described entropy decoding unit 310, residual processing unit 320, prediction unit 330, addition unit 340, and filtering unit 350 can be configured by one hardware component (for example, a decoder chipset or a processor) according to an embodiment. Further, the memory 360 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 360 as an internal / external component.

[0069] If a bitstream including video / video information is input, the decoding device 300 can restore the video corresponding to the process in which the video / video information was processed by the encoding device in FIG. 3. For example, the decoding device 300 can derive units / blocks based on the block division related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding can be, for example, a coding unit, and the coding unit can be divided according to a quad tree structure, a binary tree structure, and / or a ternary tree structure from a coding tree unit or a maximum coding unit. One or more transform units can be derived from the coding unit. Then, the restored video signal decoded and output via the decoding device 300 can be played back via a playback device.

[0070] The decoding device 300 can receive the signal output from the encoding device of FIG. 3 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information (e.g., video / video information) necessary for video restoration (or, picture restoration). The video / video information can further include information regarding various parameter sets, such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the video / video information can further include general constraint information. The decoding device can further decode a picture based on the information regarding the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for video restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded, the information adjacent to the syntax element to be decoded, the decoding information of the block to be decoded, or the information of the symbol / bin decoded in the previous step, predicts the occurrence probability of the bin according to the determined context model, performs arithmetic decoding of the bin, and can generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Among the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit 330, and the residual information that has undergone entropy decoding in the entropy decoding unit 310, that is, the quantized transform coefficient and related parameter information, can be input to the inverse quantization unit 321. Also, among the information decoded by the entropy decoding unit 310, the information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit can be a component of the entropy decoding unit 310. On the other hand, the decoding device according to this document can be called a video / video / picture decoding device, and the decoding device can be classified into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The information decoder can include the entropy decoding unit 310, and the sample decoder can include at least one of the inverse quantization unit 321, the inverse transform unit 322, the prediction unit 330, the addition unit 340, the filtering unit 350, and the memory 360.

[0071] In the inverse quantization unit 321, the quantized transform coefficient can be inverse quantized to output a transform coefficient. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the reordering can be performed based on the coefficient scan order performed in the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficient using a quantization parameter (for example, quantization step size information) to obtain a transform coefficient.

[0072] In the inverse conversion unit 322, the conversion coefficient is inversely converted to obtain a residual signal (residual block, residual sample array).

[0073] The prediction unit can perform prediction on the current block and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.

[0074] The prediction unit can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can apply not only intra prediction or inter prediction for predicting a block, but also can apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can perform intra block copy (IBC) for predicting a block. The intra block copy can be used for content video / motion video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed in the same manner as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction.

[0075] The intra prediction unit 331 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located adjacent to the current block or at a distance therefrom depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 can also determine the prediction mode to be applied to the current block using the prediction mode applied to an adjacent block.

[0076] The inter prediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between an adjacent block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent block can include a spatial neighboring block existing within the current picture and a temporal neighboring block existing in the reference picture. For example, the inter prediction unit 332 can configure a motion information candidate list based on the adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter prediction for the current block.

[0077] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the predictor. When there is no residual for the block to be processed, such as when the skip mode is applied, the predicted block can be used as the restored block.

[0078] The adder 340 can be called a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next block to be processed within the current picture, can be output after filtering as described later, or can also be used for inter prediction of the next picture.

[0079] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.

[0080] The filtering unit 350 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can send the modified restored picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0081] The (corrected) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction unit 332. The memory 360 can store the motion information of the block from which the motion information in the current picture has been derived (or decoded) and / or the motion information of the block in the already reconstructed picture. The stored motion information can be transmitted to the inter prediction unit 332 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the reconstructed samples of the reconstructed blocks in the current picture and can transmit them to the intra prediction unit 331.

[0082] In this specification, the embodiments described in the prediction unit 330, inverse quantization unit 321, inverse transform unit 322, filtering unit 350, etc. of the decoding apparatus 300 can be applied so as to be identical or corresponding to the prediction unit 220, inverse quantization unit 234, inverse transform unit 235, filtering unit 260, etc. of the encoding apparatus 200 as well.

[0083] As described above, in performing video coding, prediction is executed to increase the compression efficiency. Through this, a predicted block including prediction samples for the current block which is the block to be coded can be generated. Here, the predicted block includes prediction samples in the spatial domain (or pixel domain). The predicted block is also derived in the encoding apparatus and the decoding apparatus, and the encoding apparatus can increase the video coding efficiency by signaling information (residual information) regarding the residual between the original block and the predicted block which is not the original sample value of the original block to the decoding apparatus. The decoding apparatus can derive a residual block including residual samples based on the residual information, and can generate a reconstructed block including reconstructed samples by combining the residual block and the predicted block, and can generate a reconstructed picture including the reconstructed block.

[0084] The residual information can be generated via conversion and quantization procedures. For example, an encoding device can derive a residual block between the original block and the predicted block, execute a conversion procedure on the residual samples (residual sample array) included in the residual block to derive conversion coefficients, and execute a quantization procedure on the conversion coefficients to derive quantized conversion coefficients, so as to signal (via a bitstream) the related residual information to a decoding device. Here, the residual information can include information such as the value information, position information, conversion technique, conversion kernel, quantization parameter, etc. of the quantized conversion coefficients. The decoding device can execute an inverse quantization / inverse conversion procedure based on the residual information to derive residual samples (or a residual block). The decoding device can generate a restored picture based on the predicted block and the residual block. Also, the encoding device can inverse quantize / inverse convert the quantized conversion coefficients for reference in the inter prediction of subsequent pictures to derive a residual block, and generate a restored picture based on this.

[0085] In this document, at least one of quantization / inverse quantization and / or conversion / inverse conversion can be omitted. When the quantization / inverse quantization is omitted, the quantized conversion coefficients can be called conversion coefficients. When the conversion / inverse conversion is omitted, the conversion coefficients can also be called coefficients or residual coefficients, or, for the sake of consistency of expression, can still be called conversion coefficients.

[0086] In this document, the quantized transform coefficients and the transform coefficients can each be referred to as the transform coefficients and the scaled transform coefficients, respectively. In this case, the residual information can include information regarding the transform coefficients, and the information regarding the transform coefficients can be signaled via a residual coding syntax. The transform coefficients can be derived based on the residual information (or the information regarding the transform coefficients), and the scaled transform coefficients can be derived via an inverse transform (scaling) with respect to the transform coefficients. The residual samples can be derived based on an inverse transform (transformation) with respect to the scaled transform coefficients. This can be applied / expressed similarly in other parts of this document.

[0087] The prediction unit of the encoding device / decoding device can derive prediction samples by performing inter prediction in block units. Inter prediction can indicate a prediction derived in a method that depends on data elements (e.g., sample values or motion information) of pictures other than the current picture. When inter prediction is applied to the current block, a predicted block (prediction sample array) for the current block can be induced based on a reference block (reference sample array) specified by a motion vector on a reference picture pointed to by the index of the reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information of the current block can be predicted in block, sub-block, or sample units based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and an index of a reference picture. The motion information can further include information on the inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). When inter prediction is applied, the adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called by names such as a collocated reference block or a collocated CU (colCU), and the reference picture including the temporal neighboring block may also be called a collocated picture (colPic). For example, a candidate list of motion information can be constructed based on the adjacent blocks of the current block, and a flag or index information indicating which candidate is selected (used) can be signaled in order to derive the motion vector and / or the index of the reference picture of the current block.Inter prediction is performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the motion information of the current block may be the same as the motion information of the selected adjacent block. In the case of the skip mode, different from the merge mode, a residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of the selected adjacent block is used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.

[0088] The motion information can include L0 motion information and / or L1 motion information according to an inter-prediction type (such as L0 prediction, L1 prediction, Bi prediction, etc.). The motion vector in the L0 direction can be called the L0 motion vector or MVL0, and the motion vector in the L1 direction can be called the L1 motion vector or MVL1. The prediction based on the L0 motion vector can be called L0 prediction, the prediction based on the L1 motion vector can be called L1 prediction, and the prediction based on both the L0 motion vector and the L1 motion vector can be called bi (Bi) prediction. Here, the L0 motion vector can indicate a motion vector related to the reference picture list L0 (L0), and the L1 motion vector can indicate a motion vector related to the reference picture list L1 (L1). The reference picture list L0 can include reference pictures that are earlier in output order than the current picture, and the reference picture list L1 can include pictures that are later in output order than the current picture. The earlier picture can be called a forward (reference) picture, and the later picture can be called a backward (reference) picture. The reference picture list L0 can further include pictures that are later in output order than the current picture as reference pictures. In this case, the earlier picture can be indexed first within the reference picture list L0, and the later picture can be indexed thereafter. The reference picture list L1 can further include pictures that are earlier in output order than the current picture as reference pictures. In this case, the later picture can be indexed first within the reference picture list 1, and the earlier picture can be indexed thereafter. Here, the output order can correspond to the POC (picture order count) order (order).

[0089] FIG. 4 exemplarily shows a hierarchical structure for a coded video / picture.

[0090] Referring to FIG. 4, the coded video / video can be divided into a VCL (video coding layer) that handles the decoding process of the video / video itself, a lower system that transmits and stores the encoded information, and a NAL (network abstraction layer) that exists between the VCL and the lower system and is responsible for the network adaptation function.

[0091] In the VCL, VCL data including compressed video data (slice data) can be generated, or parameter sets including information such as a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), a Video Parameter Set (VPS), or a Supplemental Enhancement Information (SEI) message additionally required in the decoding process of the video can be generated.

[0092] In the NAL, a NAL unit can be generated by adding header information (NAL unit header) to the RBSP (Raw Byte Sequence Payload) generated in the VCL. At this time, the RBSP means slice data, parameter sets, SEI messages, etc. generated in the VCL. The NAL unit header can include NAL unit type information specified by the RBSP data included in the corresponding NAL unit.

[0093] As shown in the above drawings, the NAL unit can be divided into a VCL NAL unit and a Non-VCL NAL unit by the RBSP generated in the VCL. The VCL NAL unit can mean a NAL unit including information (slice data) for the video, and the Non-VCL NAL unit can mean a NAL unit including information (parameter set or SEI message) required for decoding the video.

[0094] The above-mentioned VCL NAL units and Non-VCL NAL units can be transmitted via a network with header information attached according to the data standards of the lower-level system. For example, the NAL unit can be transformed into a data form of a predetermined standard such as the H.266 / VVC file format, RTP (Real-time Transport Protocol), TS (Transport Stream), etc., and transmitted via various networks.

[0095] As described above, the NAL unit type can be specified by the RBSP data structure included in the corresponding NAL unit, and information regarding such NAL unit type can be stored in the NAL unit header and signaled.

[0096] For example, depending on whether the NAL unit contains information (slice data) for video, it can be roughly classified into a VCL NAL unit type and a Non-VCL NAL unit type. The VCL NAL unit type can be classified according to the nature and type of the picture included in the VCL NAL unit, and the Non-VCL NAL unit type can be classified according to the type of parameter set, etc.

[0097] The following is an example of the NAL unit type specified according to the type of parameter set included in the Non-VCL NAL unit type, etc.

[0098] - APS (Adaptation Parameter Set) NAL unit: The type for the NAL unit containing APS

[0099] - DPS (Decoding Parameter Set) NAL unit: The type for the NAL unit containing DPS

[0100] -VPS (Video Parameter Set) NAL unit: Type for NAL unit containing VPS

[0101] -SPS (Sequence Parameter Set) NAL unit: Type for NAL unit containing SPS

[0102] -PPS (Picture Parameter Set) NAL unit: Type for NAL unit containing PPS

[0103] -PH (Picture header) NAL unit: Type for NAL unit containing PH

[0104] The above-mentioned NAL unit types have syntax information for the NAL unit type, and the syntax information can be stored in the NAL unit header and signaled. For example, the syntax information is nal_unit_type, and the NAL unit type can be specified by the nal_unit_type value.

[0105] On the other hand, as described above, one picture can include a plurality of slices, and one slice can include a slice header and slice data. In this case, one picture header can be further added to the plurality of slices (slice header and slice data set) within one picture. The picture header (picture header syntax) can include information / parameters that are commonly applicable to the picture. In this document, a slice can be mixed or substituted with a tile group. Also, in this document, a slice header can be mixed or substituted with a type group header.

[0106] The slice header (slice header syntax, slice header information) can include information / parameters that are commonly applicable to the slice. The APS (APS syntax) or PPS (PPS syntax) can include information / parameters that are commonly applicable to one or more slices or pictures. The SPS (SPS syntax) can include information / parameters that are commonly applicable to one or more sequences. The VPS (VPS syntax) can include information / parameters that are commonly applicable to multiple layers. The DPS (DPS syntax) can include information / parameters that are commonly applicable to the entire video. The DPS can include information / parameters related to the concatenation of CVS (coded video sequence). In this document, the high level syntax (HLS) can include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, and slice header syntax.

[0107] In this document, the video / video information encoded from an encoding device and signaled in the form of a bitstream to a decoding device can include not only information related to partitioning within a picture, intra / inter prediction information, residual information, in-loop filtering information, etc., but also the information included in the slice header, the information included in the picture header, the information included in the APS, the information included in the PPS, the information included in the SPS, the information included in the VPS, and / or the information included in the DPS. Further, the video / video information can further include the information of the NAL unit header.

[0108] On the one hand, in order to compensate for the difference between the original video and the restored video due to errors occurring in the compression encoding process such as quantization, as described above, an in-loop filtering procedure can be executed on the restored sample or the restored picture. As described above, in-loop filtering can be executed in the filter section of the encoding device and the filter section of the decoding device, and a deblocking filter, SAO, and / or an adaptive loop filter (ALF) can be applied. For example, the ALF procedure can be executed after the deblocking filtering procedure and / or the SAO procedure is completed. However, also in this case, the deblocking filtering procedure and / or the SAO procedure can be omitted.

[0109] Specific descriptions regarding picture restoration and filtering are described below. In video / video coding, restored blocks can be generated based on intra prediction / inter prediction for each block unit, and a restored picture including the restored blocks can be generated. When the current picture / slice is an I picture / slice, the blocks included in the current picture / slice can be restored based only on intra prediction. On the other hand, when the current picture / slice is a P or B picture / slice, the blocks included in the current picture / slice can be restored based on intra prediction or inter prediction. In this case, intra prediction can be applied to some blocks within the current picture / slice, and inter prediction can also be applied to the remaining blocks.

[0110] Intra prediction can indicate a prediction that generates a predicted sample for a current block based on reference samples within a picture to which the current block belongs (hereinafter referred to as the current picture). When intra prediction is applied to the current block, adjacent reference samples to be used for the intra prediction of the current block can be derived. The adjacent reference samples of the current block can include samples adjacent to the left boundary of the current block of size nW×nH and a total of 2×nH samples adjacent to the bottom-left, samples adjacent to the top boundary of the current block and a total of 2×nW samples adjacent to the top-right, and 1 sample adjacent to the top-left of the current block. Alternatively, the adjacent reference samples of the current block can also include a plurality of columns of upper adjacent samples and a plurality of rows of left adjacent samples. Also, the adjacent reference samples of the current block can include a total of nH samples adjacent to the right boundary of the current block of size nW×nH, a total of nW samples adjacent to the bottom boundary of the current block, and 1 sample adjacent to the bottom-right of the current block.

[0111] However, some of the adjacent reference samples of the current block may not have been decoded yet or may not be available. In this case, the decoder can substitute samples that are not available with samples that are available to form adjacent reference samples to be used for prediction. Alternatively, adjacent reference samples to be used for prediction can be formed through interpolation of available samples.

[0112] When an adjacent reference sample is derived, (i) a predicted sample can be induced based on the average or interpolation of the neighboring reference samples of the current block, and (ii) the predicted sample can also be induced based on the reference samples among the adjacent reference samples of the current block that exist in a specific (predicted) direction with respect to the predicted sample. In the case of (i), it is called a non-directional mode or a non-angle mode, and in the case of (ii), it can be called a directional mode or an angular mode. Also, among the adjacent reference samples, based on the predicted sample of the current block, the predicted sample can also be generated through interpolation between the second adjacent sample and the first adjacent sample located in the opposite direction of the prediction direction of the intra prediction mode of the current block. The above-mentioned case can be called linear interpolation intra prediction (LIP). Also, a chroma predicted sample can be generated based on the luma sample using a linear model. In this case, it can be called the LM mode. Also, a temporary predicted sample of the current block is derived based on the filtered adjacent reference samples, and a predicted sample of the current block is derived by taking a weighted sum of at least one reference sample derived by the intra prediction mode among the existing adjacent reference samples, i.e., the non-filtered adjacent reference samples, and the temporary predicted sample. The above-mentioned case can be called PDPC (Position dependent intra prediction). Also, the intra prediction coding can be executed by selecting the reference sample line with the highest prediction accuracy from among the adjacent multiple reference sample lines of the current block and using the reference sample located in the prediction direction on the corresponding line, and indicating (signaling) the used reference sample line to the decoding device.In the above-described case, it can be called multi-reference line (MRL) intra prediction or MRL-based intra prediction. Also, the current block can be divided into vertical or horizontal sub-partitions, and intra prediction can be performed based on the same intra prediction mode, and adjacent reference samples can be derived and used in units of the sub-partitions. That is, in this case, the intra prediction mode for the current block is also applied to the sub-partitions, and by deriving and using adjacent reference samples in units of the sub-partitions, in some cases, the intra prediction performance can be improved. Such a prediction method can be called intra sub-partitions (ISP) or ISP-based intra prediction. The intra prediction methods described above can be called intra prediction types, distinguished from the intra prediction modes in Tables of Contents 1 and 2. The intra prediction type can be called by various terms such as an intra prediction technique or an additional intra prediction mode. For example, the intra prediction type (or, an additional intra prediction mode, etc.) can include at least one of the above-described LIP, PDPC, MRL, and ISP. A general intra prediction method excluding specific intra prediction types such as LIP, PDPC, MRL, and ISP can be called a normal intra prediction type. The normal intra prediction type can be generally applied when the above-described specific intra prediction types are not applied, and prediction can be performed based on the above-described intra prediction mode. On the other hand, if necessary, post-processing filtering for the derived prediction samples can also be performed.

[0113] Specifically, the intra prediction procedure can include an intra prediction mode / type determination step, an adjacent reference sample derivation step, and an intra prediction mode / type-based prediction sample derivation step. Also, if necessary, a post-processing filtering step for the derived prediction samples can also be performed.

[0114] FIG. 5 is a flowchart for explaining an intra prediction-based block restoration method in an encoding device. The method of FIG. 5 can include steps S500, S510, S520, S530, and S540.

[0115] S500 can be executed by the intra prediction unit 222 of the encoding device, and S510 to S530 can be executed by the residual processing unit 230 of the encoding device. Specifically, S510 can be executed by the subtraction unit 231 of the encoding device, S520 can be executed by the conversion unit 232 and the quantization unit 233 of the encoding device, and S530 can be executed by the inverse quantization unit 234 and the inverse conversion unit 235 of the encoding device. In S500, prediction information can be derived by the intra prediction unit 222 and encoded by the entropy encoding unit 240. Residual information is derived through S510 and S520 and can be encoded by the entropy encoding unit 240. The residual information is information regarding the residual sample. The residual information can include information regarding the quantized transform coefficients for the residual sample. As described above, the residual sample is derived as a transform coefficient through the conversion unit 232 of the encoding device, and the transform coefficient can be derived as a quantized transform coefficient through the quantization unit 233. Information regarding the quantized transform coefficients can be encoded by the entropy encoding unit 240 through the residual coding procedure.

[0116] The encoding device performs intra prediction on the current block (S500). The encoding device can derive an intra prediction mode for the current block and derive adjacent reference samples of the current block, and generate prediction samples within the current block based on the intra prediction mode and the adjacent reference samples. Here, the intra prediction mode determination, adjacent reference sample derivation, and prediction sample generation procedures can be executed simultaneously, or one procedure can be executed prior to the other procedures. For example, the intra prediction unit 222 of the encoding device can include a prediction mode / type determination unit, a reference sample derivation unit, and a prediction sample derivation unit. The prediction mode / type determination unit determines the intra prediction mode / type for the current block, the reference sample derivation unit derives the adjacent reference samples of the current block, and the prediction sample derivation unit can derive the motion samples of the current block. On the other hand, although not shown, when a prediction sample filtering procedure described later is executed, the intra prediction unit 222 can further include a prediction sample filter unit (not shown). The encoding device can determine the mode applied to the current block among a plurality of intra prediction modes. The encoding device can compare the RD cost for the intra prediction mode and determine the optimal intra prediction mode for the current block.

[0117] On the other hand, the encoding device can also execute a prediction sample filtering procedure. The prediction sample filtering can be referred to as post filtering. Some or all of the prediction samples can be filtered by the prediction sample filtering procedure. In some cases, the prediction sample filtering procedure can be omitted.

[0118] The encoding device derives a residual sample for the current block based on a prediction sample (S510). The encoding device can compare the prediction sample with the original sample of the current block based on phase and derive the residual sample.

[0119] The encoding device converts / quantizes the residual sample to derive quantized transform coefficients (S520), and then can perform inverse quantization / inverse transformation processing on the quantized transform coefficients again to derive a (corrected) residual sample (S530). The reason for performing inverse quantization / inverse transformation again after conversion / quantization is to derive the same residual sample as the residual sample derived by the decoding device, as described above.

[0120] The encoding device can generate a restored block including a restored sample for the current block based on the prediction sample and the (corrected) residual sample (S540). A restored picture for the current picture can be generated based on the restored block.

[0121] As described above, the encoding device can encode video information including prediction information regarding the intra prediction (for example, prediction mode information indicating a prediction mode) and residual information regarding the intra and residual samples, and output the encoded video information in the form of a bitstream. The residual information can include a residual coding syntax. The encoding device can convert / quantize the residual sample to derive quantized transform coefficients. The residual information can include information regarding the quantized transform coefficients.

[0122] FIG. 6 is a flowchart for explaining an intra prediction-based block restoration method in a decoding apparatus. The method of FIG. 6 can include steps S600, S610, S620, S630, and S640. The decoding apparatus can execute operations corresponding to the operations executed by the encoding apparatus.

[0123] S600 to S620 can be executed by an intra prediction unit 331 of the decoding apparatus, and the prediction information of S600 and the residual information of S630 can be obtained from a bitstream by an entropy decoding unit 310 of the decoding apparatus. A residual processing unit 320 of the decoding apparatus can derive residual samples for a current block based on the residual information. Specifically, an inverse quantization unit 321 of the residual processing unit 320 can perform inverse quantization to derive conversion coefficients based on the quantized conversion coefficients derived based on the residual information, and an inverse conversion unit 322 of the residual processing unit can perform inverse conversion on the conversion coefficients to derive residual samples for the current block. S640 can be executed by an addition unit 340 or a restoration unit of the decoding apparatus.

[0124] Specifically, the decoding apparatus can derive an intra prediction mode for a current block based on received prediction mode information (S600). The decoding apparatus can derive adjacent reference samples of the current block (S610). The decoding apparatus can generate prediction samples within the current block based on the intra prediction mode and the adjacent reference samples (S620). In this case, the decoding apparatus can execute a prediction sample filtering procedure. The prediction sample filtering can be called post-filtering. Some or all of the prediction samples can be filtered by the prediction sample filtering procedure. In some cases, the prediction sample filtering procedure can be omitted.

[0125] The decoding device generates residual samples for the current block based on the received residual information (S630). The decoding device can generate restored samples for the current block based on the prediction samples and the residual samples, and derive a restored block including the restored samples (S640). A restored picture for the current picture can be generated based on the restored block.

[0126] Here, the intra prediction unit 331 of the decoding device can include a prediction mode / type determination unit, a reference sample derivation unit, and a prediction sample derivation unit. The prediction mode / type determination unit determines an intra prediction mode for the current block based on the prediction mode information obtained by the entropy decoding unit 310 of the decoding device. The reference sample derivation unit derives adjacent reference samples of the current block, and the prediction sample derivation unit can derive prediction samples of the current block. On the other hand, although not shown in the figure, when the above-described prediction sample filtering procedure is executed, the intra prediction unit 331 can further include a prediction sample filter unit (not shown).

[0127] The prediction information can include intra prediction mode information and / or intra prediction type information. The intra prediction mode information can include, for example, flag information (e.g., intra_luma_mpm_flag) indicating whether the MPM (most probable mode) is applied to the current block or the remaining mode is applied. When the MPM is applied to the current block, the prediction mode information can further include index information (e.g., intra_luma_mpm_idx) indicating one of the intra prediction mode candidates (MPM candidates). The intra prediction mode candidates (MPM candidates) can be composed of an MPM candidate list or an MPM list. Also, when the MPM is not applied to the current block, the intra prediction mode information can further include remaining mode information (e.g., intra_luma_mpm_remainder) indicating one of the remaining intra prediction modes excluding the intra prediction mode candidates (MPM candidates). The decoding device can determine the intra prediction mode of the current block based on the intra prediction mode information. A separate MPM list can be configured for the above-mentioned MIP.

[0128] In addition, the intra prediction type information can be embodied in various forms. As an example, the intra prediction type information can include intra prediction type index information indicating one of the intra prediction types. As another example, the intra prediction type information includes reference sample line information (e.g., intra_luma_ref_idx) indicating whether the MRL is applied to the current block and which reference sample line is used if it is applied, ISP flag information (e.g., intra_subpartitions_mode_flag) indicating whether the ISP is applied to the current block, ISP type information (e.g., intra_subpartitions_split_flag) indicating the split type of the subpartition if the ISP is applied, flag information indicating whether PDCP is applicable, or flag information indicating whether LIP is applicable, and can include at least one of them. Further, the intra prediction type information can include an MIP flag indicating whether MIP is applied to the current block.

[0129] The intra prediction mode information and / or the intra prediction type information can be encoded / decoded through the coding method described in this document. For example, the intra prediction mode information and / or the intra prediction type information can be encoded / decoded through entropy coding (e.g., CABAC, CAVLC coding) based on truncated (rice) binary code.

[0130] The prediction unit of the encoding device / decoding device can derive a prediction sample by performing inter prediction in block units. Inter prediction can be a prediction derived in a manner that is dependent on data elements (e.g., sample values, or motion information, etc.) of picture(s) other than the current picture (Inter prediction can be a prediction derived in a manner that is dependent on data elements (e.g., sample values or motion information) of picture(s) other than the current picture). When inter prediction is applied to the current block, a predicted block (predicted sample array) for the current block can be induced based on a reference block (reference sample array) specified by a motion vector on a reference picture pointed to by a reference picture index. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information of the current block can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the adjacent block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. When inter prediction is applied, the adjacent block can include a spatial neighboring block existing within the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block can be called by names such as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block can also be called a collocated picture (colPic).For example, a motion information candidate list can be configured based on adjacent blocks of a current block, and a flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block can be signaled. Inter prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the motion information of the current block is the same as the motion information of the selected adjacent block. In the case of skip mode, different from the merge mode, a residual signal is not transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of the selected adjacent block is used as a motion vector predictor, and a motion vector difference can be signaled. In this case, the motion vector of the current block can be derived by using the sum of the motion vector predictor and the motion vector difference.

[0131] FIG. 7 is a flowchart for explaining an inter prediction-based block restoration method in an encoding device. The method of FIG. 7 can include steps S700, S710, S720, S730, and S740.

[0132] S700 can be executed by the inter-prediction unit 221 of the encoding apparatus, and S710 to S730 can be executed by the residual processing unit 230 of the encoding apparatus. Specifically, S710 can be executed by the subtraction unit 231 of the encoding apparatus, S720 can be executed by the conversion unit 232 and the quantization unit 233 of the encoding apparatus, and S730 can be executed by the inverse quantization unit 234 and the inverse conversion unit 235 of the encoding apparatus. In S700, prediction information can be derived by the inter-prediction unit 221 and encoded by the entropy encoding unit 240. Residual information is derived through S710 and S720 and can be encoded by the entropy encoding unit 240. The residual information is information regarding the residual sample. The residual information can include information regarding the quantized conversion coefficients for the residual sample. As described above, the residual sample is derived as conversion coefficients through the conversion unit 232 of the encoding apparatus, and the conversion coefficients can be derived as quantized conversion coefficients through the quantization unit 233. Information regarding the quantized conversion coefficients can be encoded by the entropy encoding unit 240 through the residual coding procedure.

[0133] The encoding device performs inter prediction on the current block (S700). The encoding device can derive the inter prediction mode and motion information of the current block and generate a predicted sample of the current block. Here, the inter prediction mode determination, motion information derivation, and predicted sample generation procedures can be executed simultaneously, or one procedure can be executed prior to the other procedures. For example, the inter prediction unit 221 of the encoding device can include a prediction mode determination unit, a motion information derivation unit, and a predicted sample derivation unit, where the prediction mode determination unit determines the prediction mode for the current block, the motion information derivation unit derives the motion information of the current block, and the predicted sample derivation unit can derive the motion sample of the current block. For example, the inter prediction unit 221 of the encoding device can search for a block similar to the current block within a certain area (search area) of the reference picture via motion estimation, and derive a reference block whose difference from the current block is the minimum or below a certain criterion. Based on this, a reference picture index indicating the reference picture where the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding device can determine the mode applied to the current block among various prediction modes. The encoding device can compare the RD cost for the various prediction modes to determine the optimal prediction mode for the current block.

[0134] For example, when the skip mode or the merge mode is applied to the current block, the encoding device configures a merge candidate list to be described later, and can derive a reference block from among the reference blocks pointed to by the merge candidates included in the merge candidate list, where the difference between the current block and the reference block is the smallest or below a certain criterion. In this case, a merge candidate associated with the derived reference block can be selected, and merge index information indicating the selected merge candidate can be generated and signaled to the decoding device. The motion information of the current block can be derived using the motion information of the selected merge candidate.

[0135] As another example, when the (A)MVP mode is applied to the current block, the encoding device configures an (A)MVP candidate list to be described later, and the motion vector of an mvp (motion vector predictor) candidate selected from among the mvp candidates included in the (A)MVP candidate list can be used as the mvp of the current block. In this case, for example, the motion vector pointing to the reference block derived by the above-described motion estimation can be used as the motion vector of the current block, and the mvp candidate having the motion vector with the smallest difference from the motion vector of the current block among the mvp candidates can be the selected mvp candidate. An MVD (motion vector difference), which is the difference obtained by subtracting the mvp from the motion vector of the current block, can be derived. In this case, information regarding the MVD can be signaled to the decoding device. Also, when the (A)MVP mode is applied, the value of the reference picture index can be configured with reference picture index information and signaled separately to the decoding device.

[0136] The encoding device can derive a residual sample based on the prediction sample (S710). The encoding device can derive the residual sample by comparing the original sample of the current block with the prediction sample.

[0137] The encoding device converts / quantizes the residual sample, derives the quantized transform coefficient (S720), and then can perform inverse quantization / inverse transformation processing on the quantized transform coefficient again to derive the (corrected) residual sample (S730). The reason for performing inverse quantization / inverse transformation again after conversion / quantization is to derive the same residual sample as the residual sample derived by the decoding device, as described above.

[0138] The encoding device can generate a restored block including a restored sample for the current block based on the prediction sample and the (corrected) residual sample (S740). A restored picture for the current picture can be generated based on the restored block.

[0139] Although not shown in the figure, as described above, the encoding device can encode video information including prediction information and residual information. The encoding device can output the encoded video information in the form of a bit stream. The prediction information is information related to the prediction procedure, and can include prediction mode information (e.g., skip flag, merge flag or mode index, etc.) and information related to motion information. The information related to the motion information can include candidate selection information (e.g., merge index, mvp flag or mvp index) for deriving a motion vector. Also, the information related to the motion information can include information related to the MVD and / or reference picture index information described above. Also, the information related to the motion information can include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information related to the residual sample. The residual information can include information related to the quantized transform coefficient for the residual sample.

[0140] The output bitstream can be stored in a (digital) storage medium and transmitted to a decoding device, or can also be transmitted to the decoding device via a network.

[0141] FIG. 8 is a flowchart for explaining an inter-prediction based block restoration method in a decoding device. The method of FIG. 8 can include steps S800, S810, S820, S830, and S840. The decoding device can perform operations corresponding to the operations executed by the encoding device.

[0142] S800 to S820 can be executed by the inter-prediction unit 332 of the decoding device, and the prediction information of S800 and the residual information of S830 can be obtained from the bitstream by the entropy decoding unit 310 of the decoding device. The residual processing unit 320 of the decoding device can derive residual samples for the current block based on the residual information. Specifically, the inverse quantization unit 321 of the residual processing unit 320 can perform inverse quantization to derive conversion coefficients based on the quantized conversion coefficients derived based on the residual information, and the inverse conversion unit 322 of the residual processing unit can perform inverse conversion on the conversion coefficients to derive residual samples for the current block. S840 can be executed by the addition unit 340 or the restoration unit of the decoding device.

[0143] Specifically, the decoding device can determine a prediction mode for the current block based on the received prediction information (S800). The decoding device can determine which inter-prediction mode is applicable to the current block based on the prediction mode information in the prediction information.

[0144] For example, the merge mode can be applied to the current block based on the merge flag, or it can be determined whether (A) the MVP mode is determined. Alternatively, one can be selected from various inter-prediction mode candidates based on the mode index. The inter-prediction mode candidates can include the skip mode, the merge mode, and / or (A) the MVP mode, or can include various inter-prediction modes described later.

[0145] The decoding device derives motion information of the current block based on the determined inter-prediction mode (S810). For example, when the skip mode or the merge mode is applied to the current block, the decoding device can construct a merge candidate list described later and select one merge candidate from the merge candidates included in the merge candidate list. The selection can be executed based on the selection information (merge index) described above. The motion information of the current block can be derived using the motion information of the selected merge candidate. The motion information of the selected merge candidate can be used as the motion information of the current block.

[0146] As another example, when the (A)MVP mode is applied to the current block, the decoding device configures an (A)MVP candidate list described below, and can use the motion vector of the mvp (motion vector predictor) candidate selected from among the mvp candidates included in the (A)MVP candidate list as the mvp of the current block. The selection can be performed based on the selection information (mvp flag or mvp index) described above. In this case, the MVD of the current block can be derived based on the information regarding the MVD, and the motion vector of the current block can be derived based on the mvp of the current block and the MVD. Also, the reference picture index of the current block can be derived based on the reference picture index information. The picture pointed to by the reference picture index within the reference picture list regarding the current block can be derived as the reference picture to be referred to for the inter prediction of the current block.

[0147] On the other hand, as will be described later, the motion information of the current block can be derived without configuring a candidate list, and in this case, the motion information of the current block can be derived by the procedure disclosed in the prediction mode described later. In this case, the candidate list configuration as described above can be omitted.

[0148] The decoding device can generate a prediction sample for the current block based on the motion information of the current block (S820). In this case, the reference picture can be derived based on the reference picture index of the current block, and the prediction sample of the current block can be derived using the sample of the reference block pointed to by the motion vector of the current block on the reference picture. In this case, as will be described later, in some cases, a prediction sample filtering procedure for all or part of the prediction samples of the current block can be further executed.

[0149] For example, the inter prediction unit 332 of the decoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode for the current block is determined based on the prediction mode information received by the prediction mode determination unit. The motion information (such as a motion vector and / or a reference picture index, etc.) of the current block is derived based on the information related to the motion information received by the motion information derivation unit. The prediction sample derivation unit can derive the prediction sample of the current block.

[0150] The decoding device generates a residual sample for the current block based on the received residual information (S830). The decoding device can generate a restored sample for the current block based on the prediction sample and the residual sample, and derive a restored block including the restored sample (S840). A restored picture for the current picture can be generated based on the restored block.

[0151] A variety of inter-prediction modes can be used for predicting the current block within a picture. For example, various modes such as the merge mode, skip mode, MVP (motion vector prediction) mode, Affine mode, sub-block merge mode, MMVD (merge with MVD) mode, etc. can be used. DMVR (Decoder side motion vector refinement) mode, AMVR (adaptive motion vector resolution) mode, Bi-prediction with CU-level weight (BCW), Bi-directional optical flow (BDOF), etc. can be used additionally or alternatively as accompanying modes. The Affine mode may also be referred to as the affine motion prediction mode. The MVP mode may also be referred to as the AMVP (advanced motion vector prediction) mode. In this document, candidates for motion information derived by some modes and / or some modes may be included as one of the candidates for motion information of other modes. For example, the HMVP candidate may be added as a merge candidate in the merge / skip mode, or may be added as an mvp candidate in the MVP mode.

[0152] Prediction mode information indicating the inter prediction mode of the current block can be signaled from an encoding device to a decoding device. The prediction mode information is included in a bitstream and can be received by the decoding device. The prediction mode information can include index information indicating one of a number of candidate modes. Alternatively, the inter prediction mode may be indicated through hierarchical signaling of flag information. In this case, the prediction mode information can include one or more flags. For example, a skip flag is signaled to indicate whether the skip mode can be applied. When the skip mode is not applied, a merge flag is signaled to indicate whether the merge mode can be applied. When the merge mode is not applied, it may be indicated that the MVP mode is to be applied, or additional flags for further classification may be signaled. The affine mode may be signaled as an independent mode, or may be signaled as a mode subordinate to the merge mode or the MVP mode, etc. For example, the affine mode can include an affine merge mode and an affine MVP mode.

[0153] On the one hand, information indicating whether the above-mentioned list0 (L0) prediction, list1 (L1) prediction, or bi-prediction is used for the current block (current coding unit) can be signaled to the current block. The said information can be referred to as motion prediction direction information, inter prediction direction information, or inter prediction indication information, and can be configured / encoded / signaled, for example, in the form of the syntax element of inter_pred_idc. That is, the syntax element of inter_pred_idc can indicate whether the above-mentioned list0 (L0) prediction, list1 (L1) prediction, or bi-prediction is used for the current block (current coding unit). In this document, for the convenience of explanation, the inter prediction type (L0 prediction, L1 prediction, or BI prediction) pointed to by the syntax element of inter_pred_idc can be represented as the motion prediction direction. L0 prediction may be represented as pred_L0, L1 prediction as pred_L1, and bi-prediction as pred_BI. For example, according to the value of the syntax element of inter_pred_idc, the prediction type as shown in the following table can be determined.

[0154]

Table 1

[0155] As described above, one picture can include one or more slices. A slice can have one of the slice types including an intra slice, a predictive slice, and a bi-predictive slice. The slice type can be indicated based on slice type information. For blocks within an intra slice, inter prediction is not used for prediction, and only intra prediction can be used. Of course, in this case as well, it is also possible to code and signal the original sample values without prediction. For blocks within a predictive slice, intra prediction or inter prediction can be used, and when inter prediction is used, only uni prediction can be used. On the other hand, for blocks within a bi-predictive slice, intra prediction or inter prediction can be used, and when inter prediction is used, up to maximum bi-prediction can be used.

[0156] L0 and L1 can include reference pictures that have been encoded / decoded before the current picture. For example, L0 can include reference pictures before and / or after the current picture in POC order, and L1 can include reference pictures after and / or before the current picture in POC order. In this case, an index of a reference picture that is relatively lower than the reference picture before the current picture in POC order may be assigned to L0, and an index of a reference picture that is relatively lower than the reference picture after the current picture in POC order may be assigned to L1. In the case of a bi-predictive slice, bi-prediction can be applied, and in this case as well, uni-directional bi-prediction can be applied, or both-directional bi-prediction can be applied. Both-directional bi-prediction can be called true bi-prediction.

[0157] As described above, a residual block (residual sample) can be derived based on a predicted block (predicted sample) derived through prediction in the encoding stage, and residual information can be generated through conversion / quantization on the residual sample. The residual information can include information on quantized conversion coefficients. The residual information can be included in video / image information, and the video / image information can be encoded and transmitted to a decoding device in the form of a bitstream. The decoding device can obtain the residual information from the bitstream and derive a residual sample based on the residual information. Specifically, the decoding device can derive quantized conversion coefficients based on the residual information and derive a residual block (residual sample) through an inverse quantization / inverse conversion procedure.

[0158] On the other hand, at least one of the (inverse) conversion and / or (inverse) quantization procedures can be omitted.

[0159] The in-loop filtering procedure executed for the restored picture is described below. Restored samples, blocks, pictures (or modified filtered samples, blocks, pictures) modified through the in-loop filtering procedure can be generated, and the modified (modified filtered) restored picture can be output as a decoded picture in a decoding apparatus, and can also be stored in a decoded picture buffer or memory of an encoding apparatus / decoding apparatus and then be used as a reference picture in an inter prediction procedure during encoding / decoding of subsequent pictures. As described above, the in-loop filtering procedure can include a deblocking filtering procedure, an SAO (sample adaptive offset) procedure, and / or an ALF (adaptive loop filter) procedure, etc. In this case, one or some of the deblocking filtering procedure, the SAO (sample adaptive offset) procedure, the ALF (adaptive loop filter) procedure, and the bilateral filter procedure can be sequentially applied, or all of them can also be sequentially applied. For example, after the deblocking filtering procedure is applied to the restored picture, the SAO procedure can be executed. Or, for example, after the deblocking filtering procedure is applied to the restored picture, the ALF procedure can be executed. This can also be executed similarly in an encoding apparatus.

[0160] Deblocking filtering is a filtering technique that removes distortions occurring at the boundaries between blocks in a restored picture. The deblocking filtering procedure can, for example, derive a target boundary in the restored picture, determine the bS (boundary strength) for the target boundary, and perform deblocking filtering for the target boundary based on the bS. The bS can be determined based on, for example, the prediction modes of two adjacent blocks of the target boundary, the motion vector difference, whether the reference pictures are the same, and whether there are non-zero valid coefficients.

[0161] SAO is a method for compensating the offset difference between a restored picture and an original picture in sample units, and can be applied based on types such as, for example, Band Offset and Edge Offset. According to SAO, samples can be classified into different categories by each SAO type, and an offset value can be added to each sample based on the category. The filtering information for SAO can include information on whether SAO can be applied, SAO type information, SAO offset value information, etc. SAO can also be applied to the restored picture after the deblocking filtering is applied.

[0162] ALF (Adaptive Loop Filter) is a technique for filtering in sample units based on filter coefficients according to a filter shape for a restored picture. The encoding device can determine, via comparison between the restored picture and the original picture, whether ALF can be applied, the ALF shape and / or the ALF filtering coefficient, etc., and can signal it to the decoding device. That is, the filtering information for ALF can include information on whether ALF can be applied, ALF filter shape information, ALF filtering coefficient information, etc. ALF can also be applied to the restored picture after the deblocking filtering is applied.

[0163] FIG. 9 shows an example of the ALF filter shape.

[0164] (a) of FIG. 9 shows a 7×7 diamond filter shape, and (b) shows a 5×5 diamond filter shape. In FIG. 9, Cn in the filter shape indicates a filter coefficient. In the case of the Cn, when n is the same, it indicates that the same filter coefficient can be assigned. In this document, the position and / or unit to which the filter coefficient is assigned by the filter shape of the ALF can be called a filter tab. At this time, one filter coefficient can be assigned to each filter tab, and the form in which the filter tabs are arranged can correspond to the filter shape. The filter tab located at the center of the filter shape can be called a center filter tab. The same filter coefficient can be assigned to two filter tabs with the same n value existing at positions corresponding to each other with reference to the center filter tab. For example, in the case of a 7×7 diamond filter shape, it includes 25 filter tabs, and since the filter coefficients of C0 to C11 are assigned in a centrosymmetric form, the filter coefficients of only 13 filter coefficients can be assigned to the 25 filter tabs. Also, for example, in the case of a 5×5 diamond filter shape, it includes 13 filter tabs, and since the filter coefficients of C0 to C5 are assigned in a centrosymmetric form, the filter coefficients of only 7 filter coefficients can be assigned to the 13 filter tabs. For example, in order to reduce the data amount of information regarding the signaled filter coefficients, among the 13 filter coefficients for the 7×7 diamond filter shape, 12 filter coefficients are (explicitly) signaled, and 1 filter coefficient can be (implicitly) derived. Also, for example, among the 7 filter coefficients for the 5×5 diamond filter shape, 6 filter coefficients are (explicitly) signaled, and 1 filter coefficient can be (implicitly) derived.

[0165] According to one embodiment of the present document, the ALF parameters used for the ALF procedure can be signaled via an APS (adaptation parameter set). The ALF parameters can be derived from filter information or ALF data for the ALF.

[0166] As described above, ALF is a type of in-loop filtering technique that can be applied in video / videotape coding. ALF can be performed using a Wiener-based adaptive filter. This is to minimize the MSE (mean square error) between the original sample and the decoded sample (or, restored sample). The high level design for the ALF tool can incorporate syntax elements accessible in the SPS and / or slice header (or, tile group header).

[0167] In one example, prior to filtering for each 4×4 luma block, geometric transformations such as rotation or diagonal and vertical flipping can be applied to the filter coefficients f(k, l) and the corresponding filter clipping values c(k, l) that depend on the slope value calculated for the block. This is the same as these transformations being applied to the samples within the filter support region. Generating other blocks to which ALF is applied and aligning these blocks according to their directions are similar.

[0168] For example, three transformations, diagonal, vertical flip, and rotation can be performed based on the following mathematical formulas.

[0169]

Equation

[0170]

Equation

[0171] [Number]

[0172] In the above equations (1) to (3), K is the size of the filter. 0 ≦ k, 1 ≦ K - 1 are the coefficient coordinates. For example, (0, 0) is the upper left corner coordinate, and / or (K - 1, K - 1) is the lower right corner coordinate. The relationship between the transformation and the four slopes in the four directions can be summarized as shown in the following table.

[0173] [Table 2]

[0174] The ALF filter parameters can be signaled with the APS and the slice header. In one APS, up to 25 luma filter coefficients and the clipping value index can be signaled. In one APS, up to 8 chroma filter coefficients and the clipping value index can be signaled. To reduce the bit overhead, the filter coefficients of different classifications for the luma component can be merged. In the slice header, the index of the APS (referred to by the current slice) used for the current slice can be signaled.

[0175] The clipping value index decoded from the APS can be used to determine the clipping value by using the luma table of the clipping value and the chroma table of the clipping value. These clipping values are dependent on the internal bit depth. More specifically, the luma table of the clipping value and the chroma table of the clipping value can be derived based on the following equations.

[0176]

Number

[0177]

Number

[0178] In the above equation, B is the internal bitdepth, and N is the number of allowed clipping values (a pre-determined number). For example, N is 4.

[0179] In the slice header, up to 7 APS indexes can be signaled to indicate the luma filter set used for the current slice. The filtering procedure can be further controlled at the CTB level. For example, a flag can be signaled to indicate whether ALF is applied to the luma CTB. The luma CTB can select one filter set from 16 fixed filter sets and the filter sets from APS. The filter set index can be signaled for the luma CTB to indicate which filter set is applied. The 16 fixed filter sets can be predefined and can be hard-coded in both the encoder and the decoder.

[0180] For the chroma component, the APS index can be signaled in the slice header to indicate the chroma filter set used for the current slice. At the CTB level, if there are two or more chroma filter sets in APS, the filter index can be signaled for each chroma CTB.

[0181] The filter coefficients can be quantized (norm) with respect to 128. To limit the multiplication complexity, bitstream conformance can be applied, and thus, the coefficient values of non - central position are within the range from 0 to 28, and / or the coefficient values at other positions are within the range from - 27 to 27 - 1. The central position coefficient can be pre - determined (considered) as 128 without being signaled in the bitstream.

[0182] When ALF is available for the current block, each sample R(i, j) can be filtered, and the filtered result R′(i, j) can be expressed as follows.

[0183] [Equation]

[0184] In the above equation, f(k, l) is the decoded filter coefficient, K(x, y) is the clipping function, and c(k, l) is the decoded clipping parameter. For example, the variables k and / or l can vary from - L / 2 to L / 2. Here, L can indicate the filter length. The clipping function K(x, y)=min(y, max(-y, x)) can correspond to the function Clip3(-y, y, x).

[0185] In one example, to reduce the line buffer requirements of ALF, modified block classification and filtering can be applied for samples adjacent to the horizontal CTU boundary. For this purpose, a virtual boundary can be defined.

[0186] FIG. 10 is a drawing for explaining a virtual boundary applied to a filtering procedure according to an embodiment of the present document. FIG. 11 shows an example of an ALF procedure using a virtual boundary according to an embodiment of the present document. FIG. 11 is explained together with FIG. 10.

[0187] Referring to FIG. 10, the virtual boundary is a line defined by shifting the horizontal CTU boundary by about N samples. In one example, N is 4 for the luma component and / or N is 2 for the chroma component.

[0188] In FIG. 10, the modified block classification can be applied to the luma component. Only the samples on the virtual boundary can be used for the 1D Laplacian slope calculation of the 4×4 blocks on the virtual boundary. Similarly, only the samples below the virtual boundary can be used for the 1D Laplacian slope calculation of the 4×4 blocks below the virtual boundary. The quantization of the activity value A can be scaled in consideration of the reduced number of samples used in the 1D Laplacian slope calculation.

[0189] For the filtering procedure, a symmetric padding operation at the virtual boundary can be used for the luma and chroma components. Referring to FIG. 10, when samples filtered below the virtual boundary are located, the adjacent samples located above the virtual boundary can be padded. On the other hand, the corresponding samples on the other side can also be padded symmetrically.

[0190] The procedure described with reference to FIG. 11 can also be used for slice, block, and / or tile boundaries when the filter is not available across the boundary. For ALF block classification, only samples contained within the same slice, block, and / or tile can be used, and the activity value can be scaled thereby. For ALF filtering, symmetric padding can be applied for each of the horizontal and / or vertical directions with respect to the horizontal and / or vertical boundaries.

[0191] FIG. 12 is a diagram for explaining a cross-component adaptive loop filtering (CCALF) procedure according to an embodiment of the present document. The CCALF procedure can also be referred to as a cross-component filtering procedure.

[0192] In one aspect, the ALF procedure can include a general ALF procedure and a CCALF procedure. That is, the CCALF procedure can refer to a partial procedure of the ALF procedure. In another aspect, the filtering procedure can include a deblocking procedure, an SAO procedure, an ALF procedure, and / or a CCALF procedure.

[0193] CC-ALF can refine each chroma component using luma sample values. CC-ALF is controlled by the (video) information of the bitstream, and the video information can include (a) information regarding filter coefficients for each chroma component and (b) information regarding a mask that controls filter application for a block of samples. The filter coefficients can be signaled in the APS, and the block size and mask can be signaled at the slice level.

[0194] Referring to FIG. 12, the CC-ALF can operate by applying a linear diamond-shaped filter (FIG. 12(b)) to the luma channel for each chroma component. The filter coefficients are sent to the APS, scaled by a factor of 210, and rounded for fixed-point representation. The application of the filter is controlled by a variable block size and can be signaled by the context coding flag received for each block of samples. The block size, together with the CC-ALF usable flag, can be received at the slice level for each chroma component. The block size (for chroma samples) is 16×16, 32×32, 64×64, or 128×128.

[0195] In the following embodiments, a method is proposed for re-filtering or modifying the restored chroma samples filtered by the ALF based on the restored luma samples.

[0196] One embodiment of this document is related to the filter on / off transmission and filter coefficient transmission of the CC-ALF. As described above, the information (syntax elements) in the syntax table disclosed in this document can be included in the video / video information, configured / encoded by the encoding device, and transmitted to the decoding device in the form of a bitstream. The decoding device can parse / decode the information (syntax elements) in the corresponding syntax table. The decoding device can execute picture / image / video decoding procedures (specifically, for example, the CC-ALF procedure) based on the decoded information. The same applies to other embodiments below.

[0197] The following table shows a part of the syntax of the slice header information according to the embodiments of this document.

[0198]

Table 3

[0199] The following table shows exemplary semantics regarding the syntax elements included in the said table.

[0200] [Table 4]

[0201] Referring to the two tables above, when sps_cross_component_alf_enabled_flag is 1 in the slice header, parsing of slice_cross_component_alf_cb_enabled_flag can be performed to determine whether CC-ALF is applicable within the corresponding slice. When slice_cross_component_alf_cb_enabled_flag is 1, CC-ALF is applied to the corresponding Cb slice, and when slice_cross_component_alf_cb_reuse_temporal_layer_filter is 1, filters of the same existing temporal layer can be reused. When slice_cross_component_alf_cb_enabled_flag is 0, CC-ALF can be applied using the filters in the corresponding APS (adaptation parameter set) id via slice_cross_component_alf_cb_aps_idparsing. slice_cross_component_alf_cb_log2_control_size_minus4 can mean the CC-ALF application block unit in the Cb slice.

[0202] For example, when the value of slice_cross_component_alf_cb_log2_control_size_minus4 is 0, the applicability of CC-ALF is determined in units of 16×16. When the value of slice_cross_component_alf_cb_log2_control_size_minus4 is 1, the applicability of CC-ALF is determined in units of 32×32. When the value of slice_cross_component_alf_cb_log2_control_size_minus4 is 2, the applicability of CC-ALF is determined in units of 64×64. When the value of slice_cross_component_alf_cb_log2_control_size_minus4 is 3, the applicability of CC-ALF is determined in units of 128×128. Also, for Cr CC-ALF, syntax with the same structure as described above is used.

[0203] The following table shows exemplary syntax regarding ALF data.

[0204]

Table 5

[0205] The following table shows exemplary semantics regarding the syntax elements included in the above table.

[0206]

Table 6-1

[0207]

Table 6-2

[0208] Referring to the two tables, the CC-ALF syntax elements are configured to be transmitted independently and applied independently regardless of the existing (general) ALF syntax structure. That is, CC-ALF can be applied even when the ALF tool on the SPS is off. Since CC-ALF must be operable independently of the existing ALF structure, a new hardware pipeline design is required. This causes an increase in hardware implementation cost and an increase in hardware delay.

[0209] Also, ALF determines whether to apply to both luma and chroma videos in CTU units, and the result of the determination is transmitted to the decoder via signaling. However, since the applicability of variable CC-ALF from 16×16 to 128×128 units is determined and applied, a collision can occur between the existing ALF structure and CC-ALF. This causes problems in hardware implementation and also increases the line buffer for various variable CC-ALF applications.

[0210] In the present invention, by integrally applying the CC-ALF syntax structure to the ALF syntax structure, an attempt is made to solve the above-mentioned hardware implementation problems of CC-ALF.

[0211] According to an embodiment of this document, in order to determine whether CC-ALF is used (applied), the sequence parameter set (SPS) can include a CC-ALF usable flag (sps_ccalf_enable_flag). The CC-ALF usable flag can be transmitted independently of the ALF usable flag (sps_alf_enabled_flag) for determining whether ALF is used (applied).

[0212] The following table shows a part of the exemplary syntax of the SPS according to this embodiment.

[0213]

Table 7

[0214] Referring to the above table, CC-ALF can only be applied when ALF is always operating. That is, the CC-ALF available flag (sps_ccalf_enabled_flag) can only be parsed when the ALF available flag (sps_alf_enabled_flag) is 1. According to the above table, CC-ALF and ALF can be combined. The CC-ALF available flag can indicate whether CC-ALF is available (which can be related).

[0215] The following table shows a part of the exemplary syntax regarding the slice header.

[0216]

Table 8

[0217] Referring to the above table, the parsing of sps_ccalf_enabled_flag can be executed only when sps_alf_enabled_flag is 1. The syntax elements included in the above table can be described based on Table 4. In one example, the video information encoded by the encoding device or acquired (received) by the decoding device can include slice header information (slice_header()). Based on the determination that the value of the CCALF available flag (sps_ccalf_flag) is 1, the slice header information can include a first flag (slice_cross_component_alf_cb_enabeld_flag) related to whether CC-ALF is available for the Cb color component of the filtered restored chroma samples, and a second flag (slice_cross_component_alf_cr_enabeld_flag) related to whether CC-ALF is available for the Cr color component of the filtered restored chroma samples.

[0218] In one example, based on the determination that the value of the first flag (slice_cross_component_alf_cb_enabeld_flag) is 1, the slice header information can include the ID information (slice_cross_component_alf_cb_aps_id) of the first APS for deriving the cross-component filter coefficients for the Cb color component. Based on the determination that the value of the second flag (slice_cross_component_alf_cr_enabeld_flag) is 1, the slice header information can include the ID information (slice_cross_component_alf_cr_aps_id) of the second APS for deriving the cross-component filter coefficients for the Cr color component.

[0219] The following table shows a part of the SPS syntax according to another example of this embodiment.

[0220]

Table 9

[0221] The following table illustratively shows a part of the slice header syntax.

[0222]

Table 10

[0223] Referring to Table 9 above, when ChromaArrayType is not 0 and the ALF available flag (sps_alf_enabled_flag) is 1, the SPS can include the CCALF available flag (sps_ccalf_enabled_flag). For example, when ChromaArrayType is not 0, the CCALF available flag can be transmitted via the SPS based on the case where the chroma format is not monochrome.

[0224] Referring to Table 9 above, based on the case where ChromaArrayType is not 0, information regarding CCALF (slice_cross_component_alf_cb_enabled_flag, slice_cross_component_alf_cb_aps_id, slice_cross_component_alf_cr_enabled_flag, slice_cross_component_alf_cr_aps_id) can be included in the slice header information.

[0225] In one example, the video information encoded by the encoding device or obtained by the decoding device may include the SPS. The SPS may include a first ALF availability flag (sps_alf_enabled_flag) related to whether ALF is available. For example, based on the determination that the value of the first ALF availability flag is 1, the SPS may include a CCALF availability flag related to whether the cross-component filtering is available. In another example, when sps_alf_enabled_flag is 1 without using sps_ccalf_enabled_flag, CCALF can always be applied (sps_ccalf_enabled_flag == 1).

[0226] The following table shows a part of the slice header syntax according to another example of this embodiment.

[0227]

Table 11

[0228] Referring to the above table, the parsing of the CCALF availability flag (sps_ccalf_enabled_flag) can be performed only when the ALF availability flag (sps_alf_enabled_flag) is 1.

[0229] The following table shows exemplary semantics regarding the syntax elements included in the above table.

[0230]

Table 12

[0231] The slice_ccalf_chroma_idc in the above table can also be explained by the semantics in the following table.

[0232]

Table 13

[0233] The following table shows a part of the slice header syntax according to other examples of this embodiment.

[0234]

Table 14

[0235] The syntax elements included in the table can be explained by Table 12 or Table 13. Also, when the chroma format is not monochrome, CCALF-related information can be included in the slice header.

[0236] The following table shows a part of the slice header syntax according to other examples of this embodiment.

[0237]

Table 15

[0238] The following table shows exemplary semantics regarding the syntax elements included in the table.

[0239]

Table 16

[0240] The following table shows a part of the slice header syntax according to other examples of this embodiment. The syntax elements included in the following table can be explained by Table 12 or Table 13.

[0241]

Table 17

[0242] Referring to the above table, it is possible to determine whether to apply slice-level ALF and CC-ALF at once via the slice_alf_enabled_flag. After parsing slice_alf_chroma_idc, if the first ALF available flag (sps_alf_enabled_flag) is 1, slice_ccalf_chroma_idc can be parsed.

[0243] Referring to the above table, it is possible to determine whether sps_ccalf_enabled_flag is 1 in the slice header information only when slice_alf_enabled_flag is 1. The slice header information can include a second ALF available flag (slice_alf_enabled_flag) related to whether ALF is available. Based on the determination that the value of the second ALF available flag (slice_alf_enabled_flag) is 1, the CCALF is available for the slice.

[0244] The following table exemplarily shows a part of the APS syntax. The syntax element adaptation_parameter_set_id can indicate the identifier information (ID information) of the APS.

[0245]

Table 18

[0246] The following table shows an exemplary syntax regarding ALF data.

[0247]

Table 19

[0248] Referring to the two tables, the APS can include ALF data (alf_data()). The APS including ALF data can be called an ALF APS (ALF type APS). That is, the type of the APS including ALF data is the ALF type. The type of the APS can be determined by information regarding the APS type or a syntax element (aps_params_type). The ALF data can include a Cb filter signal flag (alf_cross_component_cb_filter_signal_flag or alf_cc_cb_filter_signal_flag) related to whether a cross-component filter for the Cb color component is signaled. The ALF data can include a Cr filter signal flag (alf_cross_component_cr_filter_signal_flag or alf_cc_cr_filter_signal_flag) related to whether a cross-component filter for the Cr color component is signaled.

[0249] In one example, based on the Cr filter signal flag, the ALF data can include information regarding the absolute value of the cross-component filter coefficient for the Cr color component (alf_cross_component_cr_coeff_abs) and information regarding the sign of the cross-component filter coefficient for the Cr color component (alf_cross_component_cr_coeff_sign). Based on the information regarding the absolute value of the cross-component filter coefficient for the Cr color component and the information regarding the sign of the cross-component filter coefficient for the Cr color component, the cross-component filter coefficient for the Cr color component can be derived.

[0250] In one example, the ALF data can include information (alf_cross_component_cb_coeff_abs) regarding the absolute value of the cross-component filter coefficient for the Cb color component and information (alf_cross_component_cb_coeff_sign) regarding the sign of the cross-component filter coefficient for the Cb color component. Based on the information regarding the absolute value of the cross-component filter coefficient for the Cb color component and the information regarding the sign of the cross-component filter coefficient for the Cb color component, the cross-component filter coefficient for the Cb color component can be derived.

[0251] The following table shows the syntax for ALF data according to another example.

[0252] [Table 20]

[0253] Referring to the table above, after transmitting alf_cross_component_filter_signal_flag first, if alf_cross_component_filter_signal_flag is 1, the Cb / Cr filter signal flag can be transmitted. That is, alf_cross_component_filter_signal_flag determines whether to integrate Cb / Cr and transmit the CC-ALF filter coefficient.

[0254] The following table shows the syntax for ALF data according to another example.

[0255] [Table 21]

[0256] The following table shows exemplary semantics regarding the syntax elements included in the table above.

[0257]

Table 22-1

[0258]

Table 22-2

[0259] The following table shows the syntax for ALF data according to other examples.

[0260]

Table 23

[0261] The following table shows exemplary semantics for the syntax elements included in the said table.

[0262]

Table 24-1

[0263]

Table 24-2

[0264] In the said two tables, the order of exp-Golomb binarization for parsing the alf_cross_component_cb_coeff_abs[j] and alf_cross_component_cr_coeff_abs[j] syntax can be defined as one of the values from 0 to 9.

[0265] Referring to the two tables, the ALF data can include a Cb filter signal flag (alf_cross_component_cb_filter_signal_flag or alf_cc_cb_filter_signal_flag) related to whether a cross-component filter for the Cb color component is signaled. Based on the Cb filter signal flag (alf_cross_component_cb_filter_signal_flag), the ALF data can include information (ccalf_cb_num_alt_filters_minus1) related to the number of cross-component filters for the Cb color component. Based on the information related to the number of cross-component filters for the Cb color component, the ALF data can include information (alf_cross_component_cb_coeff_abs) regarding the absolute value of the cross-component filter coefficients for the Cb color component and information (alf_cross_component_cr_coeff_sign) regarding the sign of the cross-component filter coefficients for the Cb color component. Based on the information regarding the absolute value of the cross-component filter coefficients for the Cb color component and the information regarding the sign of the cross-component filter coefficients for the Cb color component, the cross-component filter coefficients for the Cb color component can be derived.

[0266] In one example, the ALF data can include a Cr filter signal flag (alf_cross_component_cr_filter_signal_flag or alf_cc_cr_filter_signal_flag) related to whether a cross-component filter for the Cr color component was signaled. Based on the Cr filter signal flag (alf_cross_component_cr_filter_signal_flag), the ALF data can include information (ccalf_cr_num_alt_filters_minus1) related to the number of cross-component filters for the Cr color component. Based on the information related to the number of cross-component filters for the Cr color component, the ALF data can include information (alf_cross_component_cr_coeff_abs) regarding the absolute value of the cross-component filter coefficients for the Cr color component and information (alf_cross_component_cr_coeff_sign) regarding the sign of the cross-component filter coefficients for the Cr color component. Based on the information regarding the absolute value of the cross-component filter coefficients for the Cr color component and the information regarding the sign of the cross-component filter coefficients for the Cr color component, the cross-component filter coefficients for the Cr color component can be derived.

[0267] The following table shows the syntax for a coding tree unit according to one embodiment of this document.

[0268] [Table 25]

[0269] The following table shows exemplary semantics for the syntax elements included in the above table.

[0270] [Table 26]

[0271] The following table shows the coding tree unit syntax according to other examples of this embodiment.

[0272] [Table 27]

[0273] Referring to the above table, CCALF can be applied in CTU units. In one example, the video information can include information related to a coding tree unit (coding_tree_unit()). The information related to the coding tree unit can include information (ccalf_ctb_flag[0]) regarding whether a cross-component filter is applied to the current block of the Cb color component, and / or information (ccalf_ctb_flag[1]) regarding whether a cross-component filter is applied to the current block of the Cr color component. Also, the information related to the coding tree unit can include information (ccalf_ctb_filter_alt_idx[0]) regarding the filter set index of the cross-component filter applied to the current block of the Cb color component, and / or information (ccalf_ctb_filter_alt_idx[1]) regarding the filter set index of the cross-component filter applied to the current block of the Cr color component. The syntax can be adaptively transmitted by the slice_ccalf_enabled_flag and slice_ccalf_chroma_idc syntax.

[0274] The following table shows the coding tree unit syntax according to other examples of this embodiment.

[0275] [Table 28]

[0276] The following table shows exemplary semantics regarding the syntax elements included in the said table.

[0277]

Table 29

[0278] The following table shows the coding tree unit syntax according to other examples of this embodiment. The syntax elements included in the following table can be explained by Table 29.

[0279]

Table 30

[0280] In one example, the video information can include information regarding a coding tree unit (coding_tree_unit()). The information regarding the coding tree unit can include information (ccalf_ctb_flag[0]) regarding whether a cross-component filter is applied to the current block of the Cb color component, and / or information (ccalf_ctb_flag[1]) regarding whether a cross-component filter is applied to the current block of the Cr color component. Also, the information regarding the coding tree unit can include information (ccalf_ctb_filter_alt_idx[0]) regarding the filter set index of the cross-component filter applied to the current block of the Cb color component, and / or information (ccalf_ctb_filter_alt_idx[1]) regarding the filter set index of the cross-component filter applied to the current block of the Cr color component.

[0281] Figures 13 and 14 schematically show an example of a video / video encoding method and related components according to an embodiment of this document.

[0282] The method disclosed in FIG. 13 can be executed by the encoding device disclosed in FIG. 2 or FIG. 14. Specifically, for example, S1300 to S1330 in FIG. 13 can be executed by the residual processing unit 230 of the encoding device in FIG. 14, S1340 in FIG. 13 can be executed by the addition unit 250 of the encoding device in FIG. 14, S1350 in FIG. 13 can be executed by the filtering unit 260 of the encoding device in FIG. 14, and S1360 in FIG. 13 can be executed by the entropy encoding unit 240 of the encoding device in FIG. 14. Also, although not shown in FIG. 13, in FIG. 13, a prediction sample or prediction-related information can be derived by the prediction unit 220 of the encoding device, and a bitstream can be generated from the residual information or prediction-related information by the entropy encoding unit 240 of the encoding device. The method disclosed in FIG. 13 can include the embodiments detailed in this document.

[0283] Referring to FIG. 13, the encoding device can derive a residual sample (S1300). The encoding device can derive a residual sample for the current block, and the residual sample for the current block can be derived based on the original sample and the prediction sample of the current block. Specifically, the encoding device can derive the prediction sample of the current block based on the prediction mode. In this case, various prediction methods disclosed in this document, such as inter prediction or intra prediction, can be applied. A residual sample can be derived based on the prediction sample and the original sample.

[0284] The encoding device can derive a transform coefficient (S1310). The encoding device can derive a transform coefficient based on a transform procedure for the residual sample. For example, the transform procedure can include at least one of DCT, DST, GBT, or CNT.

[0285] The encoding device can derive a quantized transform coefficient (S1320). The encoding device can derive a quantized transform coefficient based on a quantization procedure for the transform coefficient. The quantized transform coefficient can have a one-dimensional vector form based on the coefficient scan order.

[0286] The encoding device can generate residual information (S1330). The encoding device can generate residual information indicating the quantized transform coefficient. The residual information can be generated via various encoding methods such as exponential Golomb, CAVLC, CABAC, etc.

[0287] The encoding device can generate a restored sample (S1340). The encoding device can generate a restored sample based on the residual information. The restored sample can be generated by adding a residual sample based on the residual information to a predicted sample. Specifically, the encoding device can perform prediction (intra or inter prediction) for the current block and generate a restored sample based on the original sample and the predicted sample generated from the prediction.

[0288] The restored sample can include a restored luma sample and a restored chroma sample. Specifically, the residual sample can include a residual luma sample and a residual chroma sample. The residual luma sample can be generated based on the original luma sample and the predicted luma sample. The residual chroma sample can be generated based on the original chroma sample and the predicted chroma sample. The encoding device can derive a transform coefficient (luma transform coefficient) for the residual luma sample and / or a transform coefficient (chroma transform coefficient) for the residual chroma sample. The quantized transform coefficient can include a quantized luma transform coefficient and / or a quantized chroma transform coefficient.

[0289] The encoding device can generate ALF-related information and / or CCALF (CC-ALF)-related information for the restored sample (S1350). The encoding device can generate ALF-related information for the restored sample. The encoding device derives parameters related to ALF that can be applied for filtering the restored sample, and generates ALF-related information. For example, the ALF-related information can include the information related to ALF detailed in this document.

[0290] The encoding device can encode video / video information (S1360). The video information can include residual information, ALF-related information, and / or CCALF-related information. The encoded video / video information can be output in a bitstream form. The bitstream can be transmitted to a decoding device via a network or a storage medium.

[0291] In one example, the CCALF-related information includes a CCALF availability flag, a flag related to whether CCALF is available for a Cb (or Cr) color component, a Cb (or Cr) filter signal flag related to whether a cross-component filter for the Cb (or Cr) color component is signaled, information related to the number of cross-component filters for the Cb (or Cr) color component, information regarding the value of the cross-component filter coefficient for the Cb (or Cr) color component, information regarding the absolute value of the cross-component filter coefficient for the Cb (or Cr) color component, information regarding the sign of the cross-component filter coefficient for the Cb (or Cr) color component, and / or information related to whether a cross-component filter is applied to the current block of the Cb (or Cr) color component within information about the coding tree unit (coding tree unit syntax).

[0292] The video / video information can include various information according to the embodiments of this document. For example, the video / video information can include the information disclosed in at least one of Tables 1 to 30 described above.

[0293] In one embodiment, the video information can include header information and an adaptation parameter set (APS). The header information can include information related to an identifier of the APS including ALF data. For example, the cross-component filter coefficients can be derived based on the ALF data.

[0294] In one embodiment, the video information can include a sequence parameter set (SPS). The SPS can include a CCALF available flag related to whether the cross-component filtering is available.

[0295] In one embodiment, the SPS can include an ALF available flag (sps_alf_enabled_flag) related to whether ALF is available. Based on the determination that the value of the ALF available flag is 1, the SPS can include a CCALF available flag related to whether the cross-component filtering is available.

[0296] In one embodiment, the video information can include slice header information. The slice header information can include an ALF available flag (slice_alf_enabled_flag) related to whether ALF is available. Based on the determination that the value of the ALF available flag is 1, it can be determined whether the value of the CCALF available flag is 1. In one example, based on the determination that the value of the ALF available flag is 1, the CCALF is available for the slice.

[0297] In one embodiment, the header information (slice header information) can include a first flag related to whether CCALF can be used for the Cb color component of the filtered restored chroma samples, and a second flag related to whether CCALF can be used for the Cr color component of the filtered restored chroma samples. In other examples, based on the determination that the value of the ALF useable flag (slice_alf_enabled_flag) is 1, the header information (slice header information) can include a first flag related to whether CCALF can be used for the Cb color component of the filtered restored chroma samples, and a second flag related to whether CCALF can be used for the Cr color component of the filtered restored chroma samples.

[0298] In one embodiment, the video information can include adaptation parameter sets (APSs). In one example, the slice header information can include ID information of a first APS (information related to an identifier of a second APS) for deriving cross-component filter coefficients for the Cb color component of the filtered restored chroma samples. The slice header information can include ID information of a second APS (information related to an identifier of a second APS) for deriving cross-component filter coefficients for the Cr color component of the filtered restored chroma samples. In other examples, based on the determination that the value of the first flag is 1, the slice header information can include ID information of a first APS (information related to an identifier of a second APS) for deriving cross-component filter coefficients for the Cb color component. Based on the determination that the value of the second flag is 1, the slice header information can include ID information of a second APS (information related to an identifier of a second APS) for deriving cross-component filter coefficients for the Cr color component.

[0299] In one embodiment, the first ALF data included in the first APS can include a Cb filter signal flag related to whether a cross-component filter for the Cb color component is signaled. Based on the Cb filter signal flag, the first ALF data can include information related to the number of cross-component filters for the Cb color component. Based on the information related to the number of cross-component filters for the Cb color component, the first ALF data can include information regarding the absolute value of the cross-component filter coefficient for the Cb color component and information regarding the sign of the cross-component filter coefficient for the Cb color component. Based on the information regarding the absolute value of the cross-component filter coefficient for the Cb color component and the information regarding the sign of the cross-component filter coefficient for the Cb color component, the cross-component filter coefficient for the Cb color component can be derived.

[0300] In one embodiment, the information related to the number of cross-component filters for the Cb color component can be Golomb coded with an exponent of 0 (0 th EG).

[0301] In one embodiment, the second ALF data included in the second APS may include a Cr filter signal flag related to whether a cross-component filter for the Cr color component is signaled. Based on the Cr filter signal flag, the second ALF data may include information related to the number of cross-component filters for the Cr color component. Based on the information related to the number of cross-component filters for the Cr color component, the second ALF data may include information regarding the absolute value of the cross-component filter coefficient for the Cr color component and information regarding the sign of the cross-component filter coefficient for the Cr color component. Based on the information regarding the absolute value of the cross-component filter coefficient for the Cr color component and the information regarding the sign of the cross-component filter coefficient for the Cr color component, the cross-component filter coefficient for the Cr color component can be derived.

[0302] In one embodiment, the information related to the number of cross-component filters for the Cr color component may be Golomb coded with a zero-th order exponential (0 th EG).

[0303] In one embodiment, the video information may include information regarding a coding tree unit. The information regarding the coding tree unit may include information regarding whether a cross-component filter is applied to the current block of the Cb color component and / or information regarding whether a cross-component filter is applied to the current block of the Cr color component.

[0304] In one embodiment, the information regarding the coding tree unit can include information regarding the filter set index of the cross-component filter applied to the current block of the Cb color component, and / or information regarding the filter set index of the cross-component filter applied to the current block of the Cr color component.

[0305] FIGS. 15 and 16 schematically show an example of a video / video decoding method and related components according to an embodiment of this document.

[0306] The method disclosed in FIG. 15 can be executed by the decoding device disclosed in FIG. 3 or FIG. 16. Specifically, for example, S1500 in FIG. 15 can be executed by the entropy decoding unit 310 of the decoding device, S1510 can be executed by the addition unit 340 of the decoding device, and S1520 to S1550 can be executed by the filtering unit 350 of the decoding device. The method disclosed in FIG. 15 can include the embodiments detailed in this document.

[0307] Referring to FIG. 15, the decoding device can receive / acquire video / video information (S1500). The video / video information can include residual information. The decoding device can receive / acquire the video / video information via a bitstream. In one example, the video / video information can further include CCAL-related information. For example, the CCALF-related information can include a CCALF available flag, a flag related to whether CCALF is available for the Cb (or Cr) color component, a Cb (or Cr) filter signal flag related to whether a cross-component filter for the Cb (or Cr) color component is signaled, information related to the number of cross-component filters for the Cb (or Cr) color component, information regarding the absolute value of the cross-component filter coefficients for the Cb (or Cr) color component, information regarding the sign of the cross-component filter coefficients for the Cb (or Cr) color component, and / or information regarding whether a cross-component filter is applied to the current block of the Cb (or Cr) color component within the coding tree unit (coding tree unit syntax).

[0308] The video / video information can include various information according to the embodiments of this document. For example, the video / video information can include information disclosed in at least one of Tables 1 to 30 described above.

[0309] The decoding device can derive quantized transform coefficients. The decoding device can derive quantized transform coefficients based on the residual information. The quantized transform coefficients can have a one-dimensional vector form based on the coefficient scan order. The quantized transform coefficients can include quantized luma transform coefficients and / or quantized chroma transform coefficients.

[0310] The decoding device can derive conversion coefficients. The decoding device can derive conversion coefficients based on an inverse quantization procedure for the quantized conversion coefficients. The decoding device can derive luma conversion coefficients through inverse quantization based on the quantized luma conversion coefficients. The decoding device can derive chroma conversion coefficients through inverse quantization based on the quantized chroma conversion coefficients.

[0311] The decoding device can generate / derive residual samples. The decoding device can derive residual samples based on an inverse conversion procedure for the conversion coefficients. The decoding device can derive residual luma samples through an inverse conversion procedure based on the luma conversion coefficients. The decoding device can derive residual chroma samples through an inverse conversion procedure based on the chroma conversion coefficients.

[0312] The decoding device can generate / derive restored luma samples and / or restored chroma samples (S1510). The decoding device can generate restored luma samples and / or restored chroma samples based on the residual information. The decoding device can generate restored samples based on the residual information. The restored samples can include restored luma samples and / or restored chroma samples. The luma component of the restored samples can correspond to the restored luma samples, and the chroma component of the restored samples can correspond to the restored chroma samples. The decoding device can generate predicted luma samples and / or predicted chroma samples through a prediction procedure. The decoding device can generate restored luma samples based on the predicted luma samples and the residual luma samples. The decoding device can generate restored chroma samples based on the predicted chroma samples and the residual chroma samples.

[0313] The decoding device can derive ALF filter coefficients for the ALF procedure of the restored chroma samples (S1520). Additionally, the decoding device can derive ALF filter coefficients for the ALF procedure of the restored luma samples. The ALF filter coefficients can be derived based on the ALF parameters included in the ALF data within the APS.

[0314] The decoding device can generate filtered restored chroma samples (S1530). The decoding device can generate restored samples filtered based on the restored chroma samples and the ALF filter coefficients.

[0315] The decoding device can derive cross-component filter coefficients for the cross-component filtering (S1540). The cross-component filter coefficients can be derived based on the CCALF-related information within the ALF data included in the aforementioned APS, and the identifier (ID) information of the corresponding APS can be included in (signaled through) the slice header.

[0316] The decoding device can generate modified filtered reconstructed chroma samples (S1550). The decoding device can generate modified filtered reconstructed chroma samples based on the restored luma samples, the filtered reconstructed chroma samples, and the cross-component filter coefficients. In one example, the decoding device can derive the difference between two samples among the restored luma samples, and can multiply the difference by one of the filter coefficients among the cross-component filter coefficients. Based on the result of the multiplication and the filtered reconstructed chroma samples, the decoding device can generate the modified filtered reconstructed chroma samples. For example, the decoding device can generate the modified filtered reconstructed chroma samples based on the sum between the product and one of the samples of the filtered reconstructed chroma samples.

[0317] In one embodiment, the video information can include header information and an adaptation parameter set (APS). The header information can include information related to an identifier of the APS including ALF data. For example, the cross-component filter coefficients can be derived based on the ALF data.

[0318] In one embodiment, the video information can include a sequence parameter set (SPS). The SPS can include a CCALF available flag related to whether the cross-component filtering is available.

[0319] In one embodiment, the SPS may include an ALF availability flag (sps_alf_enabled_flag) associated with whether the ALF is available. Based on the determination that the value of the ALF availability flag is 1, the SPS may include a CCALF availability flag associated with whether the cross-component filtering is available.

[0320] In one embodiment, the video information may include slice header information. The slice header information may include an ALF availability flag (slice_alf_enabled_flag) associated with whether the ALF is available. Based on the determination that the value of the ALF availability flag is 1, it may be determined whether the value of the CCALF availability flag is 1. In one example, based on the determination that the value of the ALF availability flag is 1, the CCALF is available for the slice.

[0321] In one embodiment, the header information (slice header information) may include a first flag associated with whether the CCALF is available for the Cb color component of the filtered reconstructed chroma samples, and a second flag associated with whether the CCALF is available for the Cr color component of the filtered reconstructed chroma samples. In another example, based on the determination that the value of the ALF availability flag (slice_alf_enabled_flag) is 1, the header information (slice header information) may include a first flag associated with whether the CCALF is available for the Cb color component of the filtered reconstructed chroma samples, and a second flag associated with whether the CCALF is available for the Cr color component of the filtered reconstructed chroma samples.

[0322] In one embodiment, the video information may include adaptation parameter sets (APSs). In one example, the slice header information may include ID information of a first APS (information related to an identifier of a second APS) for deriving cross-component filter coefficients for the Cb color component of the filtered reconstructed chroma samples. The slice header information may include ID information of a second APS (information related to an identifier of a second APS) for deriving cross-component filter coefficients for the Cr color component of the filtered reconstructed chroma samples. In another example, based on the determination that the value of the first flag is 1, the slice header information may include ID information of a first APS (information related to an identifier of a second APS) for deriving cross-component filter coefficients for the Cb color component. Based on the determination that the value of the second flag is 1, the slice header information may include ID information of a second APS (information related to an identifier of a second APS) for deriving cross-component filter coefficients for the Cr color component.

[0323] In one embodiment, the first ALF data included in the first APS can include a Cb filter signal flag related to whether a cross-component filter for the Cb color component is signaled. Based on the Cb filter signal flag, the first ALF data can include information related to the number of cross-component filters for the Cb color component. Based on the information related to the number of cross-component filters for the Cb color component, the first ALF data can include information regarding the absolute value of the cross-component filter coefficient for the Cb color component and information regarding the sign of the cross-component filter coefficient for the Cb color component. Based on the information regarding the absolute value of the cross-component filter coefficient for the Cb color component and the information regarding the sign of the cross-component filter coefficient for the Cb color component, the cross-component filter coefficient for the Cb color component can be derived.

[0324] In one embodiment, the information related to the number of cross-component filters for the Cb color component can be th coded in zero-order exponential Golomb (0

[0325] In one embodiment, the second ALF data included in the second APS can include a Cr filter signal flag related to whether a cross-component filter for the Cr color component is signaled. Based on the Cr filter signal flag, the second ALF data can include information related to the number of cross-component filters for the Cr color component. Based on the information related to the number of cross-component filters for the Cr color component, the second ALF data can include information regarding the absolute value of the cross-component filter coefficient for the Cr color component and information regarding the sign of the cross-component filter coefficient for the Cr color component. Based on the information regarding the absolute value of the cross-component filter coefficient for the Cr color component and the information regarding the sign of the cross-component filter coefficient for the Cr color component, the cross-component filter coefficient for the Cr color component can be derived.

[0326] In one embodiment, the information related to the number of cross-component filters for the Cr color component can be Golomb coded with a zero order exponential (0 th EG).

[0327] In one embodiment, the video information can include information regarding a coding tree unit. The information regarding the coding tree unit can include information regarding whether a cross-component filter is applied to the current block of the Cb color component and / or information regarding whether a cross-component filter is applied to the current block of the Cr color component.

[0328] In one embodiment, the information regarding the coding tree unit can include information regarding the filter set index of the cross-component filter applied to the current block of the Cb color component and / or information regarding the filter set index of the cross-component filter applied to the current block of the Cr color component.

[0329] When there are residual samples for the current block, the decoding device can receive information regarding the residual for the current block. The information regarding the residual can include transform coefficients regarding the residual samples. The decoding device can derive the residual samples (or, residual sample array) for the current block based on the residual information. Specifically, the decoding device can derive the quantized transform coefficients based on the residual information. The quantized transform coefficients can have a one-dimensional vector form based on the coefficient scan order. The decoding device can derive the transform coefficients based on an inverse quantization procedure for the quantized transform coefficients. The decoding device can derive the residual samples based on the transform coefficients.

[0330] The decoding device can generate restored samples based on the (intra) prediction samples and the residual samples, and can derive a restored block or a restored picture based on the restored samples. Specifically, the decoding device can generate the restored samples based on the sum of the (intra) prediction samples and the residual samples. Thereafter, as described above, the decoding device can apply in-loop filtering procedures such as deblocking filtering and / or SAO procedures to the restored picture, if necessary, to improve the subjective / objective picture quality.

[0331] For example, the decoding device can decode a bitstream or encoded information and obtain video information including all or part of the aforementioned information (or syntax elements). Further, the bitstream or encoded information can be stored in a computer-readable storage medium and can be caused to perform the aforementioned decoding method.

[0332] In the foregoing embodiments, the method is described based on a flowchart as a series of steps or blocks, but the corresponding embodiments are not limited to the order of the steps, and a certain step can occur in a different order from the steps described above or simultaneously. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, and different steps can be included, or one or more steps of the flowchart can be deleted without affecting the scope of the embodiments of this document.

[0333] The method according to the embodiments of the foregoing document can be embodied in the form of software, and the encoding device and / or decoding device according to this document can be included in devices that perform video processing, such as TVs, computers, smartphones, set-top boxes, display devices, etc.

[0334] In this document, when an embodiment is implemented in software, the above-described method can be implemented by modules (processes, functions, etc.) that perform the above-described functions. The modules can be stored in a memory and executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by various well-known means. The processor can include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (read-only memory), a RAM (random access memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and performed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be implemented and performed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for implementation (e.g., information on instructions) or an algorithm can be stored in a digital storage medium.

[0335] In addition, the decoding device and the encoding device to which the embodiments of this document are applied may include a multimedia broadcast transceiver, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conferencing device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an over-the-top video (OTT) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a virtual reality (VR) device, an augmented reality (AR) device, a videophone video device, a transportation means terminal (e.g., a vehicle terminal including an autonomous vehicle terminal, an airplane terminal, a ship terminal, etc.) and a medical video device, etc., and may be used to process video signals or data signals. For example, as the over-the-top video (OTT) device, it may include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.

[0336] In addition, the processing method to which the embodiments of this document are applied can be produced in the form of a program executed by a computer and can be stored in a recording medium readable by the computer. Multimedia data having a data structure according to the embodiments of this document can also be stored in a recording medium readable by the computer. The recording medium readable by the computer includes all types of storage devices and distributed storage devices in which data readable by the computer is stored. The recording medium readable by the computer can include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. Also, the recording medium readable by the computer includes a medium embodied in the form of a carrier wave (for example, transmission via the Internet). Also, a bitstream generated by an encoding method can be stored in a recording medium readable by the computer or can be transmitted via a wired or wireless communication network.

[0337] In addition, the embodiments of this document can be embodied in a computer program product by program code, and the program code can be executed by a computer according to the embodiments of this document. The program code can be stored on a carrier readable by a computer.

[0338] FIG. 17 shows an example of a content streaming system to which the embodiments disclosed in this document can be applied.

[0339] Referring to FIG. 17, a content streaming system to which the embodiments of this document are applied can include a large encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0340] The encoding server compresses the content input from multimedia input devices such as smartphones, cameras, camcorders, etc. into digital data to generate a bitstream, and serves to transmit this to the streaming server. As another example, when a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted.

[0341] The bitstream can be generated by an encoding method or a method for generating a bitstream to which the embodiments of this document are applicable. The streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0342] The streaming server transmits multimedia data to a user device based on a user request via a web server. The web server serves as a medium to inform the user of what services are available. If the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system can include another control server. In this case, the control server serves to control commands / responses between each device within the content streaming system.

[0343] The streaming server can receive content from a media repository and / or an encoding server. For example, when it comes to receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0344] In the example of the user device, there may be a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (for example, a smartwatch, a smart glass, an HMD (head mounted display)), a digital TV, a desktop computer, a digital signage, and the like.

[0345] Each server in the content streaming system can be operated as a distributed server. In this case, the data received by each server can be distributedly processed.

[0346] The claims described in this specification can be combined in various ways. For example, the technical features of the method claims in this specification can be combined and implemented as an apparatus, and the technical features of the apparatus claims in this specification can be combined and implemented as a method. Also, the technical features of the method claims in this specification and the technical features of the apparatus claims can be combined and implemented as an apparatus, and the technical features of the method claims in this specification and the technical features of the apparatus claims can be combined and implemented as a method.

Claims

1. 1. A method for video decoding performed by a decoding device, comprising: obtaining video information including residual information via a bitstream; generating reconstructed luma samples and reconstructed chroma samples based on the residual information; deriving adaptive loop filter (ALF) filter coefficients for an ALF procedure on the reconstructed chroma samples; generating filtered reconstructed chroma samples based on the reconstructed chroma samples and the ALF filter coefficients; deriving cross-component filter coefficients for cross-component filtering; generating modified filtered reconstructed chroma samples based on the reconstructed luma samples, the filtered reconstructed chroma samples, and the cross-component filter coefficients; The video information includes a sequence parameter set (SPS), an adaptation parameter set (APS), and slice header information, The ALF data included in the APS includes flag information related to whether a cross-component filter is signaled; Based on the flag information, the ALF data includes information regarding the absolute value of the cross-component filter and information regarding the sign of the cross-component filter; The SPS includes an ALF enabled flag associated with whether the ALF procedure is enabled; Based on a determination that the ALF enable flag has a value of 1, the SPS includes a cross-component adaptive loop filter (CCALF) enable flag associated with whether the cross-component filtering is enabled; Based on a determination that the value of the ALF enabled flag in the SPS is 1, the slice header information includes an ALF enabled flag associated with whether the ALF is enabled; Based on a determination that the value of the ALF enabled flag included in the slice header information is 1 and the value of the CCALF enabled flag included in the SPS is 1, the slice header information includes information regarding whether the CCALF is enabled for the filtered reconstructed chroma samples; A method according to claim 1, wherein, based on the value of the information regarding whether the CCALF is available for the filtered reconstructed chroma sample being 1, the slice header information includes ID (identification) information of the APS associated with the CCALF for the filtered reconstructed chroma sample.

2. 1. A method of video encoding performed by an encoding device, comprising: deriving a residual sample for a current block; deriving transformation coefficients based on a transformation procedure for the residual samples; deriving quantized transform coefficients based on a quantization procedure for the transform coefficients; generating residual information associated with the quantized transform coefficients; generating a reconstruction sample based on the residual information; generating adaptive loop filter (ALF) related information and cross-component ALF (CCALF) related information for the reconstructed samples; encoding video information including the residual information, the ALF related information, and the CCALF related information; the reconstructed samples include reconstructed luma samples and reconstructed chroma samples; The video encoding method includes: deriving ALF filter coefficients for an ALF procedure on the reconstructed chroma samples; generating filtered reconstructed chroma samples based on the reconstructed chroma samples and the ALF filter coefficients; deriving cross-component filter coefficients for cross-component filtering; generating modified filtered reconstructed chroma samples based on the reconstructed luma samples, the filtered reconstructed chroma samples, and the cross-component filter coefficients; The video information includes a sequence parameter set (SPS), an adaptation parameter set (APS), and slice header information, The ALF data included in the APS includes flag information related to whether a cross-component filter is signaled; Based on the flag information, the ALF data includes information regarding the absolute value of the cross-component filter and information regarding the sign of the cross-component filter; The SPS includes an ALF enabled flag associated with whether the ALF procedure is enabled; Based on a determination that the ALF enable flag has a value of 1, the SPS includes a cross-component adaptive loop filter (CCALF) enable flag associated with whether the cross-component filtering is enabled; Based on a determination that the value of the ALF enabled flag in the SPS is 1, the slice header information includes an ALF enabled flag associated with whether the ALF is enabled; Based on the determination that the value of the ALF enabled flag included in the slice header information is 1 and the value of the CCALF enabled flag included in the SPS is 1, the slice header information includes information regarding whether the CCALF is enabled for the filtered reconstructed chroma samples; The method of claim 1, wherein based on the value of the information regarding whether the CCALF is available for the filtered reconstructed chroma sample being 1, the slice header information includes ID (identification) information of an adaptation parameter set (APS) associated with the CCALF for the filtered reconstructed chroma sample.

3. A method for transmitting data for video, comprising the steps of: generating a bitstream for the video; deriving a residual sample for a current block; deriving transformation coefficients based on a transformation procedure for the residual samples; deriving quantized transform coefficients based on a quantization procedure for the transform coefficients; generating residual information associated with the quantized transform coefficients; generating a reconstruction sample based on the residual information; generating adaptive loop filter (ALF) related information and cross-component ALF (CCALF) related information for the reconstructed samples; encoding video information including the residual information, the ALF related information, and the CCALF related information; transmitting the data including the bitstream; The video information includes a sequence parameter set (SPS), an adaptation parameter set (APS), and slice header information, The ALF data included in the APS includes flag information related to whether a cross-component filter is signaled; Based on the flag information, the ALF data includes information regarding the absolute value of the cross-component filter and information regarding the sign of the cross-component filter; The SPS includes an ALF enabled flag associated with whether the ALF procedure is enabled; Based on a determination that the ALF enable flag has a value of 1, the SPS includes a cross-component adaptive loop filter (CCALF) enable flag associated with whether cross-component filtering is enabled; Based on a determination that the value of the ALF enabled flag in the SPS is 1, the slice header information includes an ALF enabled flag associated with whether the ALF is enabled; Based on a determination that the value of the ALF enabled flag included in the slice header information is 1 and the value of the CCALF enabled flag included in the SPS is 1, the slice header information includes information regarding whether the CCALF is enabled for the filtered reconstructed chroma samples; The method of claim 1, wherein based on the value of the information regarding whether the CCALF is available for the filtered reconstructed chroma sample being 1, the slice header information includes ID (identification) information of an adaptation parameter set (APS) associated with the CCALF for the filtered reconstructed chroma sample.

Citation Information

Patent Citations

  • Image encoding device and image decoding device

    JP2021034980A

  • Apparatus and method for filtering-based video coding

    JP7656003B2

  • Cross-component filter

    US20180063527A1

  • Cross-component adaptive loop filter for chroma

    WO2021032751A1