Video coding based on scaling list data

By signaling scaling list data through an adaptation parameter set and hierarchically signaling availability flags, the method addresses inefficiencies in video/image coding for high-resolution media, enhancing compression efficiency and visual quality while managing memory usage.

JP2025142079AActive Publication Date: 2025-09-29LG ELECTRONICS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025120395
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-07-08
Filing Date
2025-07-17
Publication Date
2025-09-29
Estimated Expiration
2040-07-08

AI Technical Summary

Technical Problem

Existing video/image coding technologies face inefficiencies in compressing high-resolution, high-quality images/videos, especially for immersive media like VR and AR, leading to increased transmission and storage costs, and require improved methods for constructing and signaling scaling lists during the scaling process.

Method used

The method involves signaling scaling list data through an adaptation parameter set (APS) with specific APS ID information, using availability flag information to hierarchically signal scaling list data, and configuring SPS and PPS syntax to parse scaling list data individually in the APS, thereby improving coding efficiency.

Benefits of technology

This approach enhances overall image/video compression efficiency, improves subjective/objective visual quality, and facilitates efficient configuration and application of scaling lists, while limiting memory requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025142079000001_ABST
    Figure 2025142079000001_ABST
Patent Text Reader

Abstract

To provide a method and device for enhancing efficiency of video / image coding.SOLUTION: According to the disclosure of the present document, scaling list data delivered in an adaptation parameter set (APS) may be signaled through a hierarchical structure, and the amount of data that needs to be signaled for video / image coding may be reduced and implementation may be facilitated by placing limits on the scaling list data delivered in the APS.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present technology relates to video or image coding, for example, coding techniques based on scaling list data. [Background technology]

[0002] In recent years, demand for high-resolution, high-quality images / videos, such as 4K or 8K or higher UHD (Ultra High Definition) images / videos, has been increasing in various fields. As the resolution and quality of image / video data increases, the amount of information or bits to be transmitted increases relatively compared to existing image / video data. Therefore, when transmitting image data using existing media such as wired or wireless broadband lines or storing image / video data using existing storage media, transmission costs and storage costs increase.

[0003] In addition, interest in and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing in recent years, and the broadcast of images / videos with different image characteristics from real images, such as game images, has been increasing.

[0004] This necessitates the development of highly efficient image / video compression technologies to effectively compress and transmit, store, and play back high-resolution, high-quality image / video information that has the various characteristics described above.

[0005] Additionally, there is discussion about adaptive frequency weighting quantization techniques during the scaling process to improve compression efficiency and enhance subjective / objective visual quality. A method for signaling related information is needed to efficiently apply such techniques. Summary of the Invention [Problem to be solved by the invention]

[0006] The technical problem of this document is to provide a method and apparatus for increasing video / image coding efficiency.

[0007] Another technical problem of this document is to provide a method and apparatus for increasing coding efficiency during scaling.

[0008] Another technical problem of this document is to provide a method and apparatus for efficiently constructing a scaling list used in the scaling process.

[0009] Another technical problem of this document is to provide a method and apparatus for hierarchically signaling scaling list-related information used in the scaling process.

[0010] Another technical problem of this document is to provide a method and apparatus for efficiently applying a scaling list-based scaling process. [Means for solving the problem]

[0011] According to one embodiment of this document, scaling list data may be signaled via an adaptation parameter set (APS). Furthermore, an APS may be identified based on APS ID information included in the APS, and an APS may include scaling list data based on APS type information indicating that the APS is an APS related to the scaling list data included in the APS. Furthermore, for APS type information indicating that the APS is an APS related to the scaling list data, the value of the APS ID information may have a value within a specific range.

[0012] According to one embodiment of this document, availability flag information indicating the availability of scaling list data can be signaled hierarchically, and availability flag information in lower level syntax (e.g., picture header / slice header / type group header, etc.) can be signaled based on availability flag information signaled in higher level syntax (e.g., SPS).

[0013] According to one embodiment of this document, the SPS syntax and the PPS syntax can be configured so that scaling list data syntax is not directly signaled at the SPS or PPS level, thereby improving coding efficiency by parsing / signaling scaling list data individually in the APS based on availability flag information in the SPS.

[0014] According to an embodiment of the present document, there is provided a video / image decoding method performed by a decoding device, which may include the methods disclosed in the embodiments of the present document.

[0015] According to an embodiment of the present document, there is provided a decoding device for performing video / image decoding, which is capable of performing the method disclosed in the embodiment of the present document.

[0016] According to an embodiment of the present document, there is provided a video / image encoding method performed by an encoding device, which may include the methods disclosed in the embodiments of the present document.

[0017] According to an embodiment of the present document, an encoding device for video / image encoding is provided, which is capable of performing the method disclosed in the embodiment of the present document.

[0018] According to one embodiment of the present document, there is provided a computer-readable digital storage medium having encoded video / image information stored thereon, the encoded video / image information being generated by the video / image encoding method disclosed in at least one of the embodiments of the present document.

[0019] According to one embodiment of the present document, there is provided a computer-readable digital storage medium having stored thereon encoded information or encoded video / image information that causes a decoding device to perform the video / image decoding method disclosed in at least one of the embodiments of the present document. [Effects of the Invention]

[0020] This document may have various advantages. For example, according to an embodiment of this document, overall image / video compression efficiency may be improved. Furthermore, according to an embodiment of this document, coding efficiency may be improved by applying an efficient scaling process, thereby improving subjective / objective visual quality. Furthermore, according to an embodiment of this document, scaling lists used in the scaling process may be efficiently configured, thereby enabling scaling list-related information to be hierarchically signaled. Furthermore, according to an embodiment of this document, coding efficiency may be increased by efficiently applying a scaling process based on the scaling list. Furthermore, according to an embodiment of this document, restrictions on scaling list matrices used in the scaling process may be imposed, thereby facilitating implementation and limiting worst-case memory requirements.

[0021] The effects that can be obtained through specific embodiments of this document are not limited to the effects listed above. For example, there may be various technical effects that a person having ordinary skill in the related art can understand or derive from this document. Therefore, the specific effects of this document are not limited to those explicitly described in this document, but may include various effects that can be understood or derived from the technical features of this document. [Brief explanation of the drawings]

[0022] [Figure 1] 1 illustrates schematically an example of a video / image coding system to which embodiments of the present document may be applied. [Figure 2] 1 is a diagram illustrating a schematic configuration of a video / image encoding device to which embodiments of the present document can be applied; [Figure 3] 1 is a diagram illustrating the configuration of a video / image decoding device to which the embodiments of the present document can be applied; [Figure 4] 1 shows an example of a general video / image encoding method to which embodiments of this document can be applied. [Figure 5] 1 shows an example of a general video / image decoding method to which embodiments of this document can be applied. [Figure 6] 1 shows an exemplary hierarchical structure for coded images / videos. [Figure 7] 1 illustrates a schematic diagram of an example video / image encoding method and associated components according to an embodiment (or others) of the present document; [Figure 8] 1 illustrates a schematic diagram of an example video / image encoding method and associated components according to an embodiment (or others) of the present document; [Figure 9] 1 illustrates a schematic diagram of an example video / image decoding method and associated components according to an embodiment(s) of the present document; [Figure 10] 1 illustrates a schematic diagram of an example video / image decoding method and associated components according to an embodiment(s) of the present document; [Figure 11] 1 illustrates an example of a content streaming system to which the embodiments disclosed herein may be applied. DETAILED DESCRIPTION OF THE INVENTION

[0023] The present disclosure may be modified in various ways and may have various embodiments. Specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit the disclosure to the specific embodiments. Common terms used in this document are used merely to describe specific embodiments and are not intended to limit the technical ideas of the present disclosure. The singular expressions include the plural expressions unless the context clearly dictates otherwise. In this specification, the terms "comprise" or "have" are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0024] Meanwhile, the components in the drawings described in this disclosure are illustrated independently for the convenience of explaining the different characteristic functions, and do not mean that the components are realized by separate hardware or software. For example, two or more components may be combined into one component, or one component may be divided into multiple components. Embodiments in which the components are integrated and / or separated are also within the scope of the present disclosure as long as they do not deviate from the essence of this document.

[0025] In this document, "A or B" can mean "only A," "only B," or "both A and B." Also, in this document, "A or B" can be interpreted as "A and / or B." For example, in this document, "A, B or C" can mean "only A," "only B," "only C," or "any combination of A, B, and C."

[0026] A slash ( / ) or a comma (comma) used in this document can mean "and / or." For example, "A / B" can mean "A and / or B." Therefore, "A / B" can mean "only A," "only B," or "both A and B." For example, "A, B, C" can mean "A, B, or C."

[0027] In this document, "at least one of A and B" can mean "only A," "only B," or "both A and B." Also, in this document, the expressions "at least one of A or B" and "at least one of A and / or B" can be interpreted as "at least one of A and B."

[0028] Also, in this document, "at least one of A, B and C" can mean "only A," "only B," "only C," or "any combination of A, B and C." Also, "at least one of A, B or C" and "at least one of A, B and / or C" can mean "at least one of A, B and C."

[0029] Furthermore, parentheses used in this document may mean "for example." Specifically, when "prediction (intra prediction)" is used, "intra prediction" is proposed as an example of "prediction." In other words, "prediction" in this document is not limited to "intra prediction," and "intra prediction" is proposed as an example of "prediction." Furthermore, when "prediction (i.e., intra prediction)" is used, "intra prediction" is proposed as an example of "prediction."

[0030] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document may be applied to methods disclosed in the Versatile Video Coding (VVC) standard. The methods / embodiments disclosed in this document may also be applied to methods disclosed in the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the 2nd generation audio video coding standard (AVS2), or next generation video / image coding standards (e.g., H.267 or H.267).

[0031] Various embodiments relating to video / image coding are presented in this document, and unless otherwise stated, the embodiments may also be implemented in combination with each other.

[0032] In this document, video may refer to a collection of a series of images over time. A picture generally refers to a unit that represents one image at a specific time period, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile may include one or more coding tree units (CTUs). One picture may consist of one or more slices / tiles. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs, and the rectangular region has the same height as the picture, and the width can be specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the width of the picture.A tile scan may indicate a specific sequential ordering of CTUs partitioning a picture, in which the CTUs are ordered consecutively in a CTU raster scan in a tile, whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice includes an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that may be exclusively contained in a single NAL unit.

[0033] On the other hand, a picture can be divided into two or more sub-pictures, where a sub-picture is a rectangular region of one or more slices within a picture.

[0034] A pixel or a pel may refer to the smallest unit constituting a picture (or an image). A "sample" may also be used as a term corresponding to a pixel. A sample may generally refer to a pixel or a pixel value, or may refer to only a pixel / pixel value of a luma component, or may refer to only a pixel / pixel value of a chroma component. Alternatively, a sample may refer to a pixel value in the spatial domain, or, when such a pixel value is transformed into the frequency domain, may refer to a transform coefficient in the frequency domain.

[0035] A unit may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. The term unit may be used interchangeably with terms such as block or area. In general, an M×N block may include samples (or a sample array) consisting of M columns and N rows, or a set (or an array) of transform coefficients.

[0036] Also, in this document, at least one of quantization / dequantization and / or transform / inverse transform can be omitted. When quantization / dequantization is omitted, the quantized transform coefficients can be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients can also be referred to as coefficients or residual coefficients, or can still be referred to as transform coefficients for uniformity of expression.

[0037] In this document, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficient(s), and the information about the transform coefficient(s) may be signaled via residual coding syntax. Transform coefficients may be derived based on the residual information (or information about the transform coefficient(s)), and scaled transform coefficients may be derived through an inverse transform (scaling) of the transform coefficients. Residual samples may be derived based on an inverse transform (transform) of the scaled transform coefficients. This may be similarly applied / expressed in other parts of this document.

[0038] In this document, technical features individually described in one drawing may be realized individually or simultaneously.

[0039] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Hereinafter, the same reference numerals will be used for the same components in the drawings, and duplicated descriptions of the same components may be omitted.

[0040] FIG. 1 illustrates schematically an example of a video / image coding system to which embodiments of the present document can be applied.

[0041] 1, a video / image coding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data to the receiving device in a file or streaming format via a digital storage medium or a network.

[0042] The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may be a separate device or an external component.

[0043] A video source can acquire video / images through a video / image capture, synthesis, or generation process. A video source can include a video / image capture device and / or a video / image generation device. A video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, a virtual video / image can be generated via a computer, etc., in which case the video / image capture process can be replaced by a process in which related data is generated.

[0044] An encoding device can encode input video / images. The encoding device can perform a series of steps such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0045] The transmitter can transmit the encoded video / image information or data output in the form of a bitstream to a receiver of a receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter can include elements for generating a media file in a predetermined file format and elements for transmission via a broadcasting / communication network. The receiver can receive / extract the bitstream and transmit it to a decoding device.

[0046] The decoding device can decode the video / image by performing a series of steps such as inverse quantization, inverse transform, prediction, etc., which correspond to the operations of the encoding device.

[0047] The renderer can render the decoded video / image, and the rendered video / image can be displayed via the display unit.

[0048] 2 is a diagram for schematically illustrating the configuration of a video / image encoding device to which the embodiments of this document can be applied. Hereinafter, the encoding device may include an image encoding device and / or a video encoding device.

[0049] Referring to FIG. 2, the encoding device 200 may include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter predictor 221 and an intra predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. The image dividing unit 210, the predicting unit 220, the residual processing unit 230, the entropy encoding unit 240, the adding unit 250, and the filtering unit 260 may be configured by one or more hardware components (e.g., an encoder chipset or a processor) depending on the embodiment. Also, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.

[0050] The image division unit 210 may divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) using a QTBTTT (Quad-tree, binary-tree, ternary-tree) structure. For example, one coding unit may be divided into multiple coding units of deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and then the binary tree structure and / or ternary structure may be applied. Alternatively, the binary tree structure may be applied first. The coding procedure according to this document may be performed based on the final coding unit that is not further divided. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to image characteristics, or the coding unit may be recursively divided into coding units of lower depths as needed, and a coding unit of an optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be divided or partitioned from the final coding unit.The prediction unit is a unit of sample prediction, and the transform unit is a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0051] The term "unit" can be used interchangeably with terms such as "block" or "area." In general, an MxN block can refer to a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally refer to a pixel or pixel value, or can refer to only a pixel / pixel value of the luma component, or only a pixel / pixel value of the chroma component. A sample can also be used as a term corresponding to one pixel or pel of a picture (or image).

[0052] The encoding apparatus 200 may subtract a prediction signal (predicted block, prediction sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from an input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, a unit in the encoder 200 that subtracts the prediction signal (predicted block, prediction sample array) from the input image signal (original block, original sample array) may be referred to as a subtraction unit 231. The prediction unit may perform prediction on a current block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied in units of the current block or CU. The prediction unit may generate various information related to prediction, such as prediction mode information, and transmit the information to the entropy encoding unit 240, as will be described later in the description of each prediction mode. The prediction information can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0053] The intra prediction unit 222 may predict the current block by referring to samples in the current picture. The referenced samples may be located adjacent to or distant from the current block depending on the prediction mode. Prediction modes in intra prediction may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, DC mode and planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the granularity of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 222 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.

[0054] The inter prediction unit 221 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (such as L0 prediction, L1 prediction, or Bi prediction). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), or the like, and the reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter predictor 221 may configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes. For example, in the case of a skip mode or a merge mode, the inter predictor 221 may use motion information of neighboring blocks as motion information for the current block. In the case of the skip mode, unlike in the merge mode, a residual signal may not be transmitted.In the case of motion vector prediction (MVP) mode, the motion vector of the neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0055] The predictor 220 may generate a prediction signal based on various prediction methods, which will be described later. For example, the predictor may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as combined inter and intra prediction (CIIP). The predictor may also use an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content image / video coding, such as games, such as screen content coding (SCC). IBC basically performs prediction within a current picture, but may be performed similarly to inter prediction in deriving a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described herein. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, sample values ​​within a picture may be signaled based on information related to a palette table and a palette index.

[0056] The prediction signal generated by the prediction unit (including the inter prediction unit 221 and / or the intra prediction unit 222) may be used to generate a reconstructed signal or a residual signal. The transform unit 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, GBT refers to a transform obtained from a graph representing inter-pixel relationship information. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. The transform process may be applied to pixel blocks having the same square size or non-square blocks of variable size.

[0057] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantization unit 233 may rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 may perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit 240 may encode information required for video / image restoration (e.g., values ​​of syntax elements) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of network abstraction layer (NAL) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. In this document, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the above-described encoding procedure and included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcasting network and / or a communication network, and the digital storage medium can include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting the signal output from the entropy encoding unit 240 and / or a storage unit (not shown) for storing the signal can be configured as internal / external elements of the encoding device 200, or the transmitter can be included in the entropy encoding unit 240.

[0058] The quantized transform coefficients output from the quantizer 233 may be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) may be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantizer 234 and the inverse transformer 235. The adder 155 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter predictor 221 or the intra predictor 222. When there is no residual for the current block, such as when skip mode is applied, a predicted block may be used as the reconstructed block. The adder 250 may be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next current block in the current picture, or may be used for inter prediction of the next picture after filtering, as described below.

[0059] Meanwhile, luma mapping with chroma scaling (LMCS) can be applied during picture encoding and / or reconstruction.

[0060] The filtering unit 260 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 260 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filtering unit 260 may generate various information related to filtering and transmit it to the entropy encoding unit 240, as will be described later in the description of each filtering method. The filtering information may be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0061] The modified reconstructed picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. When inter prediction is applied through this, the encoding device can avoid prediction mismatch between the encoding device 100 and the decoding device, and can also improve coding efficiency.

[0062] The DPB of the memory 270 may store the modified reconstructed picture to be used as a reference picture in the inter predictor 221. The memory 270 may store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter predictor 221 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 270 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 222.

[0063] 3 is a diagram for explaining the configuration of a video / image decoding device to which the embodiments of this document can be applied. Hereinafter, the decoding device may include an image decoding device and / or a video decoding device.

[0064] Referring to FIG. 3, the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter predictor 331 and an intra predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. Depending on the embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be configured as a single hardware component (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) and may be configured as a digital storage medium. The hardware components may further include a memory 360 as an internal / external component.

[0065] When a bitstream including video / image information is input, the decoding device 300 can reconstruct an image corresponding to the process by which the video / image information was processed by the encoding device of FIG. 2. For example, the decoding device 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied by the encoding device. Therefore, the processing unit for decoding is, for example, a coding unit, and the coding unit can be divided from a coding tree unit or a maximal coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. The reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a playback device.

[0066] The decoding device 300 may receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal may be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 may parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The decoding device may decode pictures based on the information on the parameter sets and / or the general constraint information. Signaled / received information and / or syntax elements, which will be described later in this document, may be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 may decode information in a bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values ​​of syntax elements required for image restoration and quantized values ​​of transform coefficients related to residuals. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using information on the syntax element to be decoded, decode information on adjacent and current blocks, or information on symbols / bins decoded in previous steps, predicts the occurrence probability of bins based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values ​​of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method may update the context model using information on the decoded symbol / bin for the context model of the next symbol / bin.Among the information decoded by the entropy decoding unit 310, information related to prediction is provided to a prediction unit (inter prediction unit 332 and intra prediction unit 331), and residual values ​​entropy decoded by the entropy decoding unit 310, i.e., quantized transform coefficients and related parameter information, may be input to the residual processing unit 320. The residual processing unit 320 may derive a residual signal (residual block, residual sample, residual sample array). In addition, among the information decoded by the entropy decoding unit 310, information related to filtering may be provided to the filtering unit 350. Meanwhile, a receiving unit (not shown) that receives a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310. Meanwhile, the decoding device according to this document may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, the inverse transform unit 322, the addition unit 340, the filtering unit 350, the memory 360, the inter prediction unit 332, and the intra prediction unit 331.

[0067] The inverse quantization unit 321 may inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 321 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit 321 may perform inverse quantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.

[0068] The inverse transform unit 322 performs inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0069] The prediction unit may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on information about the prediction output from the entropy decoding unit 310, and may determine a specific intra / inter prediction mode.

[0070] The predictor 320 may generate a prediction signal based on various prediction methods, which will be described later. For example, the predictor may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as combined inter and intra prediction (CIIP). The predictor may also use an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content image / video coding, such as games, such as screen content coding (SCC). IBC basically performs prediction within a current picture, but may be performed similarly to inter prediction in that a reference block is derived within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, information regarding a palette table and a palette index may be included in the video / picture information and signaled.

[0071] The intra prediction unit 331 may predict a current block by referring to samples in a current picture. The referenced samples may be located adjacent to or distant from the current block depending on a prediction mode. Prediction modes in intra prediction may include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 may also determine a prediction mode to be applied to the current block using prediction modes applied to neighboring blocks.

[0072] The inter prediction unit 332 may derive a predicted block for the current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter prediction unit 332 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes, and the prediction information may include information indicating the inter prediction mode for the current block.

[0073] The adder 340 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to a predicted signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the current block, such as when a skip mode is applied, the predicted block may be used as the reconstructed block.

[0074] The adder 340 may be referred to as a reconstruction unit or a reconstruction block generator. The generated reconstruction signal may be used for intra prediction of a next block to be processed in the current picture, may be output after filtering as described below, or may be used for inter prediction of a next picture.

[0075] Meanwhile, LMCS (luma mapping with chroma scaling) can be applied during the picture decoding process.

[0076] The filtering unit 350 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 350 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may transmit the modified reconstructed picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.

[0077] The (modified) reconstructed picture stored in the DPB of the memory 360 may be used as a reference picture in the inter predictor 332. The memory 360 may store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter predictor 260 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 360 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 331.

[0078] In this document, the embodiments described for the filtering unit 260, inter prediction unit 221, and intra prediction unit 222 of the encoding device 200 can also be applied identically or correspondingly to the filtering unit 350, inter prediction unit 332, and intra prediction unit 331 of the decoding device 300, respectively.

[0079] As described above, prediction is performed to improve compression efficiency when performing video coding. Through this, a predicted block including predicted samples for a current block, which is a block to be coded, can be generated. Here, the predicted block includes predicted samples in the spatial domain (or pixel domain). The predicted block is derived in the same way by an encoding device and a decoding device. The encoding device can improve image coding efficiency by signaling to the decoding device information (residual information) regarding the residual between an original block and a predicted block, rather than the original sample values ​​of the original block themselves. The decoding device can derive a residual block including residual samples based on the residual information, combine the residual block with the predicted block to generate a reconstructed block including reconstructed samples, and generate a reconstructed picture including the reconstructed block.

[0080] The residual information may be generated through a transform and quantization procedure. For example, an encoding device may derive a residual block between an original block and a predicted block, perform a transform procedure on residual samples (residual sample array) included in the residual block to derive transform coefficients, and perform a quantization procedure on the transform coefficients to derive quantized transform coefficients, thereby signaling the related residual information to a decoding device (via a bitstream). Here, the residual information may include information such as value information, position information, transform technique, transform kernel, and quantization parameter of the quantized transform coefficients. The decoding device may derive residual samples (or residual blocks) by performing an inverse quantization / inverse transform procedure based on the residual information. The decoding device may generate a reconstructed picture based on the predicted block and the residual block. Furthermore, the encoding device may derive a residual block by inverse quantizing / inverse transforming the quantized transform coefficients for reference for inter-prediction of a future picture, and generate a reconstructed picture based on the residual block.

[0081] Intra prediction may refer to a prediction that generates a prediction sample for a current block based on a reference sample in a picture to which the current block belongs (hereinafter referred to as the current picture). When intra prediction is applied to the current block, neighboring reference samples used for intra prediction of the current block may be derived. The neighboring reference samples of the current block may include samples adjacent to the left boundary and bottom-left of the current block having a size of nW×nH, a total of 2×nH samples adjacent to the top boundary and top-right of the current block, and one sample adjacent to the top-left of the current block. Alternatively, the neighboring reference samples of the current block may include upper neighboring samples of multiple columns and left neighboring samples of multiple rows. In addition, the neighboring reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block of size nW x nH, a total of nW samples adjacent to the bottom boundary of the current block, and one sample adjacent to the bottom-right of the current block.

[0082] However, some of the neighboring reference samples of the current block may not yet be decoded or may not be available. In this case, the decoder may substitute the unavailable samples as available samples to construct neighboring reference samples to be used for prediction, or may construct neighboring reference samples to be used for prediction through interpolation of available samples.

[0083] When neighboring reference samples are derived, (i) a predicted sample can be derived based on an average or interpolation of neighboring reference samples of the current block, or (ii) a predicted sample can be derived based on a reference sample that exists in a specific (prediction) direction with respect to the predicted sample among the neighboring reference samples of the current block. Case (i) can be called a non-directional mode or a non-angular mode, and case (ii) can be called a directional mode or an angular mode.

[0084] In addition, a prediction sample may be generated by interpolating a first neighboring sample located in the prediction direction of the intra prediction mode of the current block and a second neighboring sample located in the opposite direction to the prediction direction based on the prediction sample of the current block among the neighboring reference samples. This case may be called linear interpolation intra prediction (LIP). Alternatively, a chroma prediction sample may be generated based on a luma sample using a linear model (LM). This case may be called an LM mode or a CCLM (chroma component LM) mode.

[0085] Alternatively, a temporary predicted sample of the current block may be derived based on the filtered neighboring reference samples, and the predicted sample of the current block may be derived by weighted summing at least one reference sample derived according to the intra prediction mode among existing neighboring reference samples, i.e., non-filtered neighboring reference samples, and the temporary predicted sample. The above case may be called Position Dependent Intra Prediction (PDPC).

[0086] In addition, intra-prediction coding can be performed by selecting a reference sample line with the highest prediction accuracy from among multiple adjacent reference sample lines of the current block, deriving a predicted sample using a reference sample located in the prediction direction of the corresponding line, and signaling the reference sample line used at this time to a decoding device. This case can be called multi-reference line intra-prediction or MRL-based intra-prediction.

[0087] In addition, the current block can be divided into vertical or horizontal sub-partitions, and intra prediction can be performed based on the same intra prediction mode, and neighboring reference samples can be derived and used in units of sub-partitions. That is, in this case, the intra prediction mode for the current block is applied to the sub-partitions in the same way, and neighboring reference samples can be derived and used in units of sub-partitions, thereby improving intra prediction performance in some cases. This prediction method can be called ISP (intra sub-partitions) based intra prediction.

[0088] The above-described intra prediction methods may be referred to as intra prediction types, distinguished from intra prediction modes. The intra prediction types may be referred to by various terms, such as intra prediction techniques or additional intra prediction modes. For example, the intra prediction types (or additional intra prediction modes, etc.) may include at least one of the above-described LIP, PDPC, MRL, and ISP. A general intra prediction method other than the specific intra prediction types, such as LIP, PDPC, MRL, and ISP, may be referred to as a normal intra prediction type. The normal intra prediction type may be generally applied when the above-described specific intra prediction types are not applied, and prediction may be performed based on the above-described intra prediction modes. Meanwhile, post-processing filtering may be performed on the derived prediction samples, if necessary.

[0089] Specifically, the intra prediction procedure may include an intra prediction mode / type determination step, a neighboring reference sample derivation step, and an intra prediction mode / type-based prediction sample derivation step. If necessary, a post-processing filtering step may be performed on the derived prediction samples.

[0090] When intra prediction is applied, the intra prediction mode applied to the current block may be determined using the intra prediction mode of a neighboring block. For example, the decoding device may select one of the most probable mode (MPM) candidates in an MPM (Most Probable Mode) list derived based on the intra prediction modes of neighboring blocks (e.g., left and / or upper neighboring blocks) of the current block and additional candidate modes based on the received MPM index, or may select one of the remaining intra prediction modes not included in the MPM candidates (and the planar mode) based on the remaining intra prediction mode information. The MPM list may be configured to include or not include the planar mode as a candidate. For example, if the MPM list includes the planar mode as a candidate, the MPM list may have six candidates, and if the MPM list does not include the planar mode as a candidate, the MPM list may have five candidates. If the MPM list does not include a planar mode as a candidate, a not planar flag (e.g., intra_luma_not_planar_flag) indicating that the intra prediction mode of the current block is not a planar mode may be signaled. For example, the MPM flag may be signaled first, and the MPM index and not planar flag may be signaled if the MPM flag has a value of 1. Also, the MPM index may be signaled if the not planar flag has a value of 1. Here, the reason why the MPM list is configured not to include a planar mode as a candidate is that, rather than the planar mode not being an MPM, the planar mode is always considered as an MPM, and therefore the flag (not planar flag) is signaled first to first confirm whether or not it is a planar mode.

[0091] For example, whether the intra prediction mode applied to the current block is among the MPM candidates (and planar mode) or among the remaining mode can be indicated based on an MPM flag (e.g., intra_luma_mpm_flag). A value of 1 for the MPM flag may indicate that the intra prediction mode for the current block is among the MPM candidates (and planar mode), and a value of 0 for the MPM flag may indicate that the intra prediction mode for the current block is not among the MPM candidates (and planar mode). A value of 0 for the not planar flag (e.g., intra_luma_not_planar_flag) may indicate that the intra prediction mode for the current block is planar mode, and a value of 1 for the not planar flag may indicate that the intra prediction mode for the current block is not planar mode. The MPM index may be signaled in the form of an mpm_idx or intra_luma_mpm_idx syntax element, and the remaining intra prediction mode information may be signaled in the form of a rem_intra_luma_pred_mode or intra_luma_mpm_remainder syntax element. For example, the remaining intra prediction mode information may index the remaining intra prediction modes not included in the MPM candidates (and planar modes) among all intra prediction modes in order of prediction mode number and point to one of them. The intra prediction mode may be an intra prediction mode for a luma component (sample). Hereinafter, the intra prediction mode information may include at least one of an MPM flag (e.g., intra_luma_mpm_flag), a not planar flag (e.g., intra_luma_not_planar_flag), an MPM index (e.g., mpm_idx or intra_luma_mpm_idx), and remaining intra prediction mode information (rem_intra_luma_pred_mode or intra_luma_mpm_remainder). In this document, the MPM list may be referred to by various terms, such as an MPM candidate list or candModeList.If MIP (matrix-based intra prediction) is applied to the current block, a separate mpm flag for MIP (ex. intra_mip_mpm_flag), mpm index (ex. intra_mip_mpm_idx), remaining intra prediction mode information (ex. intra_mip_mpm_remainder) may be signaled, and the not planar flag is not signaled.

[0092] In other words, in general, when an image is divided into blocks, a current block to be coded and neighboring blocks have similar image characteristics. Therefore, the current block and neighboring blocks are likely to have the same or similar intra prediction modes. Therefore, an encoder can use the intra prediction modes of neighboring blocks to encode the intra prediction mode of the current block.

[0093] For example, the encoder / decoder may construct an MPM (Most Probable Modes) list for the current block. The MPM list may also be referred to as an MPM candidate list. Here, MPM may refer to a mode used to improve coding efficiency by considering the similarity between the current block and neighboring blocks during intra-prediction mode coding. As described above, the MPM list may be configured to include or exclude the planar mode. For example, if the MPM list includes the planar mode, the number of candidates in the MPM list may be six. If the MPM list does not include the planar mode, the number of candidates in the MPM list may be five. The encoder / decoder may construct an MPM list including five or six MPMs.

[0094] To construct the MPM list, three types of modes may be considered: default intra modes, neighbor intra modes, and derived intra modes. For the neighbor intra modes, two neighboring blocks, i.e., a left neighboring block and an upper neighboring block, may be considered.

[0095] As described above, if the MPM list is configured not to include a planar mode, the planar mode is removed from the list, and the number of MPM list candidates can be set to five.

[0096] In addition, among the intra prediction modes, the non-directional mode (or non-angular mode) may include a DC mode based on the average of neighboring reference samples of the current block or a planar mode based on interpolation.

[0097] When inter prediction is applied, a prediction unit of an encoding / decoding device may perform inter prediction on a block-by-block basis to derive predicted samples. Inter prediction may refer to a prediction derived in a manner that is dependent on data elements (e.g., sample values ​​or motion information) of picture(s) other than the current picture. When inter prediction is applied to a current block, a predicted block (prediction sample array) for the current block may be derived based on a reference block (reference sample array) identified by a motion vector on a reference picture indicated by a reference picture index. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, motion information of the current block may be predicted on a block-by-block, sub-block-by-sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include inter prediction type information (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). When inter-prediction is applied, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring blocks may be the same or different. The temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture including the temporal neighboring blocks may be called collocated pictures (colPics).For example, a motion information candidate list may be constructed based on neighboring blocks of the current block, and flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled. Inter prediction may be performed based on various prediction modes. For example, in skip mode and merge mode, the motion information of the current block is the same as the motion information of the selected neighboring block. Unlike merge mode, in skip mode, no residual signal is transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected neighboring block is used as a motion vector predictor, and a motion vector difference may be signaled. In this case, the motion vector of the current block may be derived using the sum of the motion vector predictor and the motion vector difference.

[0098] The motion information may include L0 motion information and / or L1 motion information depending on the inter-prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction may be referred to as an L0 motion vector or MVL0, and a motion vector in the L1 direction may be referred to as an L1 motion vector or MVL1. Prediction based on an L0 motion vector may be referred to as L0 prediction, prediction based on an L1 motion vector may be referred to as L1 prediction, and prediction based on both an L0 motion vector and an L1 motion vector may be referred to as bi-prediction. Here, the L0 motion vector may indicate a motion vector associated with the reference picture list L0 (L0), and the L1 motion vector may indicate a motion vector associated with the reference picture list L1 (L1). The reference picture list L0 may include pictures that are earlier in output order than the current picture as reference pictures, and the reference picture list L1 may include pictures that are later in output order than the current picture. The earlier picture may be referred to as a forward (reference) picture, and the later picture may be referred to as a backward (reference) picture. The reference picture list L0 may further include, as reference pictures, pictures that are later in output order than the current picture. In this case, the previous picture may be indexed first in the reference picture list L0, and the later picture may be indexed next. The reference picture list L1 may further include, as reference pictures, pictures that are earlier in output order than the current picture. In this case, the later picture may be indexed first in the reference picture list L1, and the previous picture may be indexed next. Here, the output order may correspond to a picture order count (POC) order.

[0099] FIG. 4 shows an example of a general video / image encoding method to which embodiments of this document can be applied.

[0100] The method disclosed in Fig. 4 may be performed by the encoding device 200 of Fig. 2. Specifically, S400 may be performed by the inter prediction unit 221 or the intra prediction unit 222 of the encoding device 200, and S410, S420, S430, and S440 may be performed by the subtraction unit 231, the transformation unit 232, the quantization unit 233, and the entropy encoding unit 240 of the encoding device 200, respectively.

[0101] 4, the encoding device may derive a prediction sample through prediction of a current block (S400). The encoding device may determine whether to perform inter prediction or intra prediction on the current block, and may determine a specific inter prediction mode or a specific intra prediction mode based on the RD cost. The encoding device may derive a prediction sample for the current block according to the determined mode.

[0102] The encoding apparatus can derive residual samples by comparing the original samples and predicted samples for the current block (S410).

[0103] The encoding device may derive transform coefficients through a transform procedure on the residual samples (S420), and quantize the derived transform coefficients to derive quantized transform coefficients (S430).

[0104] The encoding device may encode image information including prediction information and residual information and output the encoded image information in the form of a bitstream (S440). The prediction information is information related to a prediction procedure and may include prediction mode information and information related to motion information (e.g., when inter-prediction is applied), etc. The residual information may include information related to quantized transform coefficients. The residual information may be entropy coded.

[0105] The output bitstream can be transmitted to a decoding device via a storage medium or a network.

[0106] FIG. 5 shows an example of a general video / image decoding method to which the embodiments of this document can be applied.

[0107] The method disclosed in Fig. 5 may be performed by the decoding apparatus 300 of Fig. 3 described above. Specifically, S500 may be performed by the inter prediction unit 332 or the intra prediction unit 331 of the decoding apparatus 300. The procedure of decoding prediction information included in the bitstream and deriving values ​​of related syntax elements in S500 may be performed by the entropy decoding unit 310 of the decoding apparatus 300. S510, S520, S530, and S540 may be performed by the entropy decoding unit 310, the inverse quantization unit 321, the inverse transform unit 322, and the adder 340 of the decoding apparatus 300, respectively.

[0108] As shown in Figure 5, the decoding device may perform operations corresponding to those performed by the encoding device. The decoding device may perform inter-prediction or intra-prediction on a current block based on received prediction information and derive a prediction sample (S500).

[0109] The decoding device may derive quantized transform coefficients for the current block based on the received residual information (S510). The decoding device may derive the quantized transform coefficients from the residual information through entropy decoding.

[0110] The decoding device can dequantize the quantized transform coefficients to derive the transform coefficients (S520).

[0111] The decoding device derives residual samples through an inverse transform procedure on the transform coefficients (S530).

[0112] The decoding device generates a reconstructed sample for the current block based on the predicted sample and the residual sample, and generates a reconstructed picture based on the reconstructed sample (S540). Thereafter, the in-loop filtering procedure may be further applied to the reconstructed picture, as described above.

[0113] Meanwhile, as described above, the quantization unit of the encoding device can apply quantization to the transform coefficients to derive the quantized transform coefficients, and the inverse quantization unit of the encoding device or the inverse quantization unit of the decoding device can apply inverse quantization to the quantized transform coefficients to derive the transform coefficients.

[0114] Generally, in video / image coding, the quantization rate can be changed, and the compression can be adjusted using the changed quantization rate. In terms of implementation, a quantization parameter (QP) can be used instead of directly using the quantization rate, considering the complexity. For example, an integer value of the quantization parameter from 0 to 63 can be used, and each quantization parameter value can correspond to an actual quantization rate. The quantization parameter (QP) for the luma component (luma sample) Y ) and the quantization parameter (QP C ) can be set differently.

[0115] The quantization process takes the transform coefficients (C) as input and the quantization rate (Q step ) and obtain the quantized transform coefficients (C`) based on this. At this time, considering the calculation complexity, the quantization rate is multiplied by the scale to convert it into an integer form, and a shift operation can be performed by the value corresponding to the scale value. The quantization scale can be derived by multiplying the quantization rate by the scale value. In other words, the quantization scale can be derived according to the QP. The quantization scale can also be applied to the transform coefficients (C), and the quantized transform coefficients (C`) can be derived based on this.

[0116] The inverse quantization process applies a quantization rate (Q step ) and based on this, the reconstructed transform coefficients (C``) can be obtained. In this case, a level scale can be derived depending on the quantization parameter, and the level scale can be applied to the quantized transform coefficients (C`) to derive the reconstructed transform coefficients (C``). The reconstructed transform coefficients (C``) may differ slightly from the original transform coefficients (C) due to losses in the transform and / or quantization process. Therefore, the encoding device also performs inverse quantization, just like the decoding device.

[0117] In addition, an adaptive frequency weighting quantization technique that adjusts quantization strength according to frequency may be applied. The adaptive frequency weighting quantization technique is a method of applying quantization strength differently to different frequencies. The adaptive frequency weighting can apply quantization strength differently to each frequency using a predefined quantization scaling matrix. That is, the above-mentioned quantization / dequantization process may be further performed based on the quantization scaling matrix. For example, different quantization scaling matrices may be used depending on the size of the current block and / or whether the prediction mode applied to the current block is inter prediction or intra prediction to generate a residual signal of the current block. The quantization scaling matrix may be referred to as a quantization matrix or a scaling matrix. The quantization scaling matrix may be predefined. Furthermore, for frequency adaptive scaling, frequency-specific quantization scale information for the quantization scaling matrix may be configured / encoded by an encoding device and signaled to a decoding device. The frequency-specific quantization scale information may be referred to as quantization scaling information. The frequency-specific quantization scale information may include scaling list data (scaling_list_data). A (modified) quantization scaling matrix may be derived based on the scaling list data. The frequency-specific quantization scale information may also include present flag information indicating whether scaling list data is present. Alternatively, if scaling list data is signaled at a higher level (e.g., SPS), the frequency-specific quantization scale information may further include information indicating whether the scaling list data is modified at a lower level (e.g., PPS or tile group header, etc.).

[0118] As mentioned above, scaling list data can be signaled to indicate the scaling matrix used for quantization / dequantization (frequency-based quantization).

[0119] Support for signaling default and user-defined scaling matrices exists in the HEVC standard and has now been adopted for the VVC standard. However, in the case of the VVC standard, additional support for signaling the following features has been integrated:

[0120] -Three modes for scaling matrix: OFF, DEFAULT, USER_DEFINED

[0121] -Larger size range for blocks (4x4 to 64x64 for luma, 2x2 to 32x32 for chroma)

[0122] -Triangle Transformation Blocks (TBs)

[0123] -Dependent quantization

[0124] -Multiple Transform Selection (MTS)

[0125] -Large transforms with zeroing-out high frequency coefficients

[0126] -Intra sub-block partitioning (ISP)

[0127] - Intra Block Copy (IBC) (also known as current picture referencing (CPR))

[0128] -DEFAULT scaling matrix for all TB sizes, default value is 16

[0129] It should be noted that the scaling matrix should not be applied to the Transform Skip (TS) and Secondary Transform (ST) for all sizes.

[0130] The High Level Syntax (HSL) structure for supporting scaling lists in the VVC standard will be described in detail below. First, a flag can be signaled via a Sequence Parameter Set (SPS) to indicate that a scaling list is available for the current coded video sequence (CVS) being decoded. Next, if the flag is available, an additional flag can be parsed in the SPS to indicate whether specific data exists in the scaling list. This can be shown in Table 1.

[0131] Table 1 is an excerpt from SPS to illustrate the scaling list for CVS.

[0132] [Table 1]

[0133] The semantics of the syntax elements included in the SPS syntax of Table 1 can be shown as in Table 2 below.

[0134] [Table 2]

[0135] Referring to Tables 1 and 2, scaling_list_enabled_flag can be signaled from the SPS. For example, if the value of scaling_list_enabled_flag is 1, it can indicate that a scaling list is used in the scaling process for the transform coefficients, and if the value of scaling_list_enabled_flag is 0, it can indicate that a scaling list is not used in the scaling process for the transform coefficients. In this case, if the value of scaling_list_enabled_flag is 1, sps_scaling_list_data_present_flag can be further signaled from the SPS. For example, if the value of sps_scaling_list_data_present_flag is 1, it can indicate that the scaling_list_data() syntax structure is present in the SPS, and if the value of sps_scaling_list_data_present_flag is 0, it can indicate that the scaling_list_data() syntax structure is not present in the SPS. If sps_scaling_list_data_present_flag is not present, the value of sps_scaling_list_data_present_flag can be inferred to be 0.

[0136] Also, a flag (e.g., pps_scaling_list_data_present_flag) can be parsed first in the Picture Parameter Set (PPS). If this flag is available, scaling_list_data() can be parsed in the PPS. If scaling_list_data() exists first in the SPS and is parsed later in the PPS, the data in the PPS can take precedence over the data in the SPS. Table 3 below is an excerpt from the PPS to explain the scaling list data.

[0137] [Table 3]

[0138] The semantics of the syntax elements included in the PPS syntax of Table 3 can be shown as in Table 4 below.

[0139] [Table 4]

[0140] Referring to Tables 3 and 4, pps_scaling_list_data_present_flag can be signaled from the PPS. For example, if the value of pps_scaling_list_data_present_flag is 1, it can indicate that the scaling list data used for a picture referencing the PPS is derived based on the scaling list specified by the active SPS and the scaling list specified by the PPS. If the value of pps_scaling_list_data_present_flag is 0, it can indicate that the scaling list data used for a picture referencing the PPS is inferred to be the same as the scaling list specified by the active SPS. In this case, if the value of scaling_list_enabled_flag is 0, the value of pps_scaling_list_data_present_flag must be 0. If the value of scaling_list_enabled_flag is 1, the value of sps_scaling_list_data_present_flag is 0, and the value of pps_scaling_list_data_present_flag is 0, the default scaling list data can be used to derive the scaling factor array as described in Scaling List Data Semantics.

[0141] Scaling lists can be defined in the VVC standard for the following quantization matrix sizes, as shown in Table 5 below: The supported range for quantization matrices has been extended to include 2x2 and 64x64 in the HEVC standard, with 4x4, 8x8, 16x16, and 32x32.

[0142] [Table 5]

[0143] Table 5 defines sizeIds for all quantization matrix sizes used. Using the above combinations, matrixIds can be assigned to different combinations of sizeId, coding unit prediction mode (CuPredMode), and color components. Here, CuPredModes that can be considered are inter, intra, and IBC (Intra Block Copy). Intra mode and IBC mode can be treated equally. Therefore, the same matrixId(s) can be shared for a given color component. Here, color components that can be considered are luma (Y) and two color components (Cb and Cr). The assigned matrixIds can be shown as in Table 6 below.

[0144] Table 6 shows the matrixId by sizeId, prediction mode, and color component.

[0145] [Table 6]

[0146] Table 7 below shows an example syntax structure for scaling list data (e.g., scaling_list_data()).

[0147] [Table 7]

[0148] The semantics of the syntax elements included in the syntax of Table 7 can be shown as in Table 8 below.

[0149] [Table 8-1]

[0150] [Table 8-2]

[0151] Referring to Tables 7 and 8, to extract the scaling list data (e.g., scaling_list_data()), for all sizeIds from 1 to 6 and matrixIds from 0 to 5, the scaling list data can be applied to 2x2 chroma components and 64x64 luma components. Next, a flag (e.g., scaling_list_pred_mode_flag) can be parsed to indicate whether the values ​​of the scaling list are the same as the values ​​of the reference scaling list. The reference scaling list is indicated by scaling_list_pred_matrix_id_delta[sizeId][matrixId]. However, if scaling_list_pred_mode_flag[sizeId][matrixId] is 1, the scaling list data can be explicitly signaled. If scaling_list_pred_matrix_id_delta is 0, the DEFAULT mode with default values ​​can be used, as shown in Tables 9 to 12. For other values ​​of scaling_list_pred_matrix_id_delta, refMatrixId can be determined first as shown in the semantics of Table 8 above.

[0152] In explicit signaling, i.e., in USER_DEFINED mode, the maximum number of coefficients to be signaled can be determined in advance. For quantization block sizes of 2x2, 4x4, and 8x8, all coefficients can be signaled. For sizes larger than 8x8, i.e., 16x16, 32x32, and 64x64, only 64 coefficients can be signaled. That is, an 8x8 base matrix is ​​signaled, and the remaining coefficients can be upsampled from the base matrix.

[0153] Table 9 below is an example showing the default values ​​of ScalingList[1][matrixId][i] (i=0..3).

[0154] [Table 9]

[0155] Table 10 below is an example showing the default values ​​of ScalingList[2][matrixId][i] (i=0..15).

[0156] [Table 10]

[0157] Table 11 below is an example showing the default values ​​for ScalingList[3..5][matrixId][i] (i=0..63).

[0158] [Table 11]

[0159] Table 12 below is an example showing the default values ​​of ScalingList[6][matrixId][i] (i=0..63).

[0160] [Table 12]

[0161] As described above, the default scaling list data can be used to derive a scaling factor (Scaling Factor).

[0162] The scaling factor ScalingFactor[sizeId][sizeId][matrixId][x][y] of a five-dimensional array (where x, y = 0..(1<<sizeId)-1) can represent an array of scaling factors according to the variable sizeId shown in Table 5 above and the variable matrixId shown in Table 6 above.

[0163] The following Table 13 shows an example of deriving a scaling factor based on the quantization matrix size according to the aforementioned default scaling list.

[0164]

Table 13-1

[0165]

Table 13-2

[0166] For a quantization matrix of rectangular size, the scaling factor ScalingFactor[sizeIdW][sizeIdH][matrixId][x][y] of a five-dimensional array (where x = 0..(1<<sizeIdW)-1, y = 0..(1<<sizeIdH)-1, sizeIdW!=sizeIdH) can represent an array of scaling factors according to the variables sizeIdW and sizeIdH shown in Table 15 below and can be derived as shown in Table 14 below.

[0167]

Table 14

[0168] The quantization matrix of the quadrilateral size must be zero for samples that meet the following conditions.

[0169] -x > 32

[0170] -y > 32

[0171] - The decoded TU is not coded in the default transform mode, (1 << sizeIdW) == 32 and x > 16

[0172] - The decoded TU is not coded in the default transform mode, (1 << sizeIdH) == 32 and y > 16

[0173] The following Table 15 is an example showing sizeIdW and sizeIdH according to the quantization matrix size.

[0174]

Table 15

[0175] Also, as an example, the above-mentioned scaling list data (e.g., scaling_list_data()) can be described based on the syntax structure as shown in the following Table 16 and the semantics as shown in the following Table 17. Based on the syntax elements included in the scaling list data (e.g., scaling_list_data()) disclosed in Table 16 and Table 17, as described above, the scaling list, the scaling matrix, the scaling factor, etc. can be derived, and this process can be the same as or similar procedures applied to the above-mentioned Tables 5 to 15.

[0176]

Table 16

[0177]

Table 17-1

[0178] [Table 17-2]

[0179] In the following, this document proposes a method for efficiently signaling scaling list data when applying adaptive frequency-specific weighting techniques in the quantization / dequantization process.

[0180] FIG. 6 shows an exemplary hierarchical structure for coded images / videos.

[0181] As shown in Figure 6, coded images / videos are divided into a VCL (video coding layer) that handles the image / video decoding process and itself, a lower system that transmits and stores the coded information, and a NAL (network abstraction layer) that exists between the VCL and the lower system and is responsible for network adaptation functions.

[0182] The VCL can generate VCL data containing compressed image data (slice data), or parameter sets containing information such as a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), and a Video Parameter Set (VPS), or SEI (Supplemental Enhancement Information) messages that are additionally required for the image decoding process.

[0183] In NAL, NAL units can be generated by adding header information (NAL unit header) to RBSP (Raw Byte Sequence Payload) generated by VCL. In this case, RBSP refers to slice data, parameter sets, SEI messages, etc. generated by VCL. The NAL unit header can include NAL unit type information identified by the RBSP data included in the NAL unit.

[0184] In addition, NAL units can be classified into VCL NAL units and non-VCL NAL units according to the RBSP generated by the VCL. A VCL NAL unit can refer to a NAL unit containing information about an image (slice data), and a non-VCL NAL unit can refer to a NAL unit containing information necessary for decoding an image (parameter set or SEI message).

[0185] VCL NAL units and non-VCL NAL units can be transmitted over a network with header information according to the data standard of the lower system. For example, NAL units can be transformed into a data format of a predetermined standard, such as H.266 / VVC file format, Real-time Transport Protocol (RTP), Transport Stream (TS), etc., and transmitted over various networks.

[0186] As mentioned above, the NAL unit type of an NAL unit can be identified by the RBSP data structure included in the NAL unit, and information about such NAL unit type can be stored and signaled in the NAL unit header.

[0187] For example, NAL units can be broadly classified into VCL NAL unit types and non-VCL NAL unit types depending on whether the NAL unit contains information about an image (slice data). The VCL NAL unit types can be classified according to the nature and type of pictures included in the VCL NAL unit, and the non-VCL NAL unit types can be classified according to the type of parameter set.

[0188] The following is an example of a NAL unit type identified by the type of parameter set included in the non-VCL NAL unit type.

[0189] -APS (Adaptation Parameter Set) NAL unit: Type for NAL units including APS

[0190] -DPS (Decoding Parameter Set) NAL unit: Type for NAL units containing DPS

[0191] -VPS (Video Parameter Set) NAL unit: Type for NAL unit containing VPS

[0192] -SPS (Sequence Parameter Set) NAL unit: Type for NAL unit including SPS

[0193] -PPS (Picture Parameter Set) NAL unit: Type for NAL unit including PPS

[0194] -PH (Picture header) NAL unit: Type for NAL unit including PH

[0195] The above-mentioned NAL unit type has syntax information for the NAL unit type, and the syntax information can be stored in a NAL unit header and signaled. For example, the syntax information is nal_unit_type, and the NAL unit type can be specified by the nal_unit_type value.

[0196] Meanwhile, as described above, one picture may include multiple slices, and one slice may include a slice header and slice data. In this case, one picture header may be added to multiple slices (slice header and slice data set) in one picture. The picture header (picture header syntax) may include information / parameters commonly applicable to pictures. In this document, the term "tile group" may be used interchangeably with or instead of a slice or picture. In addition, in this document, the term "tile group header" may be used interchangeably with or instead of a slice header or picture header.

[0197] A slice header (slice header syntax) can contain information / parameters commonly applicable to slices. An APS (APS syntax) or PPS (PPS syntax) can contain information / parameters commonly applicable to one or more slices or pictures. An SPS (SPS syntax) can contain information / parameters commonly applicable to one or more sequences. A VPS (VPS syntax) can contain information / parameters commonly applicable to multiple layers. A DPS (DPS syntax) can contain information / parameters commonly applicable to video in general. A DPS can contain information / parameters related to the concatenation of coded video sequences (CVSs). In this document, the term "High level syntax (HLS)" refers to at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, and slice header syntax.

[0198] In this document, image / video information encoded from an encoding device to a decoding device and signaled in the form of a bitstream may include not only intra-picture partitioning-related information, intra / inter prediction information, residual information, in-loop filtering information, etc., but also information included in a slice header, information included in a picture header, information included in an APS, information included in a PPS, information included in an SPS, information included in a VPS, and / or information included in a DPS. In addition, the image / video information may further include information of a NAL unit header.

[0199] Meanwhile, the Adaptation Parameter Set (APS) is used in the VVC standard to transmit information for the Adaptive Loop Filter (ALF) and Luma Mapping with Chroma Scaling (LMCS) procedures. The APS also has an extensible structure so that it can be used to transmit other data structures (i.e., other syntax structures). Therefore, this document proposes a method for parsing / signaling scaling list data used for frequency-specific weighting via the APS.

[0200] As described above, the scaling list data is quantization scale information for frequency-specific weighting that can be applied in the process of quantization / dequantization, and is a list that associates scale factors with each frequency index.

[0201] In one embodiment, Table 18 below shows an example of an adaptation parameter set (APS) structure used to transmit scaling list data.

[0202] [Table 18]

[0203] The semantics of the syntax elements included in the APS syntax of Table 18 can be shown as in Table 19 below.

[0204] [Table 19]

[0205] Referring to Tables 18 and 19, the adaptation_parameter_set_id syntax element can be parsed / signaled in an APS. The adaptation_parameter_set_id provides an identifier for the APS for reference by other syntax elements. That is, the APS can be identified based on the adaptation_parameter_set_id syntax element. The adaptation_parameter_set_id syntax element can be referred to as APS ID information. The APS can be shared between pictures and may be different for different tile groups within a picture.

[0206] In addition, the aps_params_type syntax element can be parsed / signaled in the APS. The aps_params_type can indicate the type of APS parameters transmitted in the APS, as shown in Table 18 below. The aps_params_type syntax element can be referred to as APS parameter type information or APS type information.

[0207] For example, Table 20 below is an example showing the types of APS parameters that can be transmitted via APS, and each APS parameter type can be indicated corresponding to the value of aps_params_type.

[0208] [Table 20]

[0209] Referring to Table 20, aps_params_type is a syntax element for classifying the type of the APS. If the value of aps_params_type is 0, the APS type is ALF_APS, and the APS can carry ALF data, which can include ALF parameters for deriving filters / filter coefficients. If the value of aps_params_type is 1, the APS type is LMCS_APS, and the APS can carry LMCS data, which can include LMCS parameters for deriving LMCS models / bin / mapping indices. If the value of aps_params_type is 2, the APS type is SCALING_APS, and the APS can carry SCALING list data, which can include scaling list data parameters for deriving frequency-based quantization scaling matrices / scaling factors / scaling list values.

[0210] For example, as shown in Table 18, the aps_params_type syntax element can be parsed / signaled in APS. In this case, if the value of aps_params_type indicates 0 (i.e., if aps_params_type indicates ALF_APS), ALF data (i.e., alf_data()) can be parsed / signaled. Alternatively, if the value of aps_params_type indicates 1 (i.e., if aps_params_type indicates LMCS_APS), LMCS data (i.e., lmcs_data()) can be parsed / signaled. Alternatively, if the value of aps_params_type indicates 2 (i.e., if aps_params_type indicates SCALING_APS), scaling list data (i.e., scaling_list_data()) can be parsed / signaled.

[0211] Also, referring to Tables 18 and 19, the aps_extension_flag syntax element can be parsed / signaled in APS. aps_extension_flag can indicate whether the APS extension data flag (aps_extension_data_flag) syntax element is present. aps_extension_flag can be used, for example, to provide an extension point for a later version of the VVC standard. The aps_extension_flag syntax element can be referred to as an APS extension flag. For example, a value of aps_extension_flag of 0 can indicate that the APS extension data flag (aps_extension_data_flag) is not present in the APS RBSP syntax structure. Alternatively, a value of aps_extension_flag of 1 can indicate that the APS extension data flag (aps_extension_data_flag) is present in the APS RBSP syntax structure.

[0212] The aps_extension_data_flag syntax element may be parsed / signaled based on the aps_extension_flag syntax element. The aps_extension_data_flag syntax element may be referred to as an APS extension data flag. For example, if the value of aps_extension_flag is 1, the aps_extension_data_flag may be parsed / signaled, and in this case, the aps_extension_data_flag may have any value.

[0213] As described above, according to one embodiment of this document, scaling list data can be efficiently transported by assigning a data type (e.g., SCALING_APS) to indicate scaling list data and parsing / signaling a syntax element (e.g., aps_params_type) that indicates the data type. In other words, according to one embodiment of this document, an APS structure that integrates scaling list data can be used.

[0214] Meanwhile, in the current VVC standard, the use of scaling list data (i.e., scaling_list_data()) can be indicated first based on the presence of a flag (i.e., sps_scaling_list_enabled_flag) indicating the availability of scaling list data in the Sequence Parameter Set (SPS). If the flag (i.e., sps_scaling_list_enabled_flag) is enabled (i.e., 1 or true, indicating that scaling list data is available), another flag (i.e., sps_scaling_list_data_present_flag) can be parsed. Also, if sps_scaling_list_data_present_flag is enabled (i.e., 1 or true, indicating that scaling list data is present in the SPS), scaling list data (i.e., scaling_list_data()) can be parsed. That is, in the current VVC standard, scaling list data is signaled by the SPS. In this case, since the SPS enables session negotiation and is generally transmitted out-of-band, it can be used during the decoding process, eliminating the need to transmit scaling list data with information related to determining the scaling factors of the transform blocks. If the decoder transmits scaling list data in the SPS, the decoder must allocate a significant amount of memory to store the information obtained from the scaling list data and must maintain this information until it is used in transform block decoding. Therefore, this process is unnecessary at the SPS level, and it is more effective to parse / signal it at a lower level. For this reason, this document proposes a hierarchical structure to effectively parse / signal scaling list data.

[0215] In one embodiment, scaling list data is not parsed / signaled from the higher level syntax SPS, but is parsed / signaled in the lower level syntax PPS, tile group header, slice header, and / or other appropriate header.

[0216] For example, the SPS syntax can be modified as shown in Table 21 below: Table 19 below shows an example of SPS syntax for describing a scaling list for a CVS.

[0217] [Table 21]

[0218] The semantics of the syntax elements included in the SPS syntax in Table 21 can be shown as in Table 22 below.

[0219] [Table 22]

[0220] Referring to Tables 21 and 22, the scaling_list_enabled_flag syntax element can be parsed / signaled in the SPS. The scaling_list_enabled_flag syntax element can indicate whether a scaling list is available based on whether its value is 0 or 1. For example, if the value of scaling_list_enabled_flag is 1, it indicates that the scaling list is used in the scaling process for the transform coefficients, and if the value of scaling_list_enabled_flag is 0, it indicates that the scaling list is not used in the scaling process for the transform coefficients.

[0221] That is, the scaling_list_enabled_flag syntax element can be called a scaling list availability flag and can be signaled at the SPS (or SPS level). That is, based on the value of scaling_list_enabled_flag signaled at the SPS level, it can be determined that a scaling list is essentially available for a picture in a CVS that references the SPS. Then, an additional availability flag can be signaled at a level lower than the SPS (e.g., a PPS, a tile group header, a slice header, and / or other appropriate header) to obtain a scaling list.

[0222] In addition, the sps_scaling_list_data_present_flag syntax element can be prevented from being parsed / signaled at the SPS. That is, by removing the sps_scaling_list_data_present_flag syntax element at the SPS, this flag information can be prevented from being parsed / signaled. The sps_scaling_list_data_present_flag syntax element is flag information indicating whether the syntax structure of scaling list data exists in the SPS, and scaling list data specified by the SPS can be parsed / signaled according to this flag information. However, by removing the sps_scaling_list_data_present_flag syntax element, scaling list data can be prevented from being parsed / signaled at the SPS level.

[0223] As described above, according to one embodiment of this document, the SPS level may be configured to explicitly signal only the scaling list availability flag (scaling_list_enabled_flag) without directly signaling the scaling list (scaling_list_data()). Thereafter, the scaling list (scaling_list_data()) can be parsed individually in the lower level syntax based on the availability flag (scaling_list_enabled_flag) in the SPS. Therefore, according to one embodiment of this document, the scaling list data can be parsed / signaled according to a hierarchical structure, thereby improving coding efficiency.

[0224] Meanwhile, the existence and use of scaling list data is conditioned by the existence of a tool enabling flag. Here, the tool enabling flag is information indicating whether a corresponding tool is enabled, and may include, for example, a scaling_list_enabled_flag syntax element. That is, the scaling_list_enabled_flag syntax element can be used to indicate whether a scaling list is enabled by indicating whether scaling list data is available. However, this tool should be syntactically constrained for the decoder. That is, there should be a constraint flag that informs the decoder that this tool is not currently being used to decode a coded video sequence (CVS). Therefore, this document proposes a method for applying a constraint flag to scaling list data.

[0225] In one embodiment, Table 23 below shows an example of syntax (eg, general restriction information syntax) for signaling scaling list data using restriction flags.

[0226] [Table 23]

[0227] The semantics of the syntax elements included in the syntax of Table 23 can be shown as in Table 22 below.

[0228] [Table 24]

[0229] Referring to Tables 23 and 24, the constraint flag can be parsed / signaled via general_constraint_info(). general_constraint_info() is called the general constraint information field or information about the constraint flag. For example, the no_scaling_list_constraint_flag syntax element can be used as the constraint flag. Here, the constraint flag can be used to specify conformance bitstream properties. For example, if the value of the no_scaling_list_constraint_flag syntax element is 1, it indicates a bitstream conformance requirement in which scaling_list_enabled_flag should be set to 0, and if the value of the no_scaling_list_constraint_flag syntax element is 0, it indicates that there is no constraint.

[0230] Meanwhile, as described above, according to one embodiment of this document, scaling list data can be transmitted via a hierarchical structure. Accordingly, this document proposes a scaling list data structure that can be parsed / signaled via a slice header. Here, the slice header can be called a tile group header, or can be mixed with or replaced by a picture header.

[0231] As an embodiment, Table 25 below shows an example of slice header syntax for signaling scaling list data.

[0232] [Table 25]

[0233] The semantics of the syntax elements included in the slice header syntax of Table 25 can be shown as in Table 26 below.

[0234] [Table 26]

[0235] Referring to Tables 25 and 26, the slice_pic_parameter_set_id syntax element can be parsed / signaled in the slice header. The slice_pic_parameter_set_id syntax element can indicate an identifier for the PPS in use. That is, the slice_pic_parameter_set_id syntax element is information for identifying the PPS referenced in the slice and can indicate the value of pps_pic_parameter_set_id. The value of slice_pic_parameter_set_id must be within the range of 0 to 63. The slice_pic_parameter_set_id syntax element can be referred to as PPS identification information or PPS ID information referenced in the slice.

[0236] In addition, the slice_scaling_list_enabled_flag syntax element can be parsed / signaled in the slice header. The slice_scaling_list_enabled_flag syntax element can indicate whether a scaling list is currently available in the slice. For example, if the value of slice_scaling_list_enabled_flag is 1, it can indicate that a scaling list is currently available in the slice, and if the value of slice_scaling_list_enabled_flag is 0, it can indicate that a scaling list is not currently available in the slice. Alternatively, if slice_scaling_list_enabled_flag is not present in the slice header, its value can be inferred to be 0.

[0237] In this case, whether or not the slice_scaling_list_enabled_flag syntax element can be parsed can be determined based on the scaling_list_enabled_flag syntax element signaled in the upper level syntax (i.e., SPS). For example, if the value of scaling_list_enabled_flag signaled in the SPS is 1 (i.e., if it is determined that scaling list data is available at the upper level), the slice_scaling_list_enabled_flag can be parsed in the slice header to determine whether or not to perform a scaling process using the scaling list in the slice.

[0238] In addition, the slice_scaling_list_aps_id syntax element can be parsed / signaled in the slice header. The slice_scaling_list_aps_id syntax element can indicate an identifier for an APS referenced by the corresponding slice. That is, the slice_scaling_list_aps_id syntax element can indicate ID information (adaptation_parameter_set_id) of an APS containing scaling list data referenced by the corresponding slice. Meanwhile, the TemporalId (i.e., TemporalID) of an APS NAL unit (i.e., an APS NAL unit containing scaling list data) having the same APS ID information (adaptation_parameter_set_id) as slice_scaling_list_aps_id must be smaller than or equal to the TemporalId (i.e., TemporalID) of the slice NAL unit being coded.

[0239] In addition, whether the slice_scaling_list_aps_id syntax element can be parsed can be determined based on the slice_scaling_list_enabled_flag syntax element. For example, if the value of slice_scaling_list_aps_id is 1 (i.e., if it is determined that the scaling list is enabled in the slice header), slice_scaling_list_aps_id can be parsed. Thereafter, scaling list data can be obtained from the APS indicated by the parsed slice_scaling_list_aps_id.

[0240] Furthermore, when multiple SCALING DATA APSs (multiple APSs including scaling list data) with the same APS ID information (adaptation_parameter_set_id) value are referenced by two or more slices in the same picture, the multiple SCALING DATA APSs with the same APS ID information (adaptation_parameter_set_id) value must contain the same content.

[0241] Furthermore, if the above-mentioned syntax elements are present, the values ​​of the slice header syntax elements slice_pic_parameter_set_id, slice_pic_order_cnt_lsb, and slice_temporal_mvp_enabled_flag must be the same in all slice headers within the picture being coded.

[0242] As described above, according to one embodiment of this document, a hierarchical structure can be used to efficiently signal scaling list data. That is, an availability flag (e.g., scaling_list_enabled_flag) indicating whether scaling list data is available is first signaled at a higher level (SPS syntax), and then an additional availability flag (e.g., slice_scaling_list_enabled_flag) is signaled at a lower level (e.g., slice header, picture header, etc.) to determine whether scaling list data is used at each lower level. Also, APS ID information (e.g., slice_scaling_list_aps_id) referenced by the slice or tile group can be signaled via a lower level (e.g., slice header, picture header, etc.), and the scaling list data can be derived from the APS identified by the APS ID information.

[0243] In addition, this document can be applied to signaling scaling list data using a hierarchical structure as proposed in Tables 23 and 24 above, or the scaling list data can be transmitted via a slice header structure such as Table 25 below.

[0244] As an embodiment, Table 27 below shows an example of slice header syntax for signaling scaling list data, where the slice header may also be called a tile group header, or may be mixed with or substituted for the picture header.

[0245] [Table 27]

[0246] The semantics of the syntax elements included in the slice header syntax of Table 27 can be shown as in Table 28 below.

[0247] [Table 28]

[0248] Referring to Tables 27 and 28, the slice_pic_parameter_set_id syntax element can be parsed / signaled in the slice header. The slice_pic_parameter_set_id syntax element can indicate an identifier for the PPS in use. That is, the slice_pic_parameter_set_id syntax element is information for identifying the PPS referenced in the slice and can indicate the value of pps_pic_parameter_set_id. The value of slice_pic_parameter_set_id must be within the range of 0 to 63. The slice_pic_parameter_set_id syntax element can be referred to as PPS identification information or PPS ID information referenced in the slice.

[0249] In addition, the slice_scaling_list_aps_id syntax element can be parsed / signaled in the slice header. The slice_scaling_list_aps_id syntax element can indicate an identifier for an APS referenced by the slice. That is, the slice_scaling_list_aps_id syntax element can indicate ID information (adaptation_parameter_set_id) of an APS containing scaling list data referenced by the slice. For example, the TemporalId (i.e., TemporalID) of an APS NAL unit (i.e., an APS NAL unit containing scaling list data) having the same APS ID information (adaptation_parameter_set_id) as slice_scaling_list_aps_id must be smaller than or equal to the TemporalId (i.e., TemporalID) of the slice NAL unit being coded.

[0250] In this case, whether the slice_scaling_list_aps_id syntax element can be parsed can be determined based on the scaling_list_enabled_flag syntax element signaled in the upper level syntax (i.e., SPS). For example, if the value of scaling_list_enabled_flag signaled in the SPS is 1 (i.e., if it is determined that scaling list data is available at the upper level), slice_scaling_list_aps_id can be parsed in the slice header. Thereafter, scaling list data can be obtained from the APS indicated by the parsed slice_scaling_list_aps_id.

[0251] That is, according to this embodiment, an APS ID including scaling list data can be parsed if the corresponding flag (e.g., scaling_list_enabled_flag) in the SPS is enabled, and therefore, as shown in Table 25 above, APS ID (e.g., slice_scaling_list_aps_id) information including scaling list data referenced at a corresponding lower level (e.g., slice header or picture header) can be parsed based on the scaling_list_enabled_flag syntax element signaled in the higher level syntax (i.e., SPS).

[0252] This document also proposes a method for using multiple APSs to signal scaling list data. The following describes a method for efficiently signaling multiple APS IDs containing scaling list data according to one embodiment of this document. This method is useful during bitstream merging.

[0253] As an embodiment, Table 29 below shows an example of slice header syntax for signaling scaling list data using multiple APSs, where the slice header may also be called a tile group header, or may be mixed with or substituted for the picture header.

[0254] [Table 29]

[0255] The semantics of the syntax elements included in the slice header syntax of Table 29 can be shown as in Table 30 below.

[0256] [Table 30]

[0257] Referring to Tables 29 and 30, the slice_pic_parameter_set_id syntax element can be parsed / signaled in the slice header. The slice_pic_parameter_set_id syntax element can indicate an identifier for the PPS in use. That is, the slice_pic_parameter_set_id syntax element is information for identifying the PPS referenced in the slice and can indicate the value of pps_pic_parameter_set_id. The value of slice_pic_parameter_set_id must be within the range of 0 to 63. The slice_pic_parameter_set_id syntax element can be referred to as PPS identification information or PPS ID information referenced in the slice.

[0258] In addition, the slice_scaling_list_enabled_flag syntax element can be parsed / signaled in the slice header. The slice_scaling_list_enabled_flag syntax element can indicate whether a scaling list is currently available in the slice. For example, if the value of slice_scaling_list_enabled_flag is 1, it can indicate that a scaling list is currently available in the slice, and if the value of slice_scaling_list_enabled_flag is 0, it can indicate that a scaling list is not currently available in the slice. Alternatively, if slice_scaling_list_enabled_flag is not present in the slice header, its value can be inferred to be 0.

[0259] In this case, whether or not the slice_scaling_list_enabled_flag syntax element can be parsed can be determined based on the scaling_list_enabled_flag syntax element signaled in the upper level syntax (i.e., SPS). For example, if the value of scaling_list_enabled_flag signaled in the SPS is 1 (i.e., if it is determined that scaling list data is available at the upper level), the slice_scaling_list_enabled_flag can be parsed in the slice header to determine whether or not to perform a scaling process using the scaling list in the slice.

[0260] In addition, the num_scaling_list_aps_ids_minus1 syntax element can be parsed / signaled in the slice header. The num_scaling_list_aps_ids_minus1 syntax element is information for indicating the number of APSs including scaling list data referenced by the slice. For example, the value of the num_scaling_list_aps_ids_minus1 syntax element plus 1 is the number of APSs. The value of num_scaling_list_aps_ids_minus1 must be within the range of 0 to 7.

[0261] Here, whether the num_scaling_list_aps_ids_minus1 syntax element can be parsed can be determined based on the slice_scaling_list_enabled_flag syntax element. For example, if the value of slice_scaling_list_enabled_flag is 1 (i.e., it is determined that scaling list data is available in the corresponding slice), num_scaling_list_aps_ids_minus1 can be parsed. In this case, the slice_scaling_list_aps_id[i] syntax element can be parsed / signaled based on the value of num_scaling_list_aps_ids_minus1.

[0262] That is, slice_scaling_list_aps_id[i] may indicate the identifier (adaptation_parameter_set_id) of the APS including the ith scaling list data (i.e., the ith SCALING LIST APS). That is, APS ID information may be signaled for as many APSs as the number indicated by the num_scaling_list_aps_ids_minus1 syntax element. Meanwhile, the TemporalId (i.e., TemporalID) of an APS NAL unit (i.e., an APS NAL unit including scaling list data) having the same APS ID information (adaptation_parameter_set_id) as slice_scaling_list_aps_id[i] must be smaller than or equal to the TemporalId (i.e., TemporalID) of the slice NAL unit to be coded.

[0263] Furthermore, when multiple SCALING DATA APSs (multiple APSs including scaling list data) with the same APS ID information (adaptation_parameter_set_id) value are referenced by two or more slices in the same picture, the multiple SCALING DATA APSs with the same APS ID information (adaptation_parameter_set_id) value must contain the same content.

[0264] This document also proposes a method for signaling scaling list data in a hierarchical structure to avoid redundant signaling. In one embodiment, signaling of scaling list data in a Picture Parameter Set (PPS) can be eliminated. The scaling list data can be fully signaled in an SPS, APS, and / or other appropriate header set.

[0265] As an embodiment, Table 31 below shows an example of PPS syntax that does not signal scaling list data in the PPS.

[0266] [Table 31]

[0267] Table 32 below is an example showing the semantics of syntax elements (e.g., pps_scaling_list_data_present_flag) that can be removed to avoid redundant signaling of scaling list data in the PPS syntax of Table 31.

[0268] [Table 32]

[0269] With reference to Tables 31 and 32, information for signaling scaling list data in a PPS, for example, the pps_scaling_list_data_present_flag syntax element, can be removed. The pps_scaling_list_data_present_flag syntax element may indicate whether scaling list data is signaled in a PPS based on whether its value is 0 or 1. For example, if the value of the pps_scaling_list_data_present_flag syntax element is 1, this may indicate that scaling list data used for a picture referencing a PPS is derived based on the scaling list specified by the active SPS and the scaling list data specified by the PPS. If the value of the pps_scaling_list_data_present_flag syntax element is 0, this may indicate that the scaling list data used for a picture referencing a PPS is inferred to be the same as that specified by the active SPS. That is, the pps_scaling_list_data_present_flag syntax element can be information indicating whether scaling list data signaled from the PPS is present.

[0270] That is, in this embodiment, in order to prevent redundant signaling of scaling list data, the pps_scaling_list_data_present_flag syntax element is removed (i.e., the pps_scaling_list_data_present_flag syntax element is not signaled) as shown in Table 31, thereby preventing the scaling list data from being parsed / signaled by the PPS.

[0271] This document also proposes a method for signaling scaling list matrices in APS. Existing methods use three modes (i.e., OFF, DEFAULT, and USER_DEFINED modes). OFF mode indicates that scaling list data is not applied to the transform block. DEFAULT mode indicates that a fixed value is used to generate the scaling matrix. USER_DEFINED mode indicates that a scaling matrix is ​​used based on the block size, prediction mode, and color component. Currently, the total number of scaling matrices supported in VVC is 44, which is a significant increase from the 28 in HEVC. Currently, scaling list data is signaled in SPS and can conditionally exist in PPS. By signaling scaling list data in APS, redundant signaling of the same data in SPS and PPS is unnecessary and can be eliminated.

[0272] For example, scaling matrices currently defined in VVC use scaling_list_enabled_flag to indicate whether they are available (enabled) in the SPS. If this flag indicates availability, the scaling list is used in the scaling process for the transform coefficients, and if this flag indicates not availability, the scaling list is not used in the scaling process for the transform coefficients (e.g., OFF mode). In addition, sps_scaling_list_data_present_flag, which indicates whether scaling list data exists in the SPS, can be parsed. In addition to signaling in the SPS, scaling list data can also exist in the PPS. If pps_scaling_list_data_present_flag indicates availability in the PPS, scaling list data can exist in the PPS. If scaling list data exists in both the SPS and the PPS, the scaling list data from the PPS can be used for frames that reference the active PPS. If scaling_list_enable_flag indicates availability but scaling list data exists only in the SPS and not in the PPS, the scaling list data from the SPS can be referenced by the frame. If scaling_list_enable_flag indicates availability and scaling list data is not present in the SPS or PPS, DEFAULT mode may be used. The use of DEFAULT mode may also be signaled within the scaling list data itself. If scaling list data is explicitly signaled, USER DEFINED mode may be used. The scaling factor for a given transform block can be determined using the information signaled in the scaling list. It is proposed to signal scaling list data in the APS.

[0273] Currently, VVC adopts and uses scaling list data, and the scaling matrix supported by VVC is wider than that of HEVC. The scaling matrix supported by VVC allows block sizes to be selected from 4x4 to 64x64 for luma and from 2x2 to 32x32 for chroma. It also integrates rectangular transform block (TB) size, dependent quantization, multiple transform selection, large transform with zeroing out high-frequency coefficients for large transform blocks, intra subblock partitioning (ISP), and intra block copy (IBC). Intra block copy (IBC) and intra coding modes can share the same scaling matrix.

[0274] Therefore, in the case of USER_DEFINED mode, the number of matrices signaled can be as follows:

[0275] ·MatrixType:30=2(2 for intra & IBC / inter)×3(Y / Cb / Cr components)×5(square TB size:from 4×4 to 64×64 for luma, from 2×2 to 32×32 for chroma)

[0276] ·MatrixType_DC:14=2(2 for intra & IBC / inter×1 for Y component)×3(TB size:16×16, 32×32, 64×64)+4(2 for intra & IBC / inter×2 for Cb / Cr components)×2(TB size:16×16, 32×32)

[0277] DC values ​​can be coded separately for scaling matrices of 16x16, 32x32, and 64x64 sizes. If the transform block (TB) size is smaller than 8x8, all elements can be signaled within one scaling matrix. If the transform block (TB) size is equal to or greater than 8x8, only 64 elements within one 8x8 scaling matrix can be signaled as the base scaling matrix. To obtain a square matrix larger than 8x8, the base 8x8 matrix can be upsampled to the required size. When the DEFAULT mode is used, the scaling matrix can be set to 16. Therefore, VVC supports 44 different matrices, while HEVC supports only 28 matrices. Because the number of scaling matrices supported by VVC is wider than HEVC, APS can be used as a more practical choice for signaling scaling list data to avoid redundant signaling in SPS and / or PPS. This can avoid unnecessary redundant signaling of scaling list data.

[0278] To solve the above problems, this document proposes a method for signaling a scaling matrix in an APS. For this purpose, scaling list data can be signaled only in an APS, without the need for it to be signaled in an SPS or conditionally present in a PPS. In addition, the APS ID can be signaled in the slice header.

[0279] In one embodiment, a flag indicating whether scaling list data is available in the SPS (e.g., scaling_list_enable_flag) may be signaled, and a flag indicating whether scaling list data is present in the PPS (e.g., pps_scaling_list_data_present_flag) may be removed, thereby signaling scaling list data in the APS. Alternatively, the APS ID may be signaled in the slice header. In this case, if the value of scaling_list_enable_flag signaled in the SPS is 1 (i.e., indicating that scaling list data is available) and the APS ID is not signaled in the slice header, the DEFAULT scaling matrix may be used. This embodiment of the present document may be implemented with the syntax and semantics shown in Tables 33 to 41 below.

[0280] Table 33 below shows an example of an APS structure used to signal scaling list data.

[0281] [Table 33]

[0282] The semantics of the syntax elements included in the APS syntax of Table 33 can be expressed as shown in Table 34 below.

[0283] [Table 34]

[0284] Referring to Tables 33 and 34, the adaptation_parameter_set_id syntax element can be parsed / signaled in an APS. The adaptation_parameter_set_id provides an identifier for the APS for reference by other syntax elements. That is, the APS can be identified based on the adaptation_parameter_set_id syntax element. The adaptation_parameter_set_id syntax element can be referred to as APS ID information. The APS can be shared between pictures and can be different in different slices within a picture.

[0285] In addition, the aps_params_type syntax element can be parsed / signaled in the APS. The aps_params_type can represent the type of APS parameters transmitted from the APS as shown in Table 35 below. The aps_params_type syntax element can be referred to as APS parameter type information or APS type information.

[0286] For example, Table 35 below is an example showing the types of APS parameters that can be transmitted via APS, and each APS parameter type can be represented corresponding to the value of aps_params_type.

[0287] [Table 35]

[0288] Referring to Table 35, aps_params_type may be a syntax element for classifying the type of the APS. If the value of aps_params_type is 0, the APS type may be ALF_APS, and the APS may carry ALF data, which may include ALF parameters for deriving a filter / filter coefficient. If the value of aps_params_type is 1, the APS type may be LMCS_APS, and the APS may carry LMCS data, which may include LMCS parameters for deriving an LMCS model / bin / mapping index. If the value of aps_params_type is 2, the APS type may be SCALING_APS, and the APS may carry SCALING list data, which may include scaling list data parameters for deriving a frequency-based quantization scaling matrix / scaling factor / scaling list value.

[0289] For example, as shown in Table 33, the aps_params_type syntax element can be parsed / signaled in APS. In this case, if the value of aps_params_type is 0 (i.e., if aps_params_type represents ALF_APS), ALF data (i.e., alf_data()) can be parsed / signaled. Alternatively, if the value of aps_params_type is 1 (i.e., if aps_params_type represents LMCS_APS), LMCS data (i.e., lmcs_data()) can be parsed / signaled. Alternatively, if the value of aps_params_type is 2 (i.e., if aps_params_type represents SCALING_APS), scaling list data (i.e., scaling_list_data()) can be parsed / signaled.

[0290] Also, referring to Tables 33 and 34, the aps_extension_flag syntax element can be parsed / signaled in APS. aps_extension_flag can indicate whether the APS extension data flag (aps_extension_data_flag) syntax element is present. aps_extension_flag can be used, for example, to provide an extension point for a later version of the VVC standard. The aps_extension_flag syntax element can be referred to as an APS extension flag. For example, a value of aps_extension_flag of 0 can indicate that the APS extension data flag (aps_extension_data_flag) is not present in the APS RBSP syntax structure. Alternatively, a value of aps_extension_flag of 1 can indicate that the APS extension data flag (aps_extension_data_flag) is present in the APS RBSP syntax structure.

[0291] The aps_extension_data_flag syntax element may be parsed / signaled based on the aps_extension_flag syntax element. The aps_extension_data_flag syntax element may be referred to as an APS extension data flag. For example, if the value of aps_extension_flag is 1, the aps_extension_data_flag may be parsed / signaled, and in this case, the aps_extension_data_flag may have any value.

[0292] As described above, according to one embodiment of this document, scaling list data can be efficiently transported by assigning a data type (e.g., SCALING_APS) to represent scaling list data and parsing / signaling a syntax element (e.g., aps_params_type) that represents the data type. In other words, according to one embodiment of this document, an APS structure that integrates scaling list data can be used.

[0293] Furthermore, when signaling scaling list data in the APS, whether scaling list data is available can be signaled in the SPS, and based on this, the scaling list data can be parsed / signaled in the APS according to the APS parameter type (e.g., aps_params_type). In addition, in one embodiment of this document, to avoid signaling redundant scaling list data in higher-level syntax, flag information indicating whether a scaling list data syntax structure (e.g., scaling_list_data()) exists in the SPS or PPS can be not parsed / signaled from the SPS or PPS, thereby preventing the scaling list data from being signaled in the SPS or PPS. This can be achieved by the syntax and semantics shown in Tables 36 to 39 below.

[0294] For example, the SPS syntax can be modified as shown in Table 36 below: Table 36 below shows an example of an SPS syntax structure that does not signal scaling list data in the SPS.

[0295] [Table 36]

[0296] The semantics of the syntax elements included in the SPS syntax of Table 36 may be modified as shown in Table 37. For example, in Tables 36 and 37, some syntax elements (e.g., sps_scaling_list_data_present_flag) may be removed from the syntax elements included in the SPS to avoid redundant signaling of scaling list data.

[0297] [Table 37]

[0298] Referring to Tables 36 and 37, the scaling_list_enabled_flag syntax element can be parsed / signaled in the SPS. The scaling_list_enabled_flag syntax element can indicate whether a scaling list is available based on whether its value is 0 or 1. For example, if the value of scaling_list_enabled_flag is 1, it indicates that the scaling list is used in the scaling process for the transform coefficients, and if the value of scaling_list_enabled_flag is 0, it indicates that the scaling list is not used in the scaling process for the transform coefficients.

[0299] That is, the scaling_list_enabled_flag syntax element can be called a scaling list availability flag and can be signaled at the SPS (or SPS level). In other words, based on the value of scaling_list_enabled_flag signaled at the SPS level, it can be determined that a scaling list is essentially available for a picture in a CVS that references the SPS. Then, an additional availability flag can be signaled at a level lower than the SPS (e.g., a PPS, a tile group header, a slice header, and / or other appropriate header) to obtain a scaling list.

[0300] In addition, the sps_scaling_list_data_present_flag syntax element can be prevented from being parsed / signaled in the SPS. That is, by removing the sps_scaling_list_data_present_flag syntax element in the SPS, this flag information can be prevented from being parsed / signaled. The sps_scaling_list_data_present_flag syntax element is flag information indicating whether the syntax structure of scaling list data exists in the SPS, and scaling list data specified by the SPS can be parsed / signaled according to this flag information. However, by removing the sps_scaling_list_data_present_flag syntax element, the SPS level can be configured to explicitly signal only the scaling list availability flag (scaling_list_enabled_flag) without directly signaling scaling list data.

[0301] Also, the signaling of scaling list data in the PPS can be removed as shown in Table 38. For example, Table 38 below shows an example of a PPS syntax structure without signaling scaling list data in the PPS.

[0302] [Table 38]

[0303] The semantics of the syntax elements included in the PPS syntax of Table 38 can be modified as shown in Table 39. As an example, Table 39 shows the semantics for syntax elements (e.g., pps_scaling_list_data_present_flag) that can be removed to avoid redundant signaling of scaling list data in the PPS syntax.

[0304] [Table 39]

[0305] Referring to Tables 38 and 39, the pps_scaling_list_data_present_flag syntax element may not be parsed / signaled in the PPS. That is, by removing the pps_scaling_list_data_present_flag syntax element in the PPS, this flag information may not be parsed / signaled. The pps_scaling_list_data_present_flag syntax element is flag information indicating whether the syntax structure of scaling list data is present in the PPS, and scaling list data designated by the PPS may be parsed / signaled according to this flag information. However, by removing the pps_scaling_list_data_present_flag syntax element, scaling list data may not be directly signaled at the PPS level.

[0306] As described above, by removing flag information indicating whether scaling list data exists at the SPS or PPS level, the SPS syntax and PPS syntax can be configured so that scaling list data syntax is not directly signaled at the SPS or PPS level. Only the scaling list enable flag (scaling_list_enabled_flag) is explicitly signaled in the SPS, and the scaling list (scaling_list_data()) can then be parsed individually in lower-level syntax (e.g., APS) based on the enable flag (scaling_list_enabled_flag) in the SPS. Therefore, according to one embodiment of this document, scaling list data can be parsed / signaled according to a hierarchical structure, thereby improving coding efficiency.

[0307] Furthermore, when signaling scaling list data in an APS, an APS ID may be signaled in a slice header, an APS may be identified based on the APS ID obtained from the slice header, and the scaling list data may be parsed / signaled from the identified APS. Here, the slice header is described as just one example, and the slice header may be mixed with or substituted for various headers such as a tile group header or a picture header.

[0308] For example, Table 40 below shows an example of slice header syntax including the APS ID syntax element for signaling scaling list data in an APS.

[0309] [Table 40]

[0310] The semantics of the syntax elements included in the slice header syntax of Table 40 can be expressed as shown in Table 41 below.

[0311] [Table 41]

[0312] Referring to Tables 40 and 41, the slice_pic_parameter_set_id syntax element can be parsed / signaled in the slice header. The slice_pic_parameter_set_id syntax element can represent an identifier for the PPS in use. That is, the slice_pic_parameter_set_id syntax element is information for identifying the PPS referenced in the slice and can represent the value of pps_pic_parameter_set_id. The value of slice_pic_parameter_set_id should be within the range of 0 to 63. The slice_pic_parameter_set_id syntax element can be said to be PPS identification information or PPSID information referenced in the slice.

[0313] In addition, the slice_scaling_list_present_flag syntax element can be parsed / signaled in the slice header. The slice_scaling_list_present_flag syntax element can be information indicating whether a scaling list matrix exists for the current slice. For example, if the value of slice_scaling_list_present_flag is 1, it can indicate that a scaling list matrix exists for the current slice, and if the value of slice_scaling_list_present_flag is 0, it can indicate that default scaling list data is used to derive a scaling factor array. Alternatively, if slice_scaling_list_present_flag is not present in the slice header, its value can be inferred to be 0.

[0314] In this case, whether or not the slice_scaling_list_present_flag syntax element can be parsed can be determined based on the scaling_list_enabled_flag syntax element signaled in the upper level syntax (i.e., SPS). For example, if the value of scaling_list_enabled_flag signaled in the SPS is 1 (i.e., if the upper level determines that scaling list data is available), the slice_scaling_list_present_flag can be parsed in the slice header to determine whether or not to perform a scaling process using the scaling list in the slice.

[0315] In addition, the slice_scaling_list_aps_id syntax element can be parsed / signaled in the slice header. The slice_scaling_list_aps_id syntax element can represent an identifier for an APS referenced by the corresponding slice. That is, the slice_scaling_list_aps_id syntax element can represent ID information (adaptation_parameter_set_id) of an APS containing scaling list data referenced by the corresponding slice. Meanwhile, the TemporalId (i.e., Temporal ID) of an APS NAL unit (i.e., an APS NAL unit containing scaling list data) having the same APS ID information (adaptation_parameter_set_id) as slice_scaling_list_aps_id must be smaller than or equal to the TemporalId (i.e., Temporal ID) of the slice NAL unit being coded.

[0316] Furthermore, whether the slice_scaling_list_aps_id syntax element can be parsed can be determined based on the slice_scaling_list_present_flag syntax element. For example, if the value of slice_scaling_list_present_flag is 1 (i.e., if a scaling list exists in the slice header), slice_scaling_list_aps_id can be parsed. Then, scaling list data can be obtained from the APS indicated by the parsed slice_scaling_list_aps_id.

[0317] Furthermore, when multiple SCALING DATA APSs (multiple APSs containing scaling list data) with the same APS ID information (adaptation_parameter_set_id) value are referenced by two or more slices in the same picture, the multiple SCALING DATA APSs with the same APS ID information (adaptation_parameter_set_id) value must contain the same content.

[0318] In Tables 40 and 41 above, it has been described that the APS ID is signaled in the slice header, but this is just one example, and in this document, the APS ID can be signaled in the picture header or tile group header, etc.

[0319] This document also proposes a general method for repositioning scaling list data in other header sets. As an embodiment, a general structure for including scaling list data in a header set is proposed. Currently, for VVC, APS is used as the appropriate header set, but it may also be possible to encapsulate scaling list data identified by Nal Unit Type (NUT) in its own header set. This can be realized as shown in Tables 42 and 43 below.

[0320] For example, Table 42 below shows an example of NAL unit types and corresponding RBSP syntax structures. As described above, the NAL unit type of an NAL unit can be identified by the RBSP data structure included in the NAL unit, and information about such NAL unit type can be stored and signaled in the NAL unit header.

[0321] [Table 42-1]

[0322] [Table 42-2]

[0323] As shown in Table 42, scaling list data can be defined as one NAL unit type (e.g., SCALING_NUT), and a specific value (e.g., 21, or one of the reserved values ​​not specified in the NAL unit type) can be specified as the value of the NAL unit type for SCALING_NUT. SCALING_NUT can be a type for an NAL unit that includes a scaling list data parameter set (e.g., Scaling_list_data_parameter set).

[0324] Additionally, the availability of SCALING_NUT may be determined at a higher level than the APS, PPS, and / or other appropriate headers, or may be determined at a lower level than other NAL unit types.

[0325] For example, Table 43 below shows the syntax of a scaling list data parameter set used to signal scaling list data.

[0326] [Table 43]

[0327] The semantics of the syntax elements included in the scaling list data parameter set syntax of Table 43 can be expressed as shown in Table 44 below.

[0328] [Table 44]

[0329] Referring to Tables 43 and 44, a scaling list data parameter set (e.g., scaling_list_data_parameter_set) may be a header set identified by the NAL unit type value for SCALING_NUT (e.g., 21). In the scaling list data parameter set, the scaling_list_data_parameter_set_id syntax element may be parsed / signaled. The scaling_list_data_parameter_set_id syntax element provides an identifier for scaling list data for reference by other syntax elements. That is, the scaling list parameter set may be identified based on the scaling_list_data_parameter_set_id syntax element. The scaling_list_data_parameter_set_id syntax element may be referred to as scaling list data parameter set ID information. A scaling list data parameter set may be shared between pictures and may be different for different slices within a picture.

[0330] Scaling list data (e.g., scaling_list_data syntax) can be parsed / signaled from the scaling list data parameter set identified by scaling_list_data_parameter_set_id.

[0331] In addition, the scaling_list_data_extension_flag syntax element can be parsed / signaled in the scaling list data parameter set. The scaling_list_data_extension_flag syntax element can indicate whether the scaling list data extension flag (scaling_list_data_extension_flag) syntax element is present in the scaling list data RBSP syntax structure. For example, a value of scaling_list_data_extension_flag of 1 can indicate that the scaling list data extension flag (scaling_list_data_extension_flag) syntax element is present in the scaling list data RBSP syntax structure. Alternatively, a value of scaling_list_data_extension_flag of 0 can indicate that the scaling list data extension flag (scaling_list_data_extension_flag) syntax element is not present in the scaling list data RBSP syntax structure.

[0332] The scaling_list_data_extension_data_flag syntax element may be parsed / signaled based on the scaling_list_data_extension_flag syntax element. The scaling_list_data_extension_data_flag syntax element may be referred to as an extension data flag for scaling list data. For example, if the value of scaling_list_data_extension_flag is 1, scaling_list_data_extension_data_flag may be parsed / signaled, and in this case, scaling_list_data_extension_data_flag may have any value.

[0333] As described above, according to one embodiment of this document, one header set structure for scaling list data can be defined and used, and the header set for scaling list data can be specified for a NAL unit type (e.g., SCALING_NUT). In this case, the header set for scaling list data can be defined as a scaling list data parameter set (e.g., scaling_list_data_parameter_set), from which the scaling list data can be obtained.

[0334] Meanwhile, this document proposes a method for effectively coding the syntax element scaling_lsit_pred_matrix_id_delta included in the scaling list data.

[0335] Currently, in VVC, the scaling_lsit_pred_matrix_id_delta syntax element is coded using an unsigned integer 0-th order Exp-Golomb-coded syntax element with the left bit first. However, to improve coding efficiency, one embodiment of this document proposes coding the scaling_lsit_pred_matrix_id_delta syntax element, which has a range of 0 to 5, using a fixed length code (e.g., u(3)), as shown in Table 45 below. In this case, to improve efficiency for the entire range, coding using only 3 bits may be sufficient.

[0336] For example, Table 45 below shows an example of a scaling list data syntax structure.

[0337] [Table 45]

[0338] The semantics of the syntax elements included in the scaling list data syntax of Table 45 can be expressed as shown in Table 46 below.

[0339] [Table 46-1]

[0340] [Table 46-2]

[0341] As shown in Tables 45 and 46, the scaling_list_pred_matrix_id_delta syntax element can be parsed / signaled from the scaling list data syntax. The scaling_list_pred_matrix_id_delta syntax element can represent a reference scaling list used to derive the scaling list. In this case, scaling_list_pred_matrix_id_delta can be parsed using a fixed length code (e.g., u(3)).

[0342] On the other hand, when signaling scaling list data in APS, it is possible to impose restrictions on the number of scaling list matrices. This document proposes a method to limit the number of scaling list matrices. The proposed method simplifies implementation and limits the worst-case memory requirement.

[0343] In one embodiment, the number of APSs containing scaling list data (i.e., APSs signaling the scaling list data syntax) can be limited. To this end, the following constraints can be added: These constraints are meant to place holders, i.e., different values ​​can be used.

[0344] For convenience of explanation, an APS containing scaling list data (i.e., an APS signaling scaling list data syntax) can be referred to as a SCALING LIST APS. In other words, as described above, if the APS parameter type (e.g., aps_params_type) transmitted from an APS is a type representing scaling list data parameters (e.g., SCALING_APS), the scaling list data transmitted from the APS can be referred to as a SCALING LIST APS.

[0345] For example, the total number of APSs containing scaling list data (i.e., SCALING LIST APSs) may be less than 3. Of course, other appropriate values ​​may be used. For example, an appropriate value within the range of 0 to 7 may be used. That is, the total number of APSs containing scaling list data (i.e., SCALING LIST APSs) may be set to the range of 0 to 7.

[0346] Also, for example, with respect to an APS containing scaling list data (ie, a SCALING LIST APS), only one SCALING LIST APS is allowed per picture.

[0347] Table 47 below shows an example of syntax elements and their semantics that represent constraints for limiting an APS that includes scaling list data as described above.

[0348] [Table 47]

[0349] Referring to Table 47, the number of APSs (i.e., SCALING LIST APSs) containing scaling list data can be limited based on the syntax element (e.g., slice_scaling_list_aps_id) representing the APS identification information (i.e., APS ID information) of the SCALING LIST APS.

[0350] For example, the syntax element slice_scaling_list_aps_id may represent APS identification information (i.e., APS ID information) of the SCALING LIST APS referenced by the slice. In this case, the value of the syntax element slice_scaling_list_aps_id may be limited to a specific value. As an example, the value of the syntax element slice_scaling_list_aps_id may be limited to a range of 0 to 3. This is just one example, and the value may be limited to have other values. As another example, the value of the syntax element slice_scaling_list_aps_id may be limited to a range of 0 to 7.

[0351] Also, the TemporalId (i.e., Temporal ID) of a SCALING LIST APS NAL unit having an APS ID (adaptation_parameter_set_id) such as slice_lmcs_aps_id must be smaller than or equal to the TemporalId (i.e., Temporal ID) of the coded slice NAL unit.

[0352] Also, for example, when multiple SCALING LIST APSs with the same APS ID (adaptation_parameter_set_id) value are referenced by two or more slices on the same picture, the multiple SCALING LIST APSs with the same APS ID (adaptation_parameter_set_id) value must have the same content.

[0353] Also, for example, only one SCALING LIST APS having the same APS ID (adaptation_parameter_set_id) value and the same content must be referenced by one or more slices in the same picture. In other words, one or more slices in the same picture must reference the same APS containing scaling list data.

[0354] The following drawings are created to illustrate a specific example of the present document. The names of specific devices and specific terms and names (e.g., names of syntax / syntax elements) shown in the drawings are provided for illustrative purposes only, and the technical features of the present document are not limited to the specific names used in the following drawings.

[0355] 7 and 8 show a schematic diagram of an example video / image encoding method and associated components according to an embodiment (or others) of this document.

[0356] The method disclosed in FIG. 7 may be performed by the encoding device 200 disclosed in FIG. 2. Specifically, step S700 of FIG. 7 may be performed by the subtraction unit 231 disclosed in FIG. 2, step S710 of FIG. 7 may be performed by the conversion unit 232 disclosed in FIG. 2, steps S720 to S730 of FIG. 7 may be performed by the quantization unit 233 disclosed in FIG. 2, and step S740 of FIG. 7 may be performed by the entropy encoding unit 240 disclosed in FIG. 2. In addition, the method disclosed in FIG. 7 may be performed including the embodiments described above in this document. Therefore, in FIG. 7, detailed descriptions of content that overlaps with the above-described embodiments will be omitted or simplified.

[0357] As shown in FIG. 7, the encoding apparatus can derive residual samples for a current block (S700).

[0358] In one embodiment, an encoding device may first determine a prediction mode for a current block and derive a prediction sample. For example, the encoding device may determine whether to perform inter prediction or intra prediction on the current block and may determine a specific inter prediction mode or a specific intra prediction mode based on an RD cost. The encoding device may perform prediction according to the determined prediction mode to derive a prediction sample for the current block. In this case, various prediction methods disclosed herein, such as inter prediction or intra prediction, may be applied. The encoding device may also generate and encode information (e.g., prediction mode information) related to the prediction applied to the current block. The encoding device may then derive a residual sample by comparing an original sample and a predicted sample for the current block.

[0359] The encoding device can derive transform coefficients based on the residual samples (S710).

[0360] In one embodiment, the encoding apparatus may derive transform coefficients through a transform process on residual samples. In this case, the encoding apparatus may determine whether to apply a transform to a current block in consideration of coding efficiency. That is, the encoding apparatus may determine whether a transform is to be applied to the residual samples. For example, if a transform is not to be applied to the residual samples, the encoding apparatus may derive the residual samples as transform coefficients. Alternatively, if a transform is to be applied to the residual samples, the encoding apparatus may derive the transform coefficients by performing a transform on the residual samples. In this case, the encoding apparatus may generate and encode transform skip flag information based on whether a transform is to be applied to the current block. The transform skip flag information may be information indicating whether a transform has been applied to the current block or whether the transform has been skipped.

[0361] The encoding device can derive quantized transform coefficients based on the transform coefficients (S720).

[0362] In one embodiment, an encoding apparatus may derive quantized transform coefficients by applying a quantization process to transform coefficients. In this case, the encoding apparatus may apply frequency-specific weighting, which adjusts quantization strength according to frequency. In this case, the quantization process may be further performed based on frequency-specific quantization scale values. The quantization scale values ​​for frequency-specific weighting may be derived using a scaling matrix. For example, the encoding apparatus / decoding apparatus may use a predefined scaling matrix, or the encoding apparatus may configure and encode frequency-specific quantization scale information for the scaling matrix and signal it to the decoding apparatus. The frequency-specific quantization scale information may include scaling list data. A (modified) scaling matrix may be derived based on the scaling list data.

[0363] In addition, the encoding device can perform the inverse quantization process, just like the decoding device. In this case, the encoding device derives a (modified) scaling matrix based on the scaling list data, and then applies inverse quantization to the quantized transform coefficients based on the scaling matrix to derive reconstructed transform coefficients. At this time, the reconstructed transform coefficients may differ from the original transform coefficients due to losses in the transform / quantization process.

[0364] Here, the scaling matrix may refer to the above-mentioned frequency-based quantization scaling matrix, and may be used interchangeably with or instead of the quantization scaling matrix, quantization matrix, scaling matrix, scaling list, etc. for convenience of explanation, and is not limited to the specific names used in this embodiment.

[0365] That is, the encoding device may further apply frequency-specific weighting when performing the quantization process, and may generate scaling list data as information about the scaling matrix. This process has been described in detail using Tables 5 to 17, so duplicated content and detailed description will be omitted in this embodiment.

[0366] The encoding device can generate residual information based on the quantized transform coefficients (S730).

[0367] Here, the residual information is information generated through a transformation and / or quantization procedure and can be information about quantized transformation coefficients, and can include, for example, information such as value information, position information, transformation technique, transformation kernel, quantization parameter, etc. of the quantized transformation coefficients.

[0368] In addition, if frequency-specific weighting is further applied to derive quantized transform coefficients during the quantization process, scaling list data for the quantized transform coefficients may be generated. The scaling list data may include scaling list parameters used to derive the quantized transform coefficients. In this case, the encoding device may generate information related to the scaling list data, for example, an APS including the scaling list data.

[0369] The encoding device may encode image information (or video information) (S740). Here, the image information may include the residual information. The image information may also include information related to the prediction used to derive the prediction sample (e.g., prediction mode information). The image information may also include information related to the scaling list data. That is, the image information may include various information derived during the encoding process, and may be encoded including such various information.

[0370] In one embodiment, the image information may include various information according to the embodiments (and the like) described above in this document, and may include information disclosed in at least one of Tables 1 through 47 above.

[0371] Furthermore, for example, the image information may include an adaptation parameter set (APS). The APS may include APS ID information (APS identification information) and APS type information (APS parameter type information). That is, an APS may be identified based on the APS ID information, and APS parameters corresponding to the type may be included in the APS based on the APS type information. For example, the APS type information may include an ALF type related to an adaptive loop filter (ALF) parameter, an LMCS type related to a luma mapping with chroma scaling (LMCS) parameter, and a scaling list type related to a scaling list data parameter. This may be represented as shown in Table 35 above. For example, if the value of the APS type information is 2, the APS type information may indicate that the APS includes scaling list data parameters.

[0372] As an example, the APS may be configured as shown in Table 33 (or Table 18). The APS ID information (APS identification information) may be the adaptation_parameter_set_id described in Tables 33 and 34 (or Tables 18 and 19). The APS type information may be the aps_params_type described in Tables 33 to 35 (or Tables 18 to 20). For example, if the APS parameter type information (e.g., aps_params_type) is the SCALING_APS type indicating that the APS includes scaling list data (or if the value of the APS parameter type information (e.g., aps_params_type) is 2), the APS may include scaling list data (e.g., scaling_list_data()). That is, the encoding device may signal scaling list data (e.g., scaling_list_data()) via the APS based on the SCALING_APS type information indicating that the APS includes scaling list data. That is, scaling list data may be included in the APS based on the APS type information (SCALING_APS type information). As described above, the scaling list data may include scaling list parameters for deriving a scaling list, scaling matrix, or scale factor used in the quantization / dequantization process. In other words, the scaling list data may include syntax elements used to configure a scaling list.

[0373] Also, as an example, the APS ID information may have a value within a specific range. For example, the value of the APS ID information may have a value within a specific range from 0 to 3 or from 0 to 7. However, these values ​​are listed as examples, and the value range of the APS ID information may have other values. Also, the value range of the APS ID information may be determined based on APS type information (e.g., aps_params_type). For example, if APS type information (e.g., SCALING_APS type) indicates an APS related to scaling list data, the value of the APS ID information may be expressed based on a syntax element (e.g., slice_scaling_list_aps_id) as shown in Table 47 above. For example, if the APS type information (e.g., aps_params_type) indicates an APS including scaling list data (e.g., SCALING_APS type), the value of the APS ID information may have a value within a range from 0 to 3 or from 0 to 7. Here, slices in one picture can refer to the same APS related to the scaling list data. Alternatively, if the APS type information (e.g., aps_params_type) indicates that the APS is related to ALF (e.g., ALF_APS type), the value of the APS ID information may have a value within the range of 0 to 7. Alternatively, if the APS type information (e.g., aps_params_type) indicates that the APS is related to LMCS (e.g., LMCS_APS type), the value of the APS ID information may have a value within the range of 0 to 3.

[0374] Also, for example, the image information may include a Sequence Parameter Set (SPS). The SPS may include first availability flag information indicating whether scaling list data is available. As an example, the SPS may be configured as shown in Table 36 (or Table 21), and the first availability flag information may be scaling_list_enabled_flag described in Tables 36 and 37 (or Tables 21 and 22). In addition, the SPS may prevent this flag information from being parsed / signaled by removing the sps_scaling_list_data_present_flag syntax element. The sps_scaling_list_data_present_flag syntax element is flag information indicating whether the syntax structure of scaling list data is present in the SPS, and scaling list data specified by the SPS may be parsed / signaled according to this flag information. However, by removing the sps_scaling_list_data_present_flag syntax element, the SPS level can be configured to explicitly signal only the scaling list available flag (scaling_list_enabled_flag) without directly signaling the scaling list data.

[0375] Further, for example, the image information may include a PPS (Picture Parameter Set). As an example, the PPS may be configured as shown in Table 38 above. In this case, the PPS may be configured not to include availability flag information indicating whether scaling list data is available. That is, by removing the pps_scaling_list_data_present_flag syntax element from the PPS, this flag information may not be parsed / signaled. The pps_scaling_list_data_present_flag syntax element is flag information indicating whether the syntax structure of scaling list data is present in the PPS, and scaling list data specified by the PPS may be parsed / signaled according to this flag information. However, by removing the pps_scaling_list_data_present_flag syntax element, scaling list data may not be directly signaled at the PPS level.

[0376] In this case, first availability flag information (e.g., scaling_list_enabled_flag) indicating whether scaling list data is available may be signaled in the SPS but not in the PPS. Therefore, based on the first availability flag information signaled in the SPS (e.g., when the value of the first availability flag information (e.g., scaling_list_enabled_flag) is 1 or true), the scaling list data included in the APS may be obtained.

[0377] Further, for example, the image information may include header information. The header information may be header information related to a slice or picture including the current block, and may include, for example, a picture header or a slice header. The header information may include scaling list data-related APS identification information. The scaling list data-related APS identification information included in the header information may represent APS ID information for an APS including scaling list data. For example, the scaling list data-related APS ID information included in the header information may be slice_scaling_list_aps_id described in Tables 40 to 41 (or Tables 25 to 28), and may be identification information for an APS (including scaling list data; i.e., SCALING LIST APS) referenced by a slice / picture including the current block. That is, an APS including scaling list data may be identified based on the scaling list data-related APS ID information (e.g., slice_scaling_list_aps_id) of the header information.

[0378] In this case, whether or not to parse / signal APS ID information for the APS including the scaling list data in the header information can be determined based on first availability flag information (scaling_list_enabled_flag) parsed / signaled in the SPS. For example, based on the first availability flag information indicating that the scaling list in the SPS is available (e.g., if the value of the first availability flag information (e.g., scaling_list_enabled_flag) is 1 or true), the header information may include APS ID information for the APS including the scaling list data.

[0379] Furthermore, for example, the header information may include second availability flag information indicating whether scaling list data is available in a picture or a slice. As an example, the second availability flag information may be slice_scaling_list_present_flag (or slice_scaling_list_enabled_flag) described in Tables 40 and 41 (or Tables 25 to 28).

[0380] In this case, whether the header information parses / signals the second availability flag information may be determined based on the first availability flag information (scaling_list_enabled_flag) parsed / signaled in the SPS. For example, the header information may include the second availability flag information based on the first availability flag information indicating that the scaling list in the SPS is available (e.g., if the value of the first availability flag information (e.g., scaling_list_enabled_flag) is 1 or true). Then, the header information may include scaling list data-related APS ID information based on the second availability flag information (e.g., if the value of the second availability flag information (e.g., slice_scaling_list_present_flag) is 1 or true).

[0381] As an example, the encoding device may signal second availability flag information (e.g., slice_scaling_list_present_flag) via header information based on first availability flag information (e.g., scaling_list_enabled_flag) signaled in the SPS, as shown in Table 40 above, and then signal APS ID information (e.g., slice_scaling_list_aps_id) for the APS including the scaling list data via header information based on the second availability flag information (e.g., slice_scaling_list_present_flag).The encoding device may then signal scaling list data from the APS indicated by the signaled APS ID information (e.g., slice_scaling_list_aps_id).

[0382] As described above, the SPS syntax and PPS syntax can be configured so that scaling list data syntax is not directly signaled at the SPS or PPS level. For example, only the scaling list enable flag (scaling_list_enabled_flag) can be explicitly signaled in the SPS, and then the scaling list (scaling_list_data()) can be parsed individually in a lower-level syntax (e.g., APS) based on the enable flag (scaling_list_enabled_flag) in the SPS. Therefore, according to one embodiment of this document, scaling list data can be parsed / signaled according to a hierarchical structure, thereby improving coding efficiency.

[0383] Image information including the various information described above can be encoded and output in the form of a bitstream. The bitstream can be transmitted to a decoding device via a network or a (digital) storage medium. Here, the network can include a broadcasting network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.

[0384] 9 and 10 show a schematic diagram of an example video / image decoding method and associated components according to an embodiment (or others) of this document.

[0385] The method disclosed in FIG. 9 may be performed by the decoding device 300 disclosed in FIG. 3. Specifically, steps S900 to S910 of FIG. 9 may be performed by the entropy decoding unit 310 disclosed in FIG. 3, step S920 of FIG. 9 may be performed by the inverse quantization unit 321 disclosed in FIG. 3, step S930 of FIG. 9 may be performed by the inverse transform unit 321 disclosed in FIG. 3, and step S940 of FIG. 9 may be performed by the addition unit 340 disclosed in FIG. 3. In addition, the method disclosed in FIG. 9 may be performed including the embodiments described above in this document. Therefore, in FIG. 9, detailed descriptions of content that overlaps with the above-described embodiments will be omitted or simplified.

[0386] As shown in FIG. 9, a decoding device can receive image information (or video information) from a bitstream (S900).

[0387] In one embodiment, a decoding device may derive information (e.g., video / image information) necessary for image restoration (or picture restoration) by parsing a bitstream. The image information may include residual information, which may include information such as value information of quantized transform coefficients, position information, a transform technique, a transform kernel, and a quantization parameter. The image information may also include information related to prediction (e.g., prediction mode information). The image information may also include information related to scaling list data. That is, the image information may include various information necessary for the decoding process, and may be decoded based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC.

[0388] In one embodiment, the image information may include various information according to the embodiments (and the like) described above in this document, and may include information disclosed in at least one of Tables 1 through 47 above.

[0389] For example, the image information may include an adaptation parameter set (APS). The APS may include APS ID information (APS identification information) and APS type information (APS parameter type information). That is, an APS may be identified based on the APS ID information, and APS parameters corresponding to the type may be included in the APS based on the APS type information. For example, the APS type information may include an ALF type related to an adaptive loop filter (ALF) parameter, an LMCS type related to a luma mapping with chroma scaling (LMCS) parameter, and a scaling list type related to a scaling list data parameter. This may be represented as shown in Table 35 above. For example, if the value of the APS type information is 2, the APS type information may indicate that the APS includes scaling list data parameters.

[0390] As an example, the APS may be configured as shown in Table 33 (or Table 18). The APS ID information (APS identification information) may be the adaptation_parameter_set_id described in Tables 33 and 34 (or Tables 18 and 19). The APS type information may be the aps_params_type described in Tables 33 to 35 (or Tables 18 to 20). For example, if the APS parameter type information (e.g., aps_params_type) is the SCALING_APS type indicating that the APS includes scaling list data (or if the value of the APS parameter type information (e.g., aps_params_type) is 2), the APS may include scaling list data (e.g., scaling_list_data()). That is, the decoding device may obtain and parse the scaling list data (e.g., scaling_list_data()) through the APS based on the SCALING_APS type information indicating that the APS includes scaling list data. That is, scaling list data may be included in the APS based on the APS type information (SCALING_APS type information). As described above, the scaling list data may include scaling list parameters for deriving a scaling list, scaling matrix, or scale factor used in the quantization / dequantization process. In other words, the scaling list data may include syntax elements used to configure a scaling list.

[0391] Also, as an example, the APS ID information may have a value within a specific range. For example, the value of the APS ID information may have a value within a specific range from 0 to 3 or from 0 to 7. However, these values ​​are merely given as examples, and the value range of the APS ID information may have other values. Also, the value range of the APS ID information may be determined based on APS type information (e.g., aps_params_type). For example, if APS type information (e.g., SCALING_APS type) indicates an APS related to scaling list data, the value of the APS ID information may be expressed based on a syntax element (e.g., slice_scaling_list_aps_id) as shown in Table 47 above. For example, if the APS type information (e.g., aps_params_type) indicates an APS including scaling list data (e.g., SCALING_APS type), the value of the APS ID information may have a value within a range from 0 to 3 or from 0 to 7. Here, slices in one picture can refer to the same APS related to the scaling list data. Alternatively, if the APS type information (e.g., aps_params_type) indicates that the APS is related to ALF (e.g., ALF_APS type), the value of the APS ID information may have a value within the range of 0 to 7. Alternatively, if the APS type information (e.g., aps_params_type) indicates that the APS is related to LMCS (e.g., LMCS_APS type), the value of the APS ID information may have a value within the range of 0 to 3.

[0392] Also, for example, the image information may include a Sequence Parameter Set (SPS). The SPS may include first availability flag information indicating whether scaling list data is available. As an example, the SPS may be configured as shown in Table 36 (or Table 21), and the first availability flag information may be scaling_list_enabled_flag described in Tables 36 and 37 (or Tables 21 and 22). In addition, the SPS may prevent this flag information from being parsed / signaled by removing the sps_scaling_list_data_present_flag syntax element. The sps_scaling_list_data_present_flag syntax element is flag information indicating whether the syntax structure of scaling list data is present in the SPS, and scaling list data specified by the SPS may be parsed / signaled according to this flag information. However, by removing the sps_scaling_list_data_present_flag syntax element, the SPS level can be configured to explicitly signal only the scaling list available flag (scaling_list_enabled_flag) without directly signaling the scaling list data.

[0393] Further, for example, the image information may include a PPS (Picture Parameter Set). As an example, the PPS may be configured as shown in Table 38 above. In this case, the PPS may be configured not to include availability flag information indicating whether scaling list data is available. That is, by removing the pps_scaling_list_data_present_flag syntax element from the PPS, this flag information may be configured not to be parsed / signaled. The pps_scaling_list_data_present_flag syntax element is flag information indicating whether the syntax structure of scaling list data is present in the PPS, and scaling list data specified by the PPS may be parsed / signaled according to this flag information. However, by removing the pps_scaling_list_data_present_flag syntax element, scaling list data may not be directly signaled at the PPS level.

[0394] In this case, first availability flag information (e.g., scaling_list_enabled_flag) indicating whether scaling list data is available may be signaled in the SPS but not in the PPS. Therefore, based on the first availability flag information signaled in the SPS (e.g., when the value of the first availability flag information (e.g., scaling_list_enabled_flag) is 1 or true), the scaling list data included in the APS may be obtained.

[0395] Further, for example, the image information may include header information. The header information may be header information related to a slice or picture including the current block, and may include, for example, a picture header or a slice header. The header information may include scaling list data-related APS identification information. The scaling list data-related APS identification information included in the header information may represent APS ID information for an APS including scaling list data. For example, the scaling list data-related APS ID information included in the header information may be slice_scaling_list_aps_id described in Tables 40 to 41 (or Tables 25 to 28), and may be identification information for an APS (including scaling list data; i.e., SCALING LIST APS) referenced by a slice / picture including the current block. That is, an APS including scaling list data may be identified based on the scaling list data-related APS ID information (e.g., slice_scaling_list_aps_id) of the header information. That is, the decoding device can identify the APS based on the APS ID information (e.g., slice_scaling_list_aps_id) in the header information, and obtain the scaling list data from the APS.

[0396] In this case, whether or not to parse / signal APS ID information for the APS including the scaling list data in the header information can be determined based on first availability flag information (scaling_list_enabled_flag) parsed / signaled in the SPS. For example, based on the first availability flag information indicating that the scaling list in the SPS is available (e.g., if the value of the first availability flag information (e.g., scaling_list_enabled_flag) is 1 or true), the header information may include APS ID information for the APS including the scaling list data.

[0397] Also, for example, the header information may include second availability flag information indicating whether scaling list data is available in the picture or slice. For example, the second availability flag information may be slice_scaling_list_present_flag (or slice_scaling_list_enabled_flag) described in Tables 40 and 41 (or Tables 25 to 28).

[0398] In this case, whether the header information parses / signals the second availability flag information may be determined based on the first availability flag information (scaling_list_enabled_flag) parsed / signaled in the SPS. For example, the header information may include the second availability flag information based on the first availability flag information indicating that the scaling list in the SPS is available (e.g., when the value of the first availability flag information (e.g., scaling_list_enabled_flag) is 1 or true). Then, the header information may include APS ID information for the APS containing the scaling list data based on the second availability flag information (e.g., when the value of the second availability flag information (e.g., slice_scaling_list_present_flag) is 1 or true).

[0399] As an example, the decoding device may acquire second availability flag information (e.g., slice_scaling_list_present_flag) via header information based on first availability flag information (e.g., scaling_list_enabled_flag) signaled in the SPS, as in Table 40 (or Table 25) above, and then acquire APS ID information (e.g., slice_scaling_list_aps_id) for the APS including scaling list data via header information based on the second availability flag information (e.g., slice_scaling_list_present_flag).The decoding device may then acquire scaling list data from the APS indicated by the APS ID information (e.g., slice_scaling_list_aps_id) acquired via the header information.

[0400] As described above, the SPS syntax and PPS syntax can be configured so that scaling list data syntax is not directly signaled at the SPS or PPS level. For example, only the scaling list enable flag (scaling_list_enabled_flag) can be explicitly signaled in the SPS, and then the scaling list (scaling_list_data()) can be parsed individually in a lower-level syntax (e.g., APS) based on the enable flag (scaling_list_enabled_flag) in the SPS. Therefore, according to one embodiment of this document, scaling list data can be parsed / signaled according to a hierarchical structure, thereby improving coding efficiency.

[0401] The decoding device can derive quantized transform coefficients for the current block based on the residual information (S910).

[0402] In one embodiment, the decoding device may acquire residual information included in the image information. As described above, the residual information may include information such as value information, position information, transform technique, transform kernel, quantization parameter, etc. The decoding device may derive quantized transform coefficients for the current block based on the quantized transform coefficient information included in the residual information.

[0403] The decoding device can derive transform coefficients based on the quantized transform coefficients (S920).

[0404] In one embodiment, a decoding device may derive transform coefficients by applying an inverse quantization process to quantized transform coefficients. In this case, the decoding device may apply frequency-specific weighting, which adjusts quantization strength according to frequency. In this case, the inverse quantization process may be further performed based on frequency-specific quantization scale values. Quantization scale values ​​for frequency-specific weighting may be derived using a scaling matrix. For example, the decoding device may use a predefined scaling matrix, or may use frequency-specific quantization scale information for a scaling matrix signaled from an encoding device. The frequency-specific quantization scale information may include scaling list data. A (modified) scaling matrix may be derived based on the scaling list data.

[0405] That is, the decoding device may further apply frequency-specific weighting when performing the inverse quantization process, and may derive transform coefficients by applying the inverse quantization process to the quantized transform coefficients based on the scaling list data.

[0406] In one embodiment, a decoding device may acquire an APS included in image information and may acquire scaling list data based on APS type information included in the APS. For example, the decoding device may acquire scaling list data included in the APS based on SCALING_APS type information, which indicates that the APS includes scaling list data. In this case, the decoding device may derive a scaling matrix based on the scaling list data, derive scaling factors based on the scaling matrix, and derive transform coefficients by applying inverse quantization based on the scaling factors. The process of performing scaling based on such scaling list data has been described in detail using Tables 5 to 17, so duplicated content and detailed description will be omitted in this embodiment.

[0407] The decoding device may also determine whether to apply frequency-specific weighting during the inverse quantization process (i.e., whether to derive transform coefficients using a scaling list (frequency-based quantization) during the inverse quantization process). For example, the decoding device may determine whether to use scaling list data based on a first availability flag obtained from an SPS included in the image information and / or second availability flag information obtained from header information included in the image information. If it is determined to use scaling list data based on the first availability flag and / or second availability flag information, the decoding device may obtain APS ID information for an APS including scaling list data via the header information, identify the APS based on the APS ID information, and obtain scaling list data from the identified APS.

[0408] The decoding device can derive residual samples based on the transform coefficients (S930).

[0409] In one embodiment, the decoding device may derive residual samples of the current block by performing an inverse transform on transform coefficients of the current block. In this case, the decoding device may obtain information indicating whether to apply an inverse transform to the current block (i.e., transform skip flag information) and derive residual samples based on the information (i.e., transform skip flag information).

[0410] For example, if an inverse transform is not applied to the transform coefficients (the value of the transform skip flag information for the current block is 1), the decoding device may derive the transform coefficients from the residual samples of the current block. Alternatively, if an inverse transform is applied to the transform coefficients (the value of the transform skip flag information for the current block is 0), the decoding device may derive the residual samples of the current block by performing an inverse transform on the transform coefficients.

[0411] The decoding device can generate reconstructed samples based on the residual samples (S940).

[0412] In one embodiment, a decoding device may determine whether to perform inter-prediction or intra-prediction on a current block based on prediction information (e.g., prediction mode information) included in image information, and may perform prediction based on the determination to derive a predicted sample for the current block. The decoding device may then generate a reconstructed sample based on the predicted sample and a residual sample. In this case, the decoding device may directly use the predicted sample as a reconstructed sample depending on the prediction mode, or may generate a reconstructed sample by adding the residual sample to the predicted sample. Furthermore, the decoding device may derive a reconstructed block or picture based on the reconstructed sample. As described above, the decoding device may then apply an in-loop filtering procedure, such as deblocking filtering and / or an SAO procedure, to the reconstructed picture to improve subjective / objective image quality as needed.

[0413] In the above-described embodiments, the method is described based on a flow chart with a series of steps or blocks, but the embodiments of this document are not limited to the order of steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flow chart are not exclusive, and other steps may be included, or one or more steps in the flow chart may be deleted without affecting the scope of this document.

[0414] The method according to the present document described above can be implemented in software form, and the encoding device and / or decoding device according to the present document can be included in a device that performs image processing, such as a TV, a computer, a smartphone, a set-top box, or a display device.

[0415] When the embodiments herein are implemented in software, the methods described above may be implemented with modules (processes, functions, etc.) that perform the functions described above. The modules may be stored in memory and executed by a processor. The memory may be internal or external to the processor and may be coupled to the processor in various well-known ways. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described herein may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in the figures may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information on instructions) or algorithms for implementation may be stored on a digital storage medium.

[0416] In addition, decoding devices and encoding devices to which this document applies may be included in multimedia broadcasting transmitting / receiving devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video interaction devices, real-time communication devices such as video communications, mobile streaming devices, storage media, camcorders, custom video (VoD) service providing devices, over-the-top (OTT) video (over-the-top) devices, internet streaming service providing devices, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, image telephone video devices, transportation terminals (e.g., vehicles (including autonomous vehicles), airplane terminals, ship terminals, etc.), medical video devices, etc., and may be used to process video signals or data signals. For example, over-the-top (OTT) video (over-the-top) devices may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.

[0417] In addition, the processing method to which this document is applied may be produced in the form of a computer-executable program and stored in a computer-readable recording medium. Multimedia data having a data structure according to this document may also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. Examples of the computer-readable recording medium include Blu-ray Discs (BDs), Universal Serial Buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media implemented in the form of carrier waves (e.g., transmission via the Internet). A bitstream generated by the encoding method may be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.

[0418] Furthermore, the embodiments of the present document may be implemented in a computer program product by program code, which may be executed by a computer in accordance with the embodiments of the present document. The program code may be stored on a computer-readable carrier.

[0419] FIG. 11 illustrates an example of a content streaming system in which the embodiments disclosed herein may be applied.

[0420] Referring to FIG. 11, the content streaming system applied to the embodiments of this document can be broadly divided into an encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.

[0421] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted.

[0422] The bitstream may be generated by an encoding method or a bitstream generation method applied to an embodiment of this document, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0423] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which controls commands and responses between devices in the content streaming system.

[0424] The streaming server can receive content from a media repository and / or an encoding server. For example, if content is received from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.

[0425] Examples of the user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, and digital signs.

[0426] Each server in the content streaming system can be operated as a distributed server, in which case data received by each server can be processed in a distributed manner.

[0427] The claims herein may be combined in various ways. For example, the technical features of the method claims herein may be combined and realized in an apparatus, the technical features of the apparatus claims herein may be combined and realized in a method, the technical features of the method claims herein and the technical features of the apparatus claims herein may be combined and realized in an apparatus, and the technical features of the method claims herein and the technical features of the apparatus claims herein may be combined and realized in a method.

Claims

1. An image decoding method performed by a decoding device, obtaining image information including residual information and prediction information from a bitstream; deriving a predicted sample for a current block by performing prediction based on the prediction information; deriving quantized transform coefficients for the current block based on the residual information; deriving transform coefficients based on the quantized transform coefficients; deriving residual samples based on the transform coefficients; generating reconstructed samples based on the predicted samples and the residual samples; the image information includes an adaptation parameter set (APS) including scaling list data; The APS includes APS ID information and APS type information, The APS is identified based on the APS ID information; the APS type information specifies whether the APS is related to an adaptive loop filter (ALF), a luma mapping with chroma scaling (LMCS), or the scaling list data; The APS includes the scaling list data based on the APS type information; the APS ID information has a value within a specific range based on the APS type information that identifies the APS as an APS related to the scaling list data; The method, wherein the specific range is predetermined for the APS type information that identifies the APS as an APS related to the scaling list data.

2. An image encoding method performed by an encoding device, deriving a predicted sample for the current block by performing a prediction; deriving a residual sample for the current block based on the predicted sample; deriving transform coefficients based on the residual samples; deriving quantized transform coefficients by applying a quantization process to the transform coefficients; generating forecast information based on the forecast; generating residual information including information about the quantized transform coefficients; encoding image information including the prediction information and the residual information; the image information includes an adaptation parameter set (APS) including scaling list data; The APS includes APS ID information and APS type information, The APS is identified based on the APS ID information; the APS type information specifies whether the APS is related to an adaptive loop filter (ALF), a luma mapping with chroma scaling (LMCS), or the scaling list data; The APS includes the scaling list data based on the APS type information; the APS ID information has a value within a specific range based on the APS type information that identifies the APS as an APS related to the scaling list data; The method, wherein the specific range is predetermined for the APS type information that identifies the APS as an APS related to the scaling list data.

3. In a method for transmitting data for an image, obtaining a bitstream for the image, the bitstream comprising: deriving a predicted sample for the current block by performing a prediction; deriving a residual sample for the current block based on the predicted sample; deriving transform coefficients based on the residual samples; deriving quantized transform coefficients by applying a quantization process to the transform coefficients; generating forecast information based on the forecast; generating residual information including information about the quantized transform coefficients; encoding image information including the prediction information and the residual information; transmitting the data including the bitstream; the image information includes an adaptation parameter set (APS) including scaling list data; The APS includes APS ID information and APS type information, The APS is identified based on the APS ID information; the APS type information specifies whether the APS is related to an adaptive loop filter (ALF), a luma mapping with chroma scaling (LMCS), or the scaling list data; The APS includes the scaling list data based on the APS type information; the APS ID information has a value within a specific range based on the APS type information that identifies the APS as an APS related to the scaling list data; The method, wherein the specific range is predetermined for the APS type information that identifies the APS as an APS related to the scaling list data.

Citation Information

Patent Citations

  • Adaptation parameter set types in video coding

    WO2020176633A1