Image or video coding based on scaling list data
By signaling scaling list data through an APS and configuring syntax to avoid direct signaling, the method addresses the inefficiencies in high-resolution image/video compression, enhancing coding efficiency and visual quality while reducing memory requirements.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2025-07-17
- Publication Date
- 2026-05-11
AI Technical Summary
The increasing demand for high-resolution and high-quality images/videos, including VR and AR content, has led to higher transmission and storage costs due to increased data volume, necessitating a more efficient image/video compression technology with improved coding efficiency and subjective/objective visual quality.
The method involves signaling scaling list data through an APS, configuring the SPS and PPS syntax to avoid direct signaling of scaling list data, and hierarchically signaling availability flag information, allowing for efficient configuration and application of scaling lists during the scaling process.
This approach enhances overall image/video compression efficiency, improves coding efficiency, and enhances subjective/objective visual quality by efficiently configuring and applying scaling lists, while limiting memory requirements and simplifying implementation.
Smart Images

Figure 0007856830000053 
Figure 0007856830000054 
Figure 0007856830000055
Abstract
Description
Technical Field
[0001] The present technology relates to video or image coding, for example, coding technology based on scaling list data.
Background Art
[0002] In recent years, the demand for high-resolution and high-quality images / videos such as 4K or UHD (Ultra High Definition) images / videos of 8K or higher has been increasing in various fields. As the image / video data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to existing image / video data. Therefore, when transmitting image data using a medium such as an existing wired or wireless broadband line or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.
[0003] In addition, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) contents, and holograms have been increasing, and the broadcast of images / videos having image characteristics different from real images, such as game images, has been increasing.
[0004] Therefore, there is a need for a highly efficient image / video compression technology to effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality images / videos having various characteristics as described above.
[0005] In addition, in order to improve the compression efficiency and enhance the subjective / objective visual quality, there is a discussion on the adaptive frequency weighting quantization technology in the scaling process. A method for signaling related information is required to efficiently apply such technology.
Summary of the Invention
Problems to be Solved by the Invention
[0006] The technical objective of this document is to provide a method and apparatus for improving video / image coding efficiency.
[0007] Another technical objective of this paper is to provide methods and apparatus for improving coding efficiency during the scaling process.
[0008] Another technical challenge addressed in this paper is to provide a method and apparatus for efficiently configuring a scaling list used in the scaling process.
[0009] Another technical challenge addressed in this paper is to provide a method and apparatus for hierarchically signaling scaling list-related information used in the scaling process.
[0010] Another technical objective of this paper is to provide a method and apparatus for efficiently applying a scaling list-based scaling process. [Means for solving the problem]
[0011] According to one embodiment of this document, scaling list data can be signaled via an APS (adaptation parameter set). Furthermore, an APS can be identified based on APS ID information contained within the APS, and an APS can contain scaling list data based on APS type information indicating that it is an APS relating to scaling list data. In addition, the value of the APS ID information can be within a specific range for the APS type information indicating that it is an APS relating to scaling list data.
[0012] According to one embodiment of this document, availability flag information indicating the availability of scaling list data can be signaled hierarchically, and availability flag information in lower-level syntax (e.g., picture header / slice header / type group header, etc.) can be signaled based on availability flag information signaled in higher-level syntax (e.g., SPS).
[0013] According to one embodiment of this document, the SPS syntax and PPS syntax can be configured so as not to directly signal the scaling list data syntax at the SPS or PPS level. This improves coding efficiency by allowing the scaling list data to be parsed / signaled individually in the APS based on availability flag information in the SPS.
[0014] According to one embodiment of this document, a video / image decoding method performed by a decoding device is provided. The video / image decoding method may include the methods disclosed in the embodiments of this document.
[0015] According to one embodiment of this document, a decoding device for video / image decoding is provided. The decoding device can perform the method disclosed in the embodiments of this document.
[0016] According to one embodiment of this document, a video / image encoding method performed by an encoding device is provided. The video / image encoding method may include the methods disclosed in the embodiments of this document.
[0017] According to one embodiment of this document, an encoding device for video / image encoding is provided. The encoding device can perform the method disclosed in the embodiments of this document.
[0018] According to one embodiment of this document, a computer-readable digital storage medium is provided which stores encoded video / image information generated by a video / image encoding method disclosed in at least one of the embodiments of this document.
[0019] According to one embodiment of this document, a computer-readable digital storage medium is provided which stores encoded information or encoded video / image information that causes a decoding device to perform a video / image decoding method disclosed in at least one of the embodiments of this document. [Effects of the Invention]
[0020] This document can have various effects. For example, according to one embodiment of this document, overall image / video compression efficiency can be increased. Also, according to one embodiment of this document, coding efficiency can be increased and subjective / objective visual quality can be improved by applying an efficient scaling process. Also, according to one embodiment of this document, the scaling list used in the scaling process can be efficiently configured, and scaling list-related information can be hierarchically signaled through it. Also, according to one embodiment of this document, coding efficiency can be increased by efficiently applying a scaling process based on a scaling list. Also, according to one embodiment of this document, by setting restrictions on the scaling list matrix used in the scaling process, implementation can be made easier and the worst-case memory requirement can be limited.
[0021] The effects that can be obtained through the specific embodiments of this document are not limited to the effects listed above. For example, there may be various technical effects that a person having ordinary skill in the related art can understand or derive from this document. Thus, the specific effects of this document are not limited to what is explicitly described in this document, and can include various effects that can be understood or derived from the technical features of this document.
Brief Description of the Drawings
[0022] [Figure 1] An example of a video / image coding system to which the embodiments of this document can be applied is schematically shown. [Figure 2] It is a diagram schematically explaining the configuration of a video / image encoding device to which the embodiments of this document can be applied. [Figure 3] It is a diagram schematically explaining the configuration of a video / image decoding device to which the embodiments of this document can be applied. [Figure 4] An example of a schematic video / image encoding method to which the embodiments of this document are applicable is shown. [Figure 5] An example of a schematic video / image decoding method to which the embodiments of this document are applicable is shown. [Figure 6] A hierarchical structure for the coded image / video is exemplarily shown. [Figure 7] An example of a video / image encoding method and related components according to the embodiments (etc.) of this document is schematically shown. [Figure 8] An example of a video / image encoding method and related components according to the embodiments (etc.) of this document is schematically shown. [Figure 9] An example of a video / image decoding method and related components according to the embodiments (etc.) of this document is schematically shown. [Figure 10] An example of a video / image decoding method and related components according to the embodiments (etc.) of this document is schematically shown. [Figure 11] Examples of content streaming systems to which the embodiments disclosed in this document may be applied are shown. [Modes for carrying out the invention]
[0023] This disclosure can be modified in various ways and may have various embodiments; however, specific embodiments are illustrated in the drawings and described in detail. This does not mean that this disclosure is limited to any particular embodiment. Terms used herein are used solely to describe specific embodiments and are not intended to limit the technical ideas of this disclosure. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as “includes” or “has” are intended to indicate the existence of features, figures, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the existence or possibility of adding one or more other features, figures, steps, actions, components, parts, or combinations thereof.
[0024] On the other hand, each configuration shown in the drawings described in this disclosure is illustrated independently for the purpose of explaining its distinct characteristic functions, and does not mean that each configuration is implemented with separate hardware or separate software. For example, two or more of the configurations may be combined to form one configuration, and one configuration may be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included in the scope of the rights of this disclosure, as long as they do not deviate from the essence of this document.
[0025] In this document, “A or B” can mean “just A,” “just B,” or “both A and B.” Furthermore, in this document, “A or B” can be interpreted as “A and / or B.” For example, in this document, “A, B or C” can mean “just A,” “just B,” “just C,” or “any combination of A, B and C.”
[0026] In this document, slashes ( / ) and commas can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "just A", "just B", or "both A and B". For example, "A, B, C" can mean "A, B or C".
[0027] In this document, "at least one of A and B" can mean "just A," "just B," or "both A and B." Furthermore, in this document, the expressions "at least one of A or B" and "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B."
[0028] Furthermore, in this document, "at least one of A, B and C" can mean "just A," "just B," "just C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" can mean "at least one of A, B and C."
[0029] Furthermore, the parentheses used in this document can mean "for example." Specifically, when displayed as "prediction (intra prediction)," "intra prediction" is proposed as an example of "prediction." In other words, "prediction" in this document is not limited to "intra prediction," but rather "intra prediction" is proposed as an example of "prediction." Similarly, when displayed as "prediction (i.e., intra prediction)," "intra prediction" is proposed as an example of "prediction."
[0030] This document relates to video / image coding. For example, the methods / embodiments disclosed herein can be applied to methods disclosed in the VVC (Versatile Video Coding) standard. Furthermore, the methods / embodiments disclosed herein can be applied to methods disclosed in the EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or next-generation video / image coding standards (e.g., H.267 or H.267).
[0031] This document presents various embodiments relating to video / image coding, and unless otherwise noted, these embodiments may be implemented in combination with each other.
[0032] In this document, "video" can mean a collection of images over time. "Picture" generally refers to a unit representing a single image at a specific time point in time, and "slice" or "tile" is a unit that constitutes part of a picture in coding. A slice or tile can contain one or more coding tree units (CTUs). A single picture can consist of one or more slices or tiles. A tile is a rectangular region of CTUs within a particular tile column and particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture, and a width specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a height equal to the height of the picture.A tile scan can represent a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in CTU raster scan in a tile, whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice includes an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that may be exclusively contained in a single NAL unit.
[0033] On the other hand, a single picture can be divided into two or more subpictures. A subpicture is a rectangular region of one or more slices within a picture.
[0034] A pixel or pel can refer to the smallest unit that makes up a picture (or image). Alternatively, the term "sample" may be used as a counterpart to pixel. A sample can generally represent a pixel or a pixel value, or it may represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. Alternatively, a sample can refer to a pixel value in the spatial domain, and if such a pixel value is converted to the frequency domain, it can also refer to the conversion coefficient in the frequency domain.
[0035] A unit can represent a basic unit of image processing. A unit can contain at least one of a specific region of a picture and information associated with that region. A unit can contain one luma block and two chroma (e.g., cb, cr) blocks. The term unit may sometimes be used interchangeably with terms such as block or area. In general, an M×N block can contain a sample (or sample array) consisting of M columns and N rows, or a set (or array) of transform coefficients.
[0036] Furthermore, in this document, at least one of quantization / inverse quantization and / or transformation / inverse transformation may be omitted. If quantization / inverse quantization is omitted, the quantized transformation coefficients may be called transformation coefficients. If transformation / inverse transformation is omitted, the transformation coefficients may also be called coefficients or residual coefficients, or, for consistency of expression, may still be called transformation coefficients.
[0037] In this document, quantized transformation coefficients and transformation coefficients can be referred to as transformation coefficients and scaled transformation coefficients, respectively. In this case, residual information may include information about the transformation coefficients, and information about the transformation coefficients can be signaled via residual coding syntax. Transformation coefficients can be derived based on residual information (or information about the transformation coefficients), and scaled transformation coefficients can be derived via inverse transformation (scaling) of the transformation coefficients. Residual samples can be derived based on inverse transformation (transformation) of the scaled transformation coefficients. This can be applied / expressed similarly in other parts of this document.
[0038] In this document, technical features described individually within a single drawing may be implemented individually or simultaneously.
[0039] Preferred embodiments of this document will be described in more detail below with reference to the attached drawings. Hereafter, the same reference numerals will be used for the same components in the drawings, and redundant descriptions of the same components may be omitted.
[0040] Figure 1 schematically shows an example of a video / image coding system that can be applied to the embodiments described herein.
[0041] Referring to Figure 1, a video / image coding system may include a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or network.
[0042] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may consist of a separate device or external component.
[0043] A video source can acquire video / images through processes such as video / image capture, synthesis, or generation. A video source may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, or a video / image archive containing previously captured video / images. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images can be generated via a computer, in which case the video / image capture process can be replaced by the process of generating the associated data.
[0044] An encoding device can encode input video / images. For compression and coding efficiency, the encoding device can perform a series of steps, including prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in bitstream format.
[0045] The transmitting unit can transmit encoded video / image information or data output in bitstream format to the receiving unit of a receiving device via a digital storage medium or network in file or streaming format. The digital storage medium can include a variety of storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitting unit may include elements for generating media files via a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0046] A decoding device can decode video / images by performing a series of steps, such as inverse quantization, inverse transformation, and prediction, corresponding to the operation of the encoding device.
[0047] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.
[0048] Figure 2 is a schematic diagram illustrating the configuration of a video / image encoding device to which the embodiments described in this document may be applied. Hereinafter, the term "encoding device" may include an image encoding device and / or a video encoding device.
[0049] Referring to Figure 2, the encoding device 200 may be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be called a rebuilder or a reconstructed block generator. The aforementioned image segmentation unit 210, prediction unit 220, residual processing unit 230, entropy encoding unit 240, addition unit 250, and filtering unit 260 can be configured by one or more hardware components (e.g., an encoder chipset or processor) depending on the embodiment. The memory 270 may also include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0050] The image splitting unit 210 can split an input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, one of these processing units may be called a coding unit (CU). In this case, a coding unit can be recursively split from a coding tree unit (CTU) or the largest coding unit (LCU) using a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be split into multiple coding units of deeper depth based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary tree structure and / or the ternary structure. Alternatively, the binary tree structure may be applied first. The coding procedure according to this document can be performed based on the final coding unit that cannot be further split. In this case, based on coding efficiency due to image characteristics, the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into lower-depth coding units so that the optimally sized coding unit is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit can each be separated or partitioned from the final coding unit described above.The prediction unit is a unit of sample prediction, and the conversion unit is a unit that derives a conversion coefficient and / or a unit that derives a residual signal from the conversion coefficient.
[0051] The term "unit" can sometimes be used interchangeably with terms such as "block" or "area." Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, or it can represent only the pixel / pixel value of the lumen component, or only the pixel / pixel value of the chroma component. A sample can be used as the term corresponding to a single picture (or image) as a pixel or pel.
[0052] The encoding device 200 can generate a residual signal (residual block, residual sample array) by subtracting the prediction signal (predicted block, predicted sample array) output from the inter-prediction unit 221 or intra-prediction unit 222 from the input image signal (original block, original sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, the unit that subtracts the prediction signal (predicted block, predicted sample array) from the input image signal (original block, original sample array) within the encoder 200 can be called the subtraction unit 231. The prediction unit can perform a prediction on the block to be processed (hereinafter referred to as the current block) and generate a predicted block that includes the predicted sample for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied on a current block or CU basis. The prediction unit can generate various information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240, as will be described later in the explanation of each prediction mode. The prediction information can be encoded by the entropy encoding unit 240 and output in bitstream format.
[0053] The intra-prediction unit 222 can predict the current block by referring to a sample in the current picture. The referenced sample may be located adjacent to the current block or at a distance, depending on the prediction mode. The prediction mode in intra-prediction can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC mode and planar mode. Depending on the degree of fineness of the prediction direction, the directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 222 can also determine the prediction mode to be applied to the current block using the prediction modes applied to adjacent blocks.
[0054] The interprediction unit 221 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, adjacent blocks may include spatially adjacent blocks existing in the current picture and temporally adjacent blocks existing in the reference picture. The reference picture containing the reference block and the reference picture containing the temporally adjacent block may be the same or different. The temporally adjacent block may be called a collocated reference block, colCU, etc., and the reference picture containing the temporally adjacent block may be called a collocated picture (colPic). For example, the interpretation unit 221 can construct a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Interpretation can be performed based on various prediction modes; for example, in skip mode and merge mode, the interpretation unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted.In motion vector prediction (MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.
[0055] The prediction unit 220 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for predictions on a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also be based on an intra-block copy (IBC) prediction mode or a palette mode for predictions on blocks. The IBC prediction mode or palette mode can be used for content image / video coding such as in games, for example, as in SCC (screen content coding). IBC basically performs predictions within the current picture, but can be performed similarly to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, sample values within the picture can be signaled based on information about the palette table and palette index.
[0056] The prediction signal generated via the prediction unit (including the inter-prediction unit 221 and / or the intra-prediction unit 222) can be used to generate a reconstructed signal or a residual signal. The transformation unit 232 can apply a transformation technique to the residual signal to generate transformation coefficients. For example, the transformation technique may include at least one of the following: DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to a transformation obtained from a graph when relational information between pixels is represented by this graph. CNT refers to a transformation obtained by generating a prediction signal using all previously reconstructed pixels and based on that. The transformation process can also be applied to pixel blocks of the same size that are square, or to non-square blocks of variable size.
[0057] The quantization unit 233 quantizes the conversion coefficients and transmits them to the entropy encoding unit 240, which can encode the quantized signal (information about the quantized conversion coefficients) and output it as a bitstream. The information about the quantized conversion coefficients can be called residual information. The quantization unit 233 can rearrange the block-form quantized conversion coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector form of the quantized conversion coefficients. The entropy encoding unit 240 can perform various encoding methods, such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). In addition to the quantized conversion coefficients, the entropy encoding unit 240 can also encode information necessary for video / image restoration (e.g., values of syntax elements) together with or separately from the quantized conversion coefficients. Encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form in units of NAL (network abstraction layer) units. The video / image information may further include information about various parameter sets, such as adaptation parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). The video / image information may also further include general constraint information. In this document, information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded via the encoding procedure described above and included in the bitstream.The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include a variety of storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be transmitted by a transmitting unit (not shown) and / or stored by a storage unit (not shown) which are configured as internal / external elements of the encoding device 200, or the transmitting unit may be included in the entropy encoding unit 240.
[0058] The quantized conversion coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transformation to the quantized conversion coefficients via the inverse quantization unit 234 and the inverse transformation unit 235. The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-prediction unit 221 or the intra-prediction unit 222. If there is no residual for the block to be processed, as in the case of skip mode, the predicted block can be used as the reconstructed block. The adder 250 can be called the reconstruction unit or reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, and can also be used for inter-prediction of the next picture after filtering, as described later.
[0059] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture encoding and / or restoration process.
[0060] The filtering unit 260 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 270, specifically in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter. The filtering unit 260 can generate various filtering information and transmit it to the entropy encoding unit 240, as will be described later in the explanation of each filtering method. The filtering information can be encoded by the entropy encoding unit 240 and output in bitstream form.
[0061] The corrected restored picture sent to memory 270 can be used as a reference picture in the interpretation unit 221. When interpretation is applied via this, the encoding device can avoid prediction mismatches between the encoding device 100 and the decoding device, and can also improve encoding efficiency.
[0062] The DPB in memory 270 can store the corrected restored picture for use as a reference picture in the inter-prediction unit 221. Memory 270 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 221 for use as motion information of spatially adjacent blocks or motion information of temporally adjacent blocks. Memory 270 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 222.
[0063] Figure 3 is a schematic diagram illustrating the configuration of a video / image decoding device to which the embodiments described in this document may be applied. Hereinafter, the decoding device may include an image decoding device and / or a video decoding device.
[0064] Referring to Figure 3, the decoding device 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-prediction unit 331 and an intra-prediction unit 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. The aforementioned entropy decoder 310, residual processor 320, predictor 330, adder 340, and filtering unit 350 can be configured by a single hardware component (e.g., a decoder chipset or processor) depending on the embodiment. The memory 360 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The aforementioned hardware component may further include memory 360 as an internal / external component.
[0065] When a bitstream containing video / image information is input, the decoding device 300 can reconstruct the image in accordance with the process by which the video / image information was processed in the encoding device shown in Figure 2. For example, the decoding device 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Thus, the decoding processing units are, for example, coding units, which can be divided from a coding tree unit or a maximum coding unit according to a quad-tree structure, a binary tree structure, and / or a terminally tree structure. One or more conversion units can be derived from the coding unit. The reconstructed image signal decoded and output via the decoding device 300 can then be reproduced via a playback device.
[0066] The decoding device 300 can receive the signal output from the encoding device shown in Figure 2 in bitstream form, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information necessary for image restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The decoding device can decode the picture based on the parameter set information and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements necessary for image reconstruction and the quantized values of conversion coefficients related to the residual. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded and the decoded information of adjacent and decoded blocks or the symbol / bin information decoded in a previous step, predicts the probability of bin occurrence based on the determined context model, and performs arithmetic decoding of the bin to generate symbols corresponding to the values of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin.Of the information decoded by the entropy decoding unit 310, information related to prediction is provided to the prediction unit (inter-prediction unit 332 and intra-prediction unit 331), and the residual values from which entropy decoding has been performed in the entropy decoding unit 310, i.e., quantized conversion coefficients and related parameter information, can be input to the residual processing unit 320. The residual processing unit 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). In addition, of the information decoded by the entropy decoding unit 310, information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives signals output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310. On the other hand, the decoding device described in this document may be called a video / image / picture decoding device, and the decoding device may also be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, inverse transformation unit 322, addition unit 340, filtering unit 350, memory 360, interpretation unit 332, and intraprediction unit 331.
[0067] The inverse quantization unit 321 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 321 can rearrange the quantized transformation coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) to obtain the transformation coefficients.
[0068] In the inverse conversion unit 322, the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).
[0069] The prediction unit can perform a prediction on the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 310, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block and can determine a specific intra / inter-prediction mode.
[0070] The prediction unit 320 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for prediction of a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also be based on an intra-block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as in games, for example, as in SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed similarly to inter-prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction. When palette mode is applied, information about the palette table and palette index can be included in the video / image information and signaled.
[0071] The intra-prediction unit 331 can predict the current block by referring to a sample in the current picture. The referenced sample may be located adjacent to the current block or at a distance, depending on the prediction mode. The prediction mode in intra-prediction may include multiple non-directional modes and multiple directional modes. The intra-prediction unit 331 can also determine the prediction mode to be applied to the current block using the prediction modes applied to adjacent blocks.
[0072] The interprediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) identified by motion vectors on a reference picture. In this case, to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in blocks, subblocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In interprediction, adjacent blocks may include spatially adjacent blocks present in the current picture and temporally adjacent blocks present in the reference picture. For example, the interprediction unit 332 can construct a motion information candidate list based on adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction can be performed based on various prediction modes, and the prediction information may include information indicating the mode of interprediction for the current block.
[0073] The summing unit 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter-prediction unit 332 and / or intra-prediction unit 331). If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the restored block.
[0074] The summing unit 340 may be called the restoration unit or restoration block generation unit. The generated restoration signal can be used for intra-prediction of the next block to be processed in the current picture, and can be output after filtering, as described later, or can be used for intra-prediction of the next picture.
[0075] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.
[0076] The filtering unit 350 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.
[0077] The (modified) restored picture stored in the DPB of memory 360 can be used as a reference picture by the inter-prediction unit 332. Memory 360 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 260 for use as motion information of spatially adjacent blocks or motion information of temporally adjacent blocks. Memory 360 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 331.
[0078] In this document, the embodiments described for the filtering unit 260, the inter-prediction unit 221, and the intra-prediction unit 222 of the encoding device 200 can be applied identically or in a corresponding manner to the filtering unit 350, the inter-prediction unit 332, and the intra-prediction unit 331 of the decoding device 300, respectively.
[0079] As mentioned above, prediction is performed to improve compression efficiency when performing video coding. Through this, a predicted block containing predicted samples for the current block, which is the block to be coded, can be generated. Here, the predicted block contains predicted samples in the spatial domain (or pixel domain). The predicted block is derived in both the encoding and decoding devices, and the encoding device can improve image coding efficiency by signaling the decoding device information about the residuals between the original block and the predicted block (residual information), which is not the original sample value of the original block itself. The decoding device can derive a residual block containing residual samples based on the residual information, and can combine the residual block and the predicted block to generate a restored block containing restored samples, and can generate a restored picture containing the restored block.
[0080] The residual information can be generated through transformation and quantization procedures. For example, an encoding device can signal the relevant residual information (via a bitstream) to a decoding device by deriving a residual block between the original block and the predicted block, performing a transformation procedure on the residual samples (residual sample array) contained in the residual block to derive transformation coefficients, and then performing a quantization procedure on the transformation coefficients to derive quantized transformation coefficients. Here, the residual information may include information such as the value information, position information, transformation technique, transformation kernel, and quantization parameters of the quantized transformation coefficients. The decoding device can derive a residual sample (or residual block) by performing an inverse quantization / inverse transformation procedure based on the residual information. The decoding device can generate a reconstructed picture based on the predicted block and the residual block. The encoding device can also derive a residual block by inverse quantization / inverse transformation of the quantized transformation coefficients for reference for subsequent interpretation of the picture, and generate a reconstructed picture based on this.
[0081] Intra prediction can represent a prediction that generates prediction samples for the current block based on reference samples within the picture to which the current block belongs (hereinafter referred to as the current picture). When intra prediction is applied to the current block, adjacent reference samples to be used for intra prediction of the current block can be derived. The adjacent reference samples of the current block can include samples adjacent to the left boundary of the nW × nH size current block, a total of 2 × nH samples adjacent to the bottom-left, samples adjacent to the top boundary of the current block, a total of 2 × nW samples adjacent to the top-right, and one sample adjacent to the top-left of the current block. Alternatively, the adjacent reference samples of the current block can include multiple columns of upper adjacent samples and multiple rows of left adjacent samples. Additionally, the adjacent reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block (nW × nH size), a total of nW samples adjacent to the bottom boundary of the current block, and one sample adjacent to the bottom-right of the current block.
[0082] However, some of the adjacent reference samples in the current block may not yet be decoded or available. In this case, the decoder can construct adjacent reference samples to use for prediction by substituting the unavailable samples as available samples, or by constructing adjacent reference samples to use for prediction through interpolation of available samples.
[0083] If neighboring reference samples are derived, (i) predicted samples can be derived based on the average or interpolation of neighboring reference samples in the current block, or (ii) predicted samples can be derived based on neighboring reference samples in the current block that are located in a specific (predicted) direction relative to the predicted sample. Case (i) can be called the non-directional mode or non-angular mode, and case (ii) can be called the directional mode or angular mode.
[0084] Furthermore, prediction samples can also be generated by interpolation between a first adjacent sample located in the prediction direction of the current block's intra-prediction mode and a second adjacent sample located in the opposite direction of the prediction direction, using the current block's prediction sample as a reference for adjacent reference samples. In the above case, this can be called linear interpolation intra-prediction (LIP). Additionally, chroma prediction samples can be generated based on chroma samples using a linear model (LM). In this case, this can be called LM mode or CCLM (chroma component LM) mode.
[0085] Alternatively, temporary predicted samples for the current block can be derived based on filtered adjacent reference samples, and the predicted samples for the current block can be derived by performing a weighted sum of the temporary predicted samples and at least one reference sample derived by the intra-prediction mode from the existing adjacent reference samples, i.e., the unfiltered adjacent reference samples. In the case described above, this can be called PDPC (Position dependent intra-prediction).
[0086] Furthermore, intra-predictive coding can be performed by selecting the reference sample line with the highest prediction accuracy from among the adjacent multi-reference sample lines in the current block, deriving a predicted sample using the reference sample located in the prediction direction on that line, and then instructing (signaling) the decoding device with the reference sample line used at that time. In the case described above, this can be called multi-reference line intra-prediction or MRL-based intra-prediction.
[0087] Furthermore, the current block can be divided into vertical or horizontal subpartitions, and intra-prediction can be performed based on the same intra-prediction mode, allowing adjacent reference samples to be derived and used on a subpartition basis. In other words, in this case, the intra-prediction mode for the current block is also applied to the subpartitions, and by deriving and using adjacent reference samples on a subpartition basis, intra-prediction performance can be improved in some cases. Such a prediction method can be called ISP (intra sub-partitions) based intra-prediction.
[0088] The intra-prediction methods described above can be distinguished from intra-prediction modes and called intra-prediction types. Intra-prediction types can be referred to by a variety of terms, such as intra-prediction techniques or additional intra-prediction modes. For example, an intra-prediction type (or additional intra-prediction mode, etc.) can include at least one of the aforementioned LIP, PDPC, MRL, and ISP. General intra-prediction methods that do not include the specific intra-prediction types such as LIP, PDPC, MRL, and ISP can be called normal intra-prediction types. Normal intra-prediction types can be generally applied when the aforementioned specific intra-prediction types are not applicable, and predictions can be performed based on the aforementioned intra-prediction modes. On the other hand, post-processing filtering can also be performed on the derived prediction samples as needed.
[0089] Specifically, the intra-prediction procedure may include an intra-prediction mode / type determination step, an adjacent reference sample derivation step, and an intra-prediction mode / type-based predictive sample derivation step. Additionally, a post-filtering step may be performed on the derived predictive samples as needed.
[0090] When intraprediction is applied, the intraprediction mode applied to the current block may be determined using the intraprediction modes of the surrounding blocks. For example, the decoder can select one of the MPM candidates in the MPM (most probable mode) list derived based on the intraprediction modes of the surrounding blocks of the current block (e.g., the left and / or upper surrounding blocks) and additional candidate modes, based on the received MPM index, or it can select one of the remaining intraprediction modes not included in the MPM candidates (and planar modes), based on the remaining intraprediction mode information. The MPM list can be configured to include or exclude planar modes as candidates. For example, if the MPM list includes planar modes as candidates, the MPM list can have 6 candidates, and if the MPM list does not include planar modes as candidates, the MPM list can have 5 candidates. If the MPM list does not include planar mode as a candidate, a not-planar flag (e.g., intra_luma_not_planar_flag) may be signaled to indicate that the current intra-prediction mode of the block is not planar mode. For example, the MPM flag may be signaled first, and the MPM index and not-planar flag may be signaled if the value of the MPM flag is 1. Also, the MPM index may be signaled if the value of the not-planar flag is 1. Here, the reason why the MPM list is configured not to include planar mode as a candidate is not because planar mode is not an MPM, but because planar mode is always considered as an MPM, so the flag (not-planar flag) is signaled first to check whether or not it is planar mode.
[0091] For example, whether the intra-prediction mode currently applied to a block is among the MPM candidates (and planar modes) or in the remaining mode can be indicated based on the MPM flag (e.g., intra_luma_mpm_flag). A value of 1 for the MPM flag indicates that the intra-prediction mode for the current block is among the MPM candidates (and planar modes), while a value of 0 for the MPM flag indicates that the intra-prediction mode for the current block is not among the MPM candidates (and planar modes). A value of 0 for the not-planar flag (e.g., intra_luma_not_planar_flag) indicates that the intra-prediction mode for the current block is planar mode, while a value of 1 for the not-planar flag indicates that the intra-prediction mode for the current block is not planar mode. The MPM index can be signaled in the form of the mpm_idx or intra_luma_mpm_idx syntax element, and remaining intra-prediction mode information can be signaled in the form of the rem_intra_luma_pred_mode or intra_luma_mpm_remainder syntax element. For example, remaining intra-prediction mode information can refer to one of the remaining intra-prediction modes from the overall intra-prediction modes that are not included in the MPM candidates (and planar modes), indexed in order of prediction mode number. The intra-prediction mode can be an intra-prediction mode for a luma component (sample). The intra-prediction mode information below may include at least one of the following: MPM flag (e.g., intra_luma_mpm_flag), not planar flag (e.g., intra_luma_not_planar_flag), MPM index (e.g., mpm_idx or intra_luma_mpm_idx), or remaining intra-prediction mode information (rem_intra_luma_pred_mode or intra_luma_mpm_remainder). In this document, the MPM list may be referred to by various terms such as MPM candidate list, candModeList, etc.If MIP (matrix-based intra prediction) is applied to the current block, a separate mpm flag (e.g., intra_mip_mpm_flag), mpm index (e.g., intra_mip_mpm_idx), and remaining intra prediction mode information (e.g., intra_mip_mpm_remainder) for MIP may be signaled, while the not-planar flag will not be signaled.
[0092] In other words, generally speaking, when an image is divided into blocks, the current block and neighboring blocks to be coded will have similar image characteristics. Therefore, there is a high probability that the current block and neighboring blocks are identical or have similar intra-prediction modes. Thus, the encoder can use the intra-prediction mode of the neighboring block to encode the intra-prediction mode of the current block.
[0093] For example, an encoder / decoder can construct an MPM (most probable modes) list for the current block. The MPM list can also be referred to as an MPM candidate list. Here, MPM can mean a mode used in intra-predictive mode coding to improve coding efficiency by considering the similarity between the current block and surrounding blocks. As mentioned above, the MPM list can include planar modes or exclude them. For example, if the MPM list includes planar modes, the number of candidates in the MPM list can be six. If the MPM list does not include planar modes, the number of candidates in the MPM list can be five. The encoder / decoder can construct an MPM list containing five or six MPMs.
[0094] To construct an MPM list, three types of modes may be considered: default intra modes, neighborhood intra modes, and derved intra modes. In this case, for neighborhood intra modes, two neighborhood blocks may be considered: a left neighborhood block and an upper neighborhood block.
[0095] As described above, if the MPM list is configured not to include planar mode, then planar mode is removed from the list, and the number of MPM list candidates can be set to 5.
[0096] Furthermore, among the intra-prediction modes, non-directional modes (or non-angle modes) may include DC modes based on the average of neighboring reference samples of the current block, or planar modes based on interpolation.
[0097] When inter-prediction is applied, the prediction unit of the encoding / decoding device can perform inter-prediction on a block-by-block basis to derive predicted samples. Inter-prediction can be a prediction derived in a manner that is dependent on data elements (e.g., sample values or motion information) of picture(s) other than the current picture. When inter-prediction is applied to the current block, a predicted block (predicted sample array) for the current block can be derived based on the reference block (reference sample array) identified by the motion vector on the reference picture pointed to by the reference picture index. In this case, in order to reduce the amount of motion information transmitted in inter-prediction mode, the motion information of the current block can be predicted on a block, subblock, or sample basis based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may also include inter-prediction type information (L0 prediction, L1 prediction, Bi prediction, etc.). When inter-prediction is applied, an adjacent block can include both a spatial neighboring block currently present in the picture and a temporal neighboring block present in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different. A temporal neighboring block can be referred to as a collocated reference block or colCU, and the reference picture containing a temporal neighboring block can be referred to as a collocated picture (colPic).For example, a list of motion information candidates can be constructed based on the adjacent blocks of the current block, and flags or index information can be signaled indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block. Interpretation can be performed based on various prediction modes; for example, in skip mode and merge mode, the motion information of the current block is the same as the motion information of the selected adjacent block. In skip mode, unlike merge mode, no residual signal is transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected adjacent block is used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.
[0098] Motion information can include L0 motion information and / or L1 motion information depending on the interpretation type (L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction can be called an L0 motion vector or MVL0, and a motion vector in the L1 direction can be called an L1 motion vector or MVL1. A prediction based on an L0 motion vector can be called an L0 prediction, a prediction based on an L1 motion vector can be called an L1 prediction, and a prediction based on both an L0 motion vector and an L1 motion vector can be called a paired (Bi) prediction. Here, an L0 motion vector can represent a motion vector associated with a reference picture list L0 (L0), and an L1 motion vector can represent a motion vector associated with a reference picture list L1 (L1). A reference picture list L0 can include pictures prior to the current picture in output order as reference pictures, and a reference picture list L1 can include pictures subsequent to the current picture in output order. Prior pictures can be called forward (reference) pictures, and subsequent pictures can be called backward (reference) pictures. The reference picture list L0 can include subsequent pictures as reference pictures relative to the current picture in terms of output order. In this case, earlier pictures can be indexed first in reference picture list L0, and subsequent pictures can be indexed after them. The reference picture list L1 can also include earlier pictures as reference pictures relative to the current picture in terms of output order. In this case, subsequent pictures can be indexed first in reference picture list L1, and earlier pictures can be indexed after them. Here, the output order can correspond to the POC (picture order count) order.
[0099] Figure 4 shows an example of a schematic video / image encoding method to which the embodiments described in this document can be applied.
[0100] The method disclosed in Figure 4 can be performed by the encoding device 200 of Figure 2 described above. Specifically, S400 can be performed by the inter-prediction unit 221 or intra-prediction unit 222 of the encoding device 200, and S410, S420, S430, and S440 can be performed by the subtraction unit 231, conversion unit 232, quantization unit 233, and entropy encoding unit 240 of the encoding device 200, respectively.
[0101] As shown in Figure 4, the encoding device can derive prediction samples through predictions for the current block (S400). The encoding device can decide whether to perform interpretation or intrapretation on the current block, and can determine a specific interpretation mode or specific intrapretation mode based on the RD cost. Based on the determined mode, the encoding device can derive prediction samples for the current block.
[0102] The encoding device can derive a residual sample by comparing the original sample and the predicted sample for the current block (S410).
[0103] The encoding device can derive conversion coefficients through a conversion procedure for residual samples (S420), and quantize the derived conversion coefficients to derive quantized conversion coefficients (S430).
[0104] The encoding device can encode image information including prediction information and residual information, and output the encoded image information in bitstream format (S440). The prediction information is information related to the prediction procedure and may include information regarding the prediction mode and motion information (e.g., when interpretation is applied). The residual information may include information regarding quantized transformation coefficients. The residual information may be entropy coded.
[0105] The output bitstream can be transmitted to a decoding device via a storage medium or network.
[0106] Figure 5 shows an example of a schematic video / image decoding method to which the embodiments described in this document can be applied.
[0107] The method disclosed in Figure 5 can be performed by the decoding device 300 of Figure 3 described above. Specifically, S500 can be performed by the inter-prediction unit 332 or the intra-prediction unit 331 of the decoding device 300. The procedure of decoding the prediction information contained in the bitstream in S500 and deriving the values of the relevant syntax elements can be performed by the entropy decoding unit 310 of the decoding device 300. S510, S520, S530, and S540 can be performed by the entropy decoding unit 310, the inverse quantization unit 321, the inverse transformation unit 322, and the addition unit 340 of the decoding device 300, respectively.
[0108] As shown in Figure 5, the decoding device can perform operations corresponding to those performed by the encoding device. Based on the received prediction information, the decoding device can perform inter-prediction or intra-prediction for the current block and derive prediction samples (S500).
[0109] The decoding device can derive quantized conversion coefficients for the current block based on the received residual information (S510). The decoding device can derive quantized conversion coefficients from the residual information via entropy decoding.
[0110] The decoding device can derive the conversion coefficients by inverse quantization of the quantized conversion coefficients (S520).
[0111] The decoding device derives the residual sample via an inverse transformation procedure for the transformation coefficients (S530).
[0112] The decoding device can generate a reconstructed sample for the current block based on the predicted sample and the residual sample, and generate a reconstructed picture based on this (S540). As described above, an in-loop filtering procedure may then be further applied to the reconstructed picture.
[0113] On the other hand, as mentioned above, the quantization unit of the encoding device can derive quantized conversion coefficients by applying quantization to the conversion coefficients, and the inverse quantization unit of the encoding device or the inverse quantization unit of the decoding device can derive conversion coefficients by applying inverse quantization to the quantized conversion coefficients.
[0114] Generally, in video / image coding, the quantization rate can be varied, and the compression can be adjusted using the varied quantization rate. From an implementation standpoint, considering complexity, instead of directly using the quantization rate, a quantization parameter (QP) can be used. For example, a quantization parameter with integer values from 0 to 63 can be used, and each quantization parameter value can correspond to an actual quantization rate. Quantization parameter (QP) for luma components (luma samples) Y ) and the quantization parameter (QP) for the chromatic component (chromatic sample) C ) can be set differently.
[0115] The quantization process takes a transformation coefficient (C) as input and a quantization rate (Q) as input. step By dividing it into parts, we can obtain quantized transformation coefficients (C') based on this. In this case, considering the computational complexity, the quantization rate is multiplied by the scale to obtain an integer form, and a shift operation can be performed for values corresponding to the scale value. The quantization scale can be derived based on the product of the quantization rate and the scale value. That is, the quantization scale can be derived by QP. It is also possible to apply the quantization scale to the transformation coefficients (C) and derive quantized transformation coefficients (C') based on this.
[0116] The inverse quantization process is the reverse process of the quantization process, where the quantized transformation coefficient (C') is converted to the quantization rate (Q). step By multiplying by ), the reconstructed transformation coefficient (C'') can be obtained. In this case, the level scale can be derived from the quantization parameters, and by applying the level scale to the quantized transformation coefficient (C''), the reconstructed transformation coefficient (C'') can be derived. The reconstructed transformation coefficient (C'') will differ slightly from the original transformation coefficient (C) due to losses during the transformation and / or quantization process. Therefore, the encoding device also performs inverse quantization, just as the decoding device.
[0117] Furthermore, adaptive frequency weighting quantization (CTS) techniques, which adjust the quantization intensity according to frequency, can be applied. Adaptive frequency weighting quantization is a method of applying different quantization intensities for different frequencies. Adaptive frequency weighting quantization can apply different frequency-specific quantization intensities using a predefined quantization scaling matrix. That is, the quantization / dequantization processes described above can be further performed based on the quantization scaling matrix. For example, different quantization scaling matrices can be used depending on whether the prediction mode applied to the current block is inter-prediction or intra-prediction to generate the current block size and / or the current block's residual signal. The quantization scaling matrix can be called a quantization matrix or scaling matrix. The quantization scaling matrix can be predefined. Also, for frequency adaptive scaling, frequency-specific quantization scale information for the quantization scaling matrix can be configured / encoded in an encoding device and signaled to a decoding device. The frequency-specific quantization scale information can be called quantization scaling information. Frequency-specific quantization scale information can include scaling list data. Based on the scaling list data, a (modified) quantization scaling matrix can be derived. The frequency-specific quantization scale information can also include present flag information indicating the presence or absence of the scaling list data. Alternatively, it may include information indicating whether the scaling list data is modified at a lower level (e.g., PPS or tile group header, etc.) if it is signaled at a higher level (e.g., SPS).
[0118] As mentioned above, scaling list data can be signaled to indicate the scaling matrix used for quantization / dequantization (frequency-based quantization).
[0119] Signaling support for default and user-defined scaling matrices is present in the HEVC standard and has now been adopted in the VVC standard. However, the VVC standard integrates additional support for signaling the following features:
[0120] - Three modes for the scaling matrix: OFF, DEFAULT, USER_DEFINED
[0121] - A larger size range for blocks (4x4 to 64x64 for Luma, 2x2 to 32x32 for Chroma)
[0122] - Rectangle transformation blocks (TBs)
[0123] -Dependent quantization
[0124] - Multiple Transform Selection (MTS)
[0125] - Transformations that zero out high-frequency coefficients.
[0126] - Intra sub-block partitioning (ISP)
[0127] - Intra-block copy (IBC) (also known as current picture referencing (CPR))
[0128] - DEFAULT scaling matrix for all TB sizes, default value is 16
[0129] It should be noted that the scaling matrix should not be applied to transform skips (TS) and secondary transforms (ST) for all sizes.
[0130] The following describes in detail the High Level Syntax (HSL) structure for supporting scaling lists in the VVC standard. First, a flag can be signaled via the Sequence Parameter Set (SPS) to indicate that a scaling list is available for the currently coded video sequence (CVS) being decoded. Next, if the aforementioned flag is available, an additional flag can be parsed in the SPS to indicate whether specific data exists in the scaling list. This can be shown as in Table 1.
[0131] Table 1 is an excerpt from SPS to illustrate the scaling list for CVS.
[0132] [Table 1]
[0133] The semantics of the syntax elements included in the SPS syntax in Table 1 can be shown in Table 2 below.
[0134] [Table 2]
[0135] Referring to Tables 1 and 2 above, the scaling_list_enabled_flag can be signaled from the SPS. For example, a value of 1 for scaling_list_enabled_flag indicates that the scaling list is used in the scaling process for the conversion coefficients, and a value of 0 for scaling_list_enabled_flag indicates that the scaling list is not used in the scaling process for the conversion coefficients. In this case, if the value of scaling_list_enabled_flag is 1, sps_scaling_list_data_present_flag can be further signaled from the SPS. For example, a value of sps_scaling_list_data_present_flag is 1 indicates that the scaling_list_data() syntax structure exists in the SPS, and a value of sps_scaling_list_data_present_flag is 0 indicates that the scaling_list_data() syntax structure does not exist in the SPS. If sps_scaling_list_data_present_flag does not exist, its value can be inferred to be 0.
[0136] Additionally, a flag (e.g., pps_scaling_list_data_present_flag) can be parsed first in the Picture Parameter Set (PPS). If this flag is available, scaling_list_data() can be parsed in the PPS. If scaling_list_data() exists first in the SPS and is parsed later in the PPS, the data in the PPS takes precedence over the data in the SPS. Table 3 below is an excerpt from the PPS to illustrate the scaling list data.
[0137] [Table 3]
[0138] The semantics of the syntax elements included in the PPS syntax in Table 3 can be shown in Table 4 below.
[0139] [Table 4]
[0140] Referring to Tables 3 and 4 above, the pps_scaling_list_data_present_flag can be signaled from the PPS. For example, if the value of pps_scaling_list_data_present_flag is 1, it can indicate that the scaling list data used for the picture referencing the PPS is derived based on the scaling list identified by the active SPS and the scaling list identified by the PPS. If the value of pps_scaling_list_data_present_flag is 0, it can indicate that the scaling list data used for the picture referencing the PPS is inferred to be the same as the scaling list identified by the active SPS. In this case, if the value of scaling_list_enabled_flag is 0, the value of pps_scaling_list_data_present_flag must be 0. If the value of scaling_list_enabled_flag is 1, the value of sps_scaling_list_data_present_flag is 0, and the value of pps_scaling_list_data_present_flag is 0, then the default scaling list data can be used to derive the Scaling Factor array, as described in Scaling List Data Semantics.
[0141] The scaling list can be defined in the VVC standard for the following quantization matrix sizes. This can be shown in Table 5 below. The range of supported quantization matrices has been extended in the HEVC standard to include 2x2 and 64x64, with 4x4, 8x8, 16x16, and 32x32.
[0142] [Table 5]
[0143] Table 5 defines the sizeId for all the quantization matrix sizes used. Using the combinations described above, matrixId can be assigned to different combinations of sizeId, coding unit prediction mode (CuPredMode), and color components. The CuPredModes that can be considered are inter, intra, and IBC (Intra Block Copy). Intra mode and IBC mode can be treated the same. Therefore, the same matrixId(s) can be shared for a given color component. The color components that can be considered are luma (Y) and two color components (Cb and Cr). The assigned matrixId can be shown in Table 6 below.
[0144] Table 6 shows the sizeId, prediction mode, and matrixId for each color component.
[0145] [Table 6]
[0146] Table 7 below shows an example of the syntax structure for scaling list data (e.g., scaling_list_data()).
[0147] [Table 7]
[0148] The semantics of the syntax elements included in the syntax in Table 7 can be shown in Table 8 below.
[0149] [Table 8-1]
[0150] [Table 8-2]
[0151] Referring to Tables 7 and 8, in order to extract scaling list data (e.g., scaling_list_data()), the scaling list data can be applied to 2x2 chroma components and 64x64 luma components for all sizeIds from 1 to 6 and matrixIds from 0 to 5. Next, a flag (e.g., scaling_list_pred_mode_flag) can be parsed to indicate whether the scaling list value is the same as the reference scaling list value. The reference scaling list is indicated by scaling_list_pred_matrix_id_delta[sizeId][matrixId]. However, if scaling_list_pred_mode_flag[sizeId][matrixId] is 1, the scaling list data can be explicitly signaled. If scaling_list_pred_matrix_id_delta is 0, the DEFAULT mode with default values can be used, as shown in Tables 9 to 12. If scaling_list_pred_matrix_id_delta is any other value, refMatrixId can be determined first, as shown in the semantics in Table 8 above.
[0152] In explicit signaling, i.e., in USER_DEFINED mode, the maximum number of coefficients to be signaled can be determined in advance. For quantization block sizes of 2x2, 4x4, and 8x8, all coefficients can be signaled. For sizes larger than 8x8, i.e., 16x16, 32x32, and 64x64, only 64 coefficients can be signaled. That is, the 8x8 base matrix is signaled, and the remaining coefficients can be upsampled from the base matrix.
[0153] Table 9 below shows an example of the default values for ScalingList[1][matrixId][i](i=0..3).
[0154] [Table 9]
[0155] Table 10 below shows an example of the default values for ScalingList[2][matrixId][i](i=0..15).
[0156] [Table 10]
[0157] Table 11 below shows an example of the default values for ScalingList[3..5][matrixId][i](i=0..63).
[0158] [Table 11]
[0159] Table 12 below shows an example of the default values for ScalingList[6][matrixId][i](i=0..63).
[0160] [Table 12]
[0161] As described above, the default scaling list data can be used to derive a scaling factor (Scaling Factor).
[0162] The scaling factor ScalingFactor[sizeId][sizeId][matrixId][x][y] of the 5D array (where x, y = 0..(1<<sizeId)-1) can represent an array of scaling factors depending on the variable sizeId shown in Table 5 above and the variable matrixId shown in Table 6 above.
[0163] The following Table 13 shows an example of deriving a scaling factor based on the quantization matrix size according to the aforementioned default scaling list.
[0164]
Table 13-1
[0165]
Table 13-2
[0166] For a quantization matrix of rectangular size, the scaling factor ScalingFactor[sizeIdW][sizeIdH][matrixId][x][y] of the 5D array (where x = 0..(1<<sizeIdW)-1, y = 0..(1<<sizeIdH)-1, sizeIdW!= sizeIdH) can represent an array of scaling factors depending on the variables sizeIdW and sizeIdH shown in Table 15 below and can be derived as shown in Table 14 below.
[0167]
Table 14
[0168] The quantization matrix of the quadrilateral size must be set to zero for samples that meet the following conditions.
[0169] -x > 32
[0170] -y > 32
[0171] - The decoded TU is not coded in the default transform mode, (1 << sizeIdW) == 32 and x > 16
[0172] - The decoded TU is not coded in the default transform mode, (1 << sizeIdH) == 32 and y > 16
[0173] The following Table 15 is an example showing sizeIdW and sizeIdH according to the quantization matrix size.
[0174]
Table 15
[0175] Also, as an example, the above-mentioned scaling list data (e.g., scaling_list_data()) can be described based on the syntax structure as shown in Table 16 below and the semantics as shown in Table 17 below. Based on the syntax elements included in the scaling list data (e.g., scaling_list_data()) disclosed in Table 16 and Table 17, as described above, scaling lists, scaling matrices, scaling factors, etc. can be derived, and this process is the same as or similar procedures can be applied to those in Table 5 to Table 15 mentioned above.
[0176]
Table 16
[0177]
Table 17-1
[0178] [Table 17-2]
[0179] In this document, we propose an efficient method for signaling scaling list data when applying adaptive frequency-based weighting techniques during the quantization / inverse quantization process.
[0180] Figure 6 illustrates the hierarchical structure for coded images / videos.
[0181] As shown in Figure 6, coded images / videos are divided into the VCL (video coding layer), which handles the decoding process and the images / videos themselves; the lower-level systems that transmit and store the coded information; and the NAL (network abstraction layer), which exists between the VCL and the lower-level systems and is responsible for network adaptation functions.
[0182] VCL can generate VCL data containing compressed image data (slice data), or generate parameter sets containing information such as Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS), or SEI (Supplemental Enhancement Information) messages additionally required during the image decoding process.
[0183] In NAL, a NAL unit can be generated by adding header information (NAL unit header) to an RBSP (Raw Byte Sequence Payload) generated by VCL. In this case, the RBSP refers to slice data, parameter sets, SEI messages, etc., generated by VCL. The NAL unit header can include NAL unit type information identified by the RBSP data contained in the NAL unit.
[0184] Furthermore, NAL units can be classified into VCL NAL units and Non-VCL NAL units based on the RBSP generated by VCL. VCL NAL units can represent NAL units that contain information about the image (slice data), while Non-VCL NAL units can represent NAL units that contain information necessary for decoding the image (parameter set or SEI message).
[0185] VCL NAL units and Non-VCL NAL units can be transmitted over a network with header information according to the data standards of the underlying system. For example, NAL units can be transformed into data formats of predetermined standards such as H.266 / VVC file format, RTP (Real-time Transport Protocol), and TS (Transport Stream) and transmitted over various networks.
[0186] As mentioned above, the NAL unit type can be identified by the RBSP data structure contained within the NAL unit, and information about such NAL unit types can be stored in the NAL unit header and signaled.
[0187] For example, NAL units can be broadly classified into VCL NAL unit types and Non-VCL NAL unit types depending on whether or not they contain information (slice data) about the image. VCL NAL unit types can be further classified by the nature and type of picture they contain, while Non-VCL NAL unit types can be further classified by the type of parameter set.
[0188] The following is an example of a NAL unit type identified by the type of parameter set included in the Non-VCL NAL unit type.
[0189] -APS (Adaptation Parameter Set) NAL unit: Type for NAL units that include APS
[0190] -DPS (Decoding Parameter Set) NAL unit: Type for NAL units including DPS
[0191] -VPS (Video Parameter Set) NAL unit: Type for NAL unit including VPS
[0192] -SPS (Sequence Parameter Set) NAL unit: Type for NAL units that include SPS
[0193] -PPS (Picture Parameter Set) NAL unit: Type for NAL units that include PPS
[0194] -PH (Picture header) NAL unit: Type for NAL units that include PH
[0195] The aforementioned NAL unit type has syntax information for the NAL unit type, and this syntax information can be stored in the NAL unit header and signaled. For example, the syntax information is nal_unit_type, and the NAL unit type can be identified by the nal_unit_type value.
[0196] On the other hand, as mentioned above, a single picture can contain multiple slices, and a single slice can contain a slice header and slice data. In this case, a single picture header can be added to each of the multiple slices (slice headers and slice data sets) within a single picture. A picture header (picture header syntax) can contain information / parameters that are commonly applicable to pictures. In this document, tile groups can be mixed with or substituted for slices or pictures. Also in this document, tile group headers can be mixed with or substituted for slice headers or picture headers.
[0197] A slice header (slice header syntax) can contain information / parameters that are commonly applicable to slices. An APS (APS syntax) or PPS (PPS syntax) can contain information / parameters that are commonly applicable to one or more slices or pictures. An SPS (SPS syntax) can contain information / parameters that are commonly applicable to one or more sequences. A VPS (VPS syntax) can contain information / parameters that are commonly applicable to multiple layers. A DPS (DPS syntax) can contain information / parameters that are commonly applicable to video in general. A DPS can contain information / parameters related to the concatenation of CVS (coded video sequence). In this document, High-level syntax (HLS) can include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, or slice header syntax.
[0198] In this document, the image / video information encoded from the encoding device to the decoding device and signaled in bitstream form may include not only partitioning-related information, intra / inter prediction information, residual information, and in-loop filtering information within the picture, but also information contained in the slice header, picture header, APS, PPS, SPS, VPS, and / or DPS. Furthermore, the image / video information may further include information from the NAL unit header.
[0199] On the other hand, the Adaptation Parameter Set (APS) is used in the VVC standard to transmit information for the Adaptive Loop Filter (ALF) and Luma Mapping with Chroma Scaling (LMCS) procedures. Furthermore, the APS has an extensible structure that allows it to be used to transmit other data structures (i.e., other syntax structures). Therefore, this document proposes a method for parsing / signaling scaling list data used for frequency-dependent weighting via the APS.
[0200] As mentioned above, the scaling list data is quantization scale information for frequency-dependent weighting that can be applied during the quantization / dequantization process, and is a list that associates the scale factor with each frequency index.
[0201] As one embodiment, Table 18 below shows an example of an APS (adaptation parameter set) structure used to transmit scaling list data.
[0202] [Table 18]
[0203] The semantics of the syntax elements included in the APS syntax in Table 18 can be shown in Table 19 below.
[0204] [Table 19]
[0205] Referring to Tables 18 and 19 above, the `adaptation_parameter_set_id` syntax element can be parsed / signaled in an APS. `adaptation_parameter_set_id` provides an identifier for the APS for reference to other syntax elements. That is, an APS can be identified based on the `adaptation_parameter_set_id` syntax element. The `adaptation_parameter_set_id` syntax element can be called APS ID information. APS can be shared between pictures and may differ in other tile groups within a picture.
[0206] Additionally, the `aps_params_type` syntax element can be parsed / signaled in APS. `aps_params_type` can indicate the type of APS parameter sent in APS, as shown in Table 18 below. The `aps_params_type` syntax element can be referred to as APS parameter type information or APS type information.
[0207] For example, Table 20 below is an example showing the types of APS parameters that can be transmitted via APS, and each APS parameter type can be shown corresponding to the value of aps_params_type.
[0208] [Table 20]
[0209] Referring to Table 20 above, aps_params_type is a syntax element for classifying the type of the APS. If the value of aps_params_type is 0, the APS type is ALF_APS, and the APS can carry ALF data, which may include ALF parameters for deriving filters / filter coefficients. If the value of aps_params_type is 1, the APS type is LMCS_APS, and the APS can carry LMCS data, which may include LMCS parameters for deriving LMCS models / bins / mapping indices. If the value of aps_params_type is 2, the APS type is SCALING_APS, and the APS can carry SCALING list data, which may include scaling list data parameters for deriving frequency-based quantization scaling matrix / scaling factor / scaling list values.
[0210] For example, as shown in Table 18 above, the aps_params_type syntax element can be parsed / signaled in APS. In this case, if the value of aps_params_type is 0 (i.e., aps_params_type indicates ALF_APS), the ALF data (i.e., alf_data()) can be parsed / signaled. Alternatively, if the value of aps_params_type is 1 (i.e., aps_params_type indicates LMCS_APS), the LMCS data (i.e., lmcs_data()) can be parsed / signaled. Alternatively, if the value of aps_params_type is 2 (i.e., aps_params_type indicates SCALING_APS), the scaling list data (i.e., scaling_list_data()) can be parsed / signaled.
[0211] Furthermore, referring to Tables 18 and 19 above, the aps_extension_flag syntax element can be parsed / signaled in APS. aps_extension_flag can indicate whether the APS extension data flag (aps_extension_data_flag) syntax element exists. aps_extension_flag can be used, for example, to provide extension points for later versions of the VVC standard. The aps_extension_flag syntax element can be called an APS extension flag. For example, a value of aps_extension_flag of 0 can indicate that the APS extension data flag (aps_extension_data_flag) does not exist in the APS RBSP syntax structure. Or, a value of aps_extension_flag of 1 can indicate that the APS extension data flag (aps_extension_data_flag) exists in the APS RBSP syntax structure.
[0212] The `aps_extension_data_flag` syntax element can be parsed / signaled based on the `aps_extension_flag` syntax element. The `aps_extension_data_flag` syntax element can be called an APS extension data flag. For example, if the value of `aps_extension_flag` is 1, then `aps_extension_data_flag` can be parsed / signaled, in which case `aps_extension_data_flag` can have any value.
[0213] As mentioned above, according to one embodiment of this document, scaling list data can be efficiently transported by assigning a data type (e.g., SCALING_APS) to represent scaling list data and parsing / signaling a syntax element (e.g., aps_params_type) that indicates the data type. In other words, according to one embodiment of this document, an APS structure that integrates scaling list data can be used.
[0214] On the other hand, the current VVC standard can be used to signal the availability of scaling list data (i.e., scaling_list_data()) based on whether a flag (i.e., sps_scaling_list_enabled_flag) exists in the SPS (Sequence Parameter Set). If the aforementioned flag (i.e., sps_scaling_list_enabled_flag) is enabled (i.e., it indicates that scaling list data is available and is 1 or true), then another flag (i.e., sps_scaling_list_data_present_flag) can be parsed. Also, if sps_scaling_list_data_present_flag is enabled (i.e., it indicates that scaling list data exists in the SPS and is 1 or true), then scaling list data (i.e., scaling_list_data()) can be parsed. In other words, the current VVC standard signals scaling list data using the SPS. In this case, since the SPS enables session negotiation and is generally transmitted out of band, it can be used during the decoding process, eliminating the need to transmit scaling list data with information related to determining the scaling factor of the transformed block. If the decoder transmits scaling list data in the SPS, the decoder needs to allocate a considerable amount of memory to store the information obtained from the scaling list data, and also needs to retain this information until it is used in the transformed block decoding. Therefore, such a process is unnecessary at the SPS level and is more effectively parsed / signaled at a lower level. Accordingly, this document proposes a hierarchical structure for effectively parsing / signaling scaling list data.
[0215] In one embodiment, scaling list data can be parsed / signaled using lower-level syntax such as PPS, tile group headers, slice headers, and / or other suitable headers, without having to parsed / signalize it using a higher-level syntax such as SPS.
[0216] For example, the SPS syntax can be modified as shown in Table 21 below. Table 19 below shows an example of SPS syntax for describing a scaling list for CVS.
[0217] [Table 21]
[0218] The semantics of the syntax elements included in the SPS syntax in Table 21 can be shown in Table 22 below.
[0219] [Table 22]
[0220] Referring to Tables 21 and 22 above, the scaling_list_enabled_flag syntax element can be parsed / signaled in SPS. The scaling_list_enabled_flag syntax element can indicate whether a scaling list is available based on whether its value is 0 or 1. For example, a value of 1 for scaling_list_enabled_flag indicates that the scaling list is used in the scaling process for the conversion coefficients, while a value of 0 for scaling_list_enabled_flag indicates that the scaling list is not used in the scaling process for the conversion coefficients.
[0221] In other words, the `scaling_list_enabled_flag` syntax element can be called a scaling list availability flag and can be signaled at the SPS (or SPS level). That is, based on the value of the `scaling_list_enabled_flag` signaled at the SPS level, it can be determined that scaling lists are essentially available for pictures in CVS that reference that SPS. And scaling lists can be obtained by signaling additional availability flags at lower levels than the SPS (e.g., PPS, tile group headers, slice headers, and / or other appropriate headers).
[0222] Furthermore, the sps_scaling_list_data_present_flag syntax element can be prevented from being parsed / signaled in SPS. In other words, by removing the sps_scaling_list_data_present_flag syntax element in SPS, this flag information can be prevented from being parsed / signaled. The sps_scaling_list_data_present_flag syntax element is flag information that indicates whether or not the syntax structure of the scaling list data exists in SPS, and the scaling list data specified by SPS can be parsed / signaled according to this flag information. However, by removing the sps_scaling_list_data_present_flag syntax element, the scaling list data can be prevented from being parsed / signaled at the SPS level.
[0223] As mentioned above, according to one embodiment of this document, the SPS can be configured to explicitly signal only the scaling list enabled flag (scaling_list_enabled_flag) without directly signaling the scaling list (scaling_list_data()). Subsequently, the scaling list (scaling_list_data()) can be individually parsed at lower-level syntax based on the enabled flag (scaling_list_enabled_flag) at the SPS. Therefore, according to one embodiment of this document, the scaling list data can be parsed / signaled by a hierarchical structure, thereby improving coding efficiency.
[0224] On the other hand, the existence and use of scaling list data depend on the presence of a tool enabling flag. Here, the tool enabling flag is information indicating whether to enable the tool in question, and can include, for example, the `scaling_list_enabled_flag` syntax element. That is, the `scaling_list_enabled_flag` syntax element can be used to indicate whether to enable the scaling list by indicating whether the scaling list data is available. However, this tool should have syntactic constraints on the decoder. That is, there should be a constraint flag that informs the decoder that this tool is not currently being used to decode a CVS (coded video sequence). Therefore, this document proposes a method for applying a constraint flag to scaling list data.
[0225] As one embodiment, Table 23 below shows an example of syntax for signaling scaling list data using restriction flags (e.g., general restriction information syntax).
[0226] [Table 23]
[0227] The semantics of the syntactic elements included in the syntax in Table 23 can be shown in Table 22 below.
[0228] [Table 24]
[0229] Referring to Tables 23 and 24 above, constraint flags can be parsed / signaled via general_constraint_info(). general_constraint_info() is called the general constraint information field or information about constraint flags. For example, the no_scaling_list_constraint_flag syntax element can be used as a constraint flag. Here, constraint flags can be used to specify conformance bitstream properties. For example, a value of 1 for the no_scaling_list_constraint_flag syntax element indicates a bitstream conformance requirement where scaling_list_enabled_flag should be specified as 0, and a value of 0 for the no_scaling_list_constraint_flag syntax element indicates no constraints.
[0230] On the other hand, as mentioned above, according to one embodiment of this document, scaling list data can be transmitted via a hierarchical structure. Accordingly, this document proposes a structure for scaling list data that can be parsed / signaled via slice headers. Here, the slice header may also be called a tile group header, or it may be mixed with or replaced by a picture header.
[0231] As one embodiment, Table 25 below shows an example of slice header syntax for signaling scaling list data.
[0232] [Table 25]
[0233] The semantics of the syntax elements included in the slice header syntax in Table 25 can be shown in Table 26 below.
[0234] [Table 26]
[0235] Referring to Tables 25 and 26 above, the slice_pic_parameter_set_id syntax element can be parsed / signaled in the slice header. The slice_pic_parameter_set_id syntax element can indicate an identifier for the PPS in use. That is, the slice_pic_parameter_set_id syntax element is information for identifying the PPS referenced in the slice, and can indicate the value of pps_pic_parameter_set_id. The value of slice_pic_parameter_set_id must be within the range of 0 to 63. The slice_pic_parameter_set_id syntax element can be called PPS identification information or PPS ID information referenced in the slice.
[0236] Additionally, the `slice_scaling_list_enabled_flag` syntax element can be parsed / signaled in the slice header. The `slice_scaling_list_enabled_flag` syntax element can indicate whether scaling lists are currently available in the slice. For example, a value of `slice_scaling_list_enabled_flag` of 1 indicates that scaling lists are currently available in the slice, while a value of 0 indicates that scaling lists are not currently available in the slice. Alternatively, if `slice_scaling_list_enabled_flag` is not present in the slice header, its value can be inferred to be 0.
[0237] In this case, the parsing feasibility of the slice_scaling_list_enabled_flag syntax element can be determined based on the scaling_list_enabled_flag syntax element signaled in the higher-level syntax (i.e., SPS). For example, if the value of scaling_list_enabled_flag signaled in SPS is 1 (i.e., it has been determined at the higher level that scaling list data is available), the slice header can be parsed to determine whether to use the scaling list to perform the scaling process in that slice.
[0238] Additionally, the slice_scaling_list_aps_id syntax element can be parsed / signaled in the slice header. The slice_scaling_list_aps_id syntax element can indicate an identifier for the APS referenced in that slice. That is, the slice_scaling_list_aps_id syntax element can indicate the ID information (adaptation_parameter_set_id) of the APS containing the scaling list data referenced in that slice. On the other hand, the TemporalId (i.e., TemporalID) of an APS NAL unit (i.e., an APS NAL unit containing scaling list data) having the same APS ID information (adaptation_parameter_set_id) as slice_scaling_list_aps_id must be smaller than or equal to the TemporalId (i.e., TemporalID) of the slice NAL unit being coded.
[0239] Furthermore, the parsing capability of the `slice_scaling_list_aps_id` syntax element can be determined based on the `slice_scaling_list_enabled_flag` syntax element. For example, if the value of `slice_scaling_list_aps_id` is 1 (i.e., if the slice header determines that the scaling list is enabled), then `slice_scaling_list_aps_id` can be parsed. Subsequently, scaling list data can be obtained from the APS indicated by the parsed `slice_scaling_list_aps_id`.
[0240] Furthermore, if multiple SCALING DATA APS (multiple APS containing scaling list data) with the same APS ID information (adaptation_parameter_set_id) are referenced by two or more slices within the same picture, the multiple SCALING DATA APS with the same APS ID information (adaptation_parameter_set_id) must contain the same content.
[0241] Furthermore, if the aforementioned syntax elements exist, the values of the slice header syntax elements slice_pic_parameter_set_id, slice_pic_order_cnt_lsb, and slice_temporal_mvp_enabled_flag must be the same for all slice headers in the coded picture.
[0242] As mentioned above, according to one embodiment of this document, a hierarchical structure can be used to efficiently signal scaling list data. Specifically, an availability flag (e.g., scaling_list_enabled_flag) indicating the availability of scaling list data is first signaled at the higher level (SPS syntax), and then additional availability flags (e.g., slice header, picture header, etc.) are signaled at the lower levels (e.g., slice header, picture header, etc.), thereby determining whether to use scaling list data at each lower level. Furthermore, APS ID information (e.g., slice_scaling_list_aps_id) referenced in the slice or tile group can be signaled via the lower levels (e.g., slice header, picture header, etc.), and scaling list data can be derived from the APS identified by the APS ID information.
[0243] In addition, this document can also be applied in signaling scaling list data by means of a hierarchical structure, such as the methods proposed in Table 23 and Table 24 described above, and can also transmit scaling list data through the structure of a slice header as shown in Table 25 below.
[0244] As an embodiment, Table 27 below shows an example of a slice header syntax for signaling scaling list data. Here, the slice header can also be called a tile group header, or can be mixed or substituted in the picture header.
[0245]
Table 27
[0246] The semantics of the syntax elements included in the slice header syntax of Table 27 above can be shown as in Table 28 below.
[0247]
Table 28
[0248] Referring to Table 27 and Table 28 above, the slice_pic_parameter_set_id syntax element can be parsed / signaled in the slice header. The slice_pic_parameter_set_id syntax element can indicate an identifier for the PPS in use. That is, the slice_pic_parameter_set_id syntax element is information for identifying the PPS referred to in the slice, and can indicate the value of pps_pic_parameter_set_id. The value of slice_pic_parameter_set_id must be within the range of 0 to 63. The slice_pic_parameter_set_id syntax element can be referred to as slice reference PPS identification information or PPS ID information.
[0249] Additionally, the slice header can be parsed / signaled with the `slice_scaling_list_aps_id` syntax element. The `slice_scaling_list_aps_id` syntax element can indicate an identifier for the APS referenced in that slice. Specifically, the `slice_scaling_list_aps_id` syntax element can indicate the ID information (adaptation_parameter_set_id) of the APS containing the scaling list data referenced in that slice. For example, the TemporalId (i.e., TemporalID) of an APS NAL unit (i.e., an APS NAL unit containing scaling list data) having the same APS ID information (adaptation_parameter_set_id) as `slice_scaling_list_aps_id` must be smaller than or equal to the TemporalId (i.e., TemporalID) of the coded slice NAL unit.
[0250] In this case, the parsing feasibility of the slice_scaling_list_aps_id syntax element can be determined based on the scaling_list_enabled_flag syntax element signaled in the higher-level syntax (i.e., SPS). For example, if the value of scaling_list_enabled_flag signaled in SPS is 1 (i.e., it has been determined at the higher level that scaling list data is available), then slice_scaling_list_aps_id can be parsed in the slice header. Subsequently, scaling list data can be obtained from the APS indicated by the parsed slice_scaling_list_aps_id.
[0251] In other words, according to this embodiment, an APS ID containing scaling list data can be parsed if the corresponding flag in the SPS (e.g., scaling_list_enabled_flag) is enabled. Therefore, as shown in Table 25 above, it is possible to parse the APS ID (e.g., slice_scaling_list_aps_id) information containing scaling list data referenced at the corresponding lower level (e.g., slice header or picture header) based on the scaling_list_enabled_flag syntax element signaled at the higher level syntax (i.e., SPS).
[0252] Furthermore, this document proposes a method for using multiple APSs to signal scaling list data. Below, one embodiment of this document describes how to efficiently signal multiple APS IDs containing scaling list data. This method is useful during bitstream merging.
[0253] As one embodiment, Table 29 below shows an example of slice header syntax for signaling scaling list data using multiple APSs. Here, the slice header may also be called a tile group header, or it may be mixed with or replaced by a picture header.
[0254] [Table 29]
[0255] The semantics of the syntax elements included in the slice header syntax in Table 29 can be shown in Table 30 below.
[0256] [Table 30]
[0257] Referring to Tables 29 and 30 above, the slice_pic_parameter_set_id syntax element can be parsed / signaled in the slice header. The slice_pic_parameter_set_id syntax element can indicate an identifier for the PPS in use. That is, the slice_pic_parameter_set_id syntax element is information for identifying the PPS referenced in the slice and can indicate the value of pps_pic_parameter_set_id. The value of slice_pic_parameter_set_id must be within the range of 0 to 63. The slice_pic_parameter_set_id syntax element can be called PPS identification information or PPS ID information referenced in the slice.
[0258] Additionally, the `slice_scaling_list_enabled_flag` syntax element can be parsed / signaled in the slice header. The `slice_scaling_list_enabled_flag` syntax element can indicate whether scaling lists are currently available in the slice. For example, a value of `slice_scaling_list_enabled_flag` of 1 indicates that scaling lists are currently available in the slice, while a value of 0 indicates that scaling lists are not currently available in the slice. Alternatively, if `slice_scaling_list_enabled_flag` is not present in the slice header, its value can be inferred to be 0.
[0259] In this case, the parsing feasibility of the slice_scaling_list_enabled_flag syntax element can be determined based on the scaling_list_enabled_flag syntax element signaled in the higher-level syntax (i.e., SPS). For example, if the value of scaling_list_enabled_flag signaled in SPS is 1 (i.e., it has been determined at the higher level that scaling list data is available), the slice header can be parsed to determine whether to use the scaling list to perform the scaling process in that slice.
[0260] Additionally, the num_scaling_list_aps_ids_minus1 syntax element can be parsed / signaled in the slice header. The num_scaling_list_aps_ids_minus1 syntax element indicates the number of APS containing the scaling list data referenced by that slice. For example, the number of APS is the value of the num_scaling_list_aps_ids_minus1 syntax element plus 1. The value of num_scaling_list_aps_ids_minus1 must be within the range of 0 to 7.
[0261] Here, the parsing eligibility of the num_scaling_list_aps_ids_minus1 syntax element can be determined based on the slice_scaling_list_enabled_flag syntax element. For example, if the value of slice_scaling_list_enabled_flag is 1 (i.e., it is determined that scaling list data is available in that slice), then num_scaling_list_aps_ids_minus1 can be parsed. In this case, the slice_scaling_list_aps_id[i] syntax element can be parsed / signaled based on the value of num_scaling_list_aps_ids_minus1.
[0262] That is, slice_scaling_list_aps_id[i] can indicate the identifier (adaptation_parameter_set_id) of the APS containing the i-th scaling list data (i.e., the i-th SCALING LIST APS). In other words, APS ID information can be signaled up to the number of APS indicated by the num_scaling_list_aps_ids_minus1 syntax element. On the other hand, the TemporalId (i.e., TemporalID) of an APS NAL unit (i.e., an APS NAL unit containing scaling list data) having the same APS ID information (adaptation_parameter_set_id) as slice_scaling_list_aps_id[i] must be less than or equal to the TemporalId (i.e., TemporalID) of the coded slice NAL unit.
[0263] Also, when multiple SCALING DATA APSs (multiple APSs including scaling list data) having the same value of APS ID information (adaptation_parameter_set_id) are referred to by two or more slices within the same picture, the multiple SCALING DATA APSs having the same value of APS ID information (adaptation_parameter_set_id) must contain the same content.
[0264] In addition, this document proposes a solution to avoid duplicate signaling when signaling scaling list data in a hierarchical structure. As one embodiment, signaling of scaling list data can be removed in the PPS (Picture Parameter Set). The scaling list data can be sufficiently signaled in the SPS, or APS, and / or other appropriate header sets.
[0265] As one embodiment, Table 31 below shows an example of the PPS syntax that does not signal scaling list data.
[0266]
Table 31
[0267] Table 32 below is an example showing the semantics for syntax elements (e.g., pps_scaling_list_data_present_flag) that can be removed to avoid duplicate signaling of scaling list data in the PPS syntax of Table 31.
[0268]
Table 32
[0269] By referring to Tables 31 and 32 above, information for signaling scaling list data in PPS, such as the pps_scaling_list_data_present_flag syntax element, can be removed. The pps_scaling_list_data_present_flag syntax element can indicate whether or not to signal scaling list data in PPS based on whether its value is 0 or 1. For example, if the value of the pps_scaling_list_data_present_flag syntax element is 1, it can indicate that the scaling list data used for a picture referencing PPS is derived based on the scaling list specified by the active SPS and the scaling list data specified by PPS. If the value of the pps_scaling_list_data_present_flag syntax element is 0, it can indicate that the scaling list data used for a picture referencing PPS is inferred to be the same as that specified by the active SPS. In other words, the pps_scaling_list_data_present_flag syntax element can represent information indicating whether or not scaling list data signaled from the PPS exists.
[0270] In other words, in this embodiment, in order to prevent redundant signaling of scaling list data, the PPS can be prevented from parsing / signaling the scaling list data by removing the pps_scaling_list_data_present_flag syntax element (i.e., by not signaling the pps_scaling_list_data_present_flag syntax element), as shown in Table 31.
[0271] This document also proposes a scheme for signaling scaling list matrices in APS. Existing schemes use three modes (i.e., OFF, DEFAULT, and USER_DEFINED modes). OFF mode indicates that scaling list data is not applied to the transformation block. DEFAULT mode indicates that fixed values are used to generate the scaling matrix. USER_DEFINED mode indicates that the scaling matrix is used based on block size, prediction mode, and color component. The total number of scaling matrices currently supported by VVC is 44, which is a significant increase from 28 in HEVC. Currently, scaling list data is signaled in SPS and can conditionally exist in PPS. By signaling scaling list data in APS, redundant signaling of the same data in SPS and PPS can be eliminated.
[0272] For example, the scaling matrix currently defined in VVC uses the `scaling_list_enabled_flag` to indicate whether it is available (enabled) in the SPS. If this flag indicates availability, the scaling list is used in the scaling process for the conversion coefficients; if this flag indicates it is not available, the scaling list is not used in the scaling process for the conversion coefficients (e.g., OFF mode). Additionally, the `sps_scaling_list_data_present_flag` may be parsed to indicate whether the scaling list data exists in the SPS. Besides signaling in the SPS, the scaling list data can also exist in the PPS. If `pps_scaling_list_data_present_flag` indicates availability in the PPS, the scaling list data can exist in the PPS. If the scaling list data exists in both the SPS and the PPS, the scaling list data from the PPS can be used in frames that reference the active PPS. If `scaling_list_enable_flag` indicates availability, but the scaling list data exists only in the SPS and not in the PPS, the scaling list data from the SPS can be referenced by the frame. The `scaling_list_enable_flag` indicates that it is available, and if scaling list data does not exist in the SPS or PPS, the DEFAULT mode may be used. The use of DEFAULT mode can also be signaled within the scaling list data itself. The USER DEFINED mode may be used if the scaling list data is explicitly signaled. The scaling factor for a given transformation block can be determined using information signaled in the scaling list. This suggests that the scaling list data should be signaled in the APS.
[0273] Currently, VVC uses scaling list data, and the scaling matrices supported by VVC are more extensive than those supported by HEVC. The scaling matrices supported by VVC allow for block sizes ranging from 4x4 to 64x64 for luma and from 2x2 to 32x32 for chroma. It also integrates rectangular transform block (TB) size, dependent quantization, multiple transform selection, large transform with zeroing out high-frequency coefficients, intra subblock partitioning (ISP), and intra block copy (IBC). Intra-block copy (IBC) and intra-coding modes can share the same scaling matrix.
[0274] Therefore, in USER_DEFINED mode, the number of matrices being signaled can be as follows:
[0275] ·MatrixType:30=2(2 for intra & IBC / inter)×3(Y / Cb / Cr components)×5(square TB size:from 4×4 to 64×64 for luma, from 2×2 to 32×32 for chroma)
[0276] ·MatrixType_DC:14=2(2 for intra & IBC / inter×1 for Y component)×3(TB size:16×16, 32×32, 64×64)+4(2 for intra & IBC / inter×2 for Cb / Cr components)×2(TB size:16×16, 32×32)
[0277] DC values can be coded separately for 16x16, 32x32, and 64x64 sized scaling matrices. If the Transform Block (TB) size is smaller than 8x8, all elements can be signaled within a single scaling matrix. If the Transform Block (TB) size is larger than or equal to 8x8, only 64 elements can be signaled as the base scaling matrix within a single 8x8 scaling matrix. To obtain a square matrix larger than 8x8, the base 8x8 matrix can be upsampled to the required size. When DEFAULT mode is used, the scaling matrix can be set to 16. Thus, VVC supports 44 distinct matrices, while HEVC supports only 28. Because the number of scaling matrices supported by VVC is wider than that of HEVC, APS can be used as a more practical choice for signaling scaling list data to avoid redundant signaling in SPS and / or PPS. This avoids unnecessary and redundant signaling of scaling list data.
[0278] To address the aforementioned issues, this document proposes a method for signaling the scaling matrix in APS. For this purpose, scaling list data can be signaled only in APS, without needing to be signaled in SPS or conditionally present in PPS. Additionally, the APS ID can be signaled in the slice header.
[0279] In one embodiment, a flag indicating whether scaling list data is available (e.g., scaling_list_enable_flag) is signaled in the SPS, and a flag indicating whether scaling list data exists (e.g., pps_scaling_list_data_present_flag) is removed in the PPS, thereby allowing the APS to signal scaling list data without signaling it. Additionally, the APS ID can be signaled in the slice header. In this case, if the value of scaling_list_enable_flag signaled in the SPS is 1 (i.e., indicating that scaling list data is available), and the APS ID is not signaled in the slice header, the DEFAULT scaling matrix may be used. Such an embodiment of this document can be implemented with the syntax and semantics shown in Tables 33 to 41.
[0280] Table 33 below shows an example of an APS structure used to signal scaling list data.
[0281] [Table 33]
[0282] The semantics of the syntax elements included in the APS syntax in Table 33 can be expressed as shown in Table 34.
[0283] [Table 34]
[0284] Referring to Tables 33 and 34 above, the `adaptation_parameter_set_id` syntax element can be parsed / signaled in an APS. `adaptation_parameter_set_id` provides an identifier for the APS for reference to other syntax elements. That is, an APS can be identified based on the `adaptation_parameter_set_id` syntax element. The `adaptation_parameter_set_id` syntax element can be called APS ID information. An APS can be shared between pictures and can be different in other slices within a picture.
[0285] Additionally, the `aps_params_type` syntax element can be parsed / signaled in APS. `aps_params_type` can represent the type of APS parameter sent from APS, as shown in Table 35 below. The `aps_params_type` syntax element can be referred to as APS parameter type information or APS type information.
[0286] For example, Table 35 below is an example of the types of APS parameters that can be transmitted via APS, and each APS parameter type can be represented in correspondence with the value of aps_params_type.
[0287] [Table 35]
[0288] Referring to Table 35 above, aps_params_type can be a syntax element for classifying the type of the APS. If the value of aps_params_type is 0, the APS type can be ALF_APS, and the APS can carry ALF data, which may include ALF parameters for deriving filters / filter coefficients. If the value of aps_params_type is 1, the APS type can be LMCS_APS, and the APS can carry LMCS data, which may include LMCS parameters for deriving LMCS models / bins / mapping indices. If the value of aps_params_type is 2, the APS type can be SCALING_APS, and the APS can carry SCALING list data, which may include scaling list data parameters for deriving frequency-based quantization scaling matrix / scaling, factor / scaling list values.
[0289] For example, as shown in Table 33 above, the aps_params_type syntax element can be parsed / signaled in APS. In this case, if the value of aps_params_type represents 0 (i.e., aps_params_type represents ALF_APS), the ALF data (i.e., alf_data()) can be parsed / signaled. Alternatively, if the value of aps_params_type represents 1 (i.e., aps_params_type represents LMCS_APS), the LMCS data (i.e., lmcs_data()) can be parsed / signaled. Alternatively, if the value of aps_params_type represents 2 (i.e., aps_params_type represents SCALING_APS), the scaling list data (i.e., scaling_list_data()) can be parsed / signaled.
[0290] Furthermore, referring to Tables 33 and 34 above, the aps_extension_flag syntax element can be parsed / signaled in APS. aps_extension_flag can indicate whether or not the APS extension data flag (aps_extension_data_flag) syntax element exists. aps_extension_flag can be used, for example, to provide an extension point for a later version of the VVC standard. The aps_extension_flag syntax element can be called an APS extension flag. For example, a value of aps_extension_flag of 0 can indicate that the APS extension data flag (aps_extension_data_flag) does not exist in the APS RBSP syntax structure. Or, a value of aps_extension_flag of 1 can indicate that the APS extension data flag (aps_extension_data_flag) exists in the APS RBSP syntax structure.
[0291] The `aps_extension_data_flag` syntax element can be parsed / signaled based on the `aps_extension_flag` syntax element. The `aps_extension_data_flag` syntax element can be called an APS extension data flag. For example, if the value of `aps_extension_flag` is 1, then `aps_extension_data_flag` can be parsed / signaled, in which case `aps_extension_data_flag` can have any value.
[0292] As described above, according to one embodiment of this document, scaling list data can be efficiently carried by assigning a data type (e.g., SCALING_APS) to represent scaling list data and parsing / signaling a syntax element (e.g., aps_params_type) that represents the data type. In other words, according to one embodiment of this document, an APS structure that integrates scaling list data can be used.
[0293] Furthermore, when signaling scaling list data in APS, SPS can signal whether or not scaling list data is available, and based on this, APS can parse / signal the scaling list data using the APS parameter type (e.g., aps_params_type). In one embodiment of this document, to avoid redundant signaling of scaling list data in higher-level syntax, SPS or PPS can not signal scaling list data by not parsing / signaling flag information indicating whether or not a scaling list data syntax structure (e.g., scaling_list_data()) exists in SPS or PPS. This can be achieved with the syntax and semantics shown in Tables 36 to 39 below.
[0294] For example, the SPS syntax can be modified as shown in Table 36. Table 36 shows an example of an SPS syntax structure that does not signal scaling list data in SPS.
[0295] [Table 36]
[0296] The semantics of the syntax elements included in the SPS syntax in Table 36 can be modified as shown in Table 37. For example, in Tables 36 and 37, some syntax elements (e.g., sps_scaling_list_data_present_flag) included in the SPS may be removed to avoid redundant signaling of scaling list data.
[0297] [Table 37]
[0298] Referring to Tables 36 and 37 above, the scaling_list_enabled_flag syntax element can be parsed / signaled in SPS. The scaling_list_enabled_flag syntax element can indicate whether a scaling list is available or not based on whether its value is 0 or 1. For example, a value of 1 for scaling_list_enabled_flag indicates that the scaling list is used in the scaling process for the conversion coefficients, while a value of 0 for scaling_list_enabled_flag indicates that the scaling list is not used in the scaling process for the conversion coefficients.
[0299] In other words, the `scaling_list_enabled_flag` syntax element can be called a scaling list availability flag and can be signaled at the SPS (or SPS level). To put it another way, based on the value of the `scaling_list_enabled_flag` signaled at the SPS level, it can be determined whether the scaling list is essentially available for pictures in CVS that reference that SPS. And, additional availability flags can be signaled at lower levels than the SPS (e.g., PPS, tile group headers, slice headers, and / or other appropriate headers) to obtain the scaling list.
[0300] Furthermore, the sps_scaling_list_data_present_flag syntax element can be prevented from being parsed / signaled in SPS. That is, by removing the sps_scaling_list_data_present_flag syntax element in SPS, this flag information can be prevented from being parsed / signaled. The sps_scaling_list_data_present_flag syntax element is flag information that indicates whether or not the syntax structure of the scaling list data exists in SPS, and the scaling list data specified by SPS can be parsed / signaled according to this flag information. However, by removing the sps_scaling_list_data_present_flag syntax element, it is possible to configure SPS to explicitly signal only the scaling list enabled flag (scaling_list_enabled_flag) without directly signaling the scaling list data.
[0301] Furthermore, signaling of scaling list data in PPS can be removed as shown in Table 38. For example, Table 38 shows an example of a PPS syntax structure that does not signal scaling list data in PPS.
[0302] [Table 38]
[0303] The semantics of the syntax elements included in the PPS syntax in Table 38 above can be modified as shown in Table 39. As an example, Table 39 shows the semantics for syntax elements (e.g., pps_scaling_list_data_present_flag) that can be removed in the PPS syntax to avoid redundant signaling of scaling list data.
[0304] [Table 39]
[0305] Referring to Tables 38 and 39 above, the pps_scaling_list_data_present_flag syntax element can be prevented from being parsed / signaled in PPS. That is, by removing the pps_scaling_list_data_present_flag syntax element in PPS, it is possible to configure it so that this flag information is not parsed / signaled. The pps_scaling_list_data_present_flag syntax element is flag information that indicates whether or not the syntax structure of the scaling list data exists in PPS, and the scaling list data specified by PPS can be parsed / signaled according to this flag information. However, by removing the pps_scaling_list_data_present_flag syntax element, the scaling list data can not be directly signaled at the PPS level.
[0306] As described above, by removing flag information indicating whether scaling list data exists at the SPS or PPS level, the SPS and PPS syntax can be configured so that scaling list data syntax is not directly signaled at the SPS or PPS level. Only the scaling list enabled flag (scaling_list_enabled_flag) is explicitly signaled at the SPS, and thereafter, scaling lists (scaling_list_data()) can be individually parsed at lower-level syntax (e.g., APS) based on the enabled flag (scaling_list_enabled_flag) at the SPS. Therefore, according to one embodiment of this document, scaling list data can be parsed / signaled by a hierarchical structure, thereby improving coding efficiency.
[0307] Furthermore, when signaling scaling list data with APS, the APS ID can be signaled in the slice header, the APS can be identified based on the APS ID obtained from the slice header, and the scaling list data can be parsed / signaled from the identified APS. Here, the slice header is described only as one example, and the slice header can be mixed with or replaced with various other headers such as tile group headers or picture headers.
[0308] For example, Table 40 below shows an example of slice header syntax that includes an APS ID syntax element for signaling scaling list data in APS.
[0309] [Table 40]
[0310] The semantics of the syntax elements included in the slice header syntax of Table 40 can be expressed as shown in Table 41.
[0311] [Table 41]
[0312] Referring to Tables 40 and 41 above, the slice_pic_parameter_set_id syntax element can be parsed / signaled in the slice header. The slice_pic_parameter_set_id syntax element can represent an identifier for the PPS in use. That is, the slice_pic_parameter_set_id syntax element is information for identifying the PPS referenced in the slice, and can represent the value of pps_pic_parameter_set_id. The value of slice_pic_parameter_set_id should be in the range of 0 to 63. The slice_pic_parameter_set_id syntax element can be said to be PPS identification information or PPSID information referenced in the slice.
[0313] Additionally, the `slice_scaling_list_present_flag` syntax element can be parsed / signaled in the slice header. The `slice_scaling_list_present_flag` syntax element can indicate whether or not a scaling list matrix currently exists for the slice. For example, a value of `slice_scaling_list_present_flag` of 1 indicates that a scaling list matrix currently exists for the slice, while a value of 0 indicates that the default scaling list data is used to derive the ScalingFactor array. Alternatively, if `slice_scaling_list_present_flag` is not present in the slice header, its value can be inferred to be 0.
[0314] In this case, the parsing feasibility of the slice_scaling_list_present_flag syntax element may be determined based on the scaling_list_enabled_flag syntax element signaled in the higher-level syntax (i.e., SPS). For example, if the value of scaling_list_enabled_flag signaled in SPS is 1 (i.e., it has been determined at the higher level that scaling list data is available), the slice header can be parsed to determine whether or not to use the scaling list to perform the scaling process in that slice.
[0315] Additionally, the slice_scaling_list_aps_id syntax element can be parsed / signaled in the slice header. The slice_scaling_list_aps_id syntax element can represent an identifier for the APS referenced in that slice. That is, the slice_scaling_list_aps_id syntax element can represent the ID information (adaptation_parameter_set_id) of the APS containing the scaling list data referenced in that slice. On the other hand, the TemporalId (i.e., Temporal ID) of an APS NAL unit (i.e., an APS NAL unit containing scaling list data) having the same APS ID information (adaptation_parameter_set_id) as slice_scaling_list_aps_id must be smaller than or equal to the TemporalId (i.e., Temporal ID) of the coded slice NAL unit.
[0316] Furthermore, the parsing feasibility of the `slice_scaling_list_aps_id` syntax element can be determined based on the `slice_scaling_list_present_flag` syntax element. For example, if the value of `slice_scaling_list_present_flag` is 1 (i.e., a scaling list exists in the slice header), then `slice_scaling_list_aps_id` can be parsed. Subsequently, the scaling list data can be obtained from the APS indicated by the parsed `slice_scaling_list_aps_id`.
[0317] Furthermore, if multiple SCALING DATA APS (multiple APS containing scaling list data) with the same APS ID information (adaptation_parameter_set_id) are referenced by two or more slices within the same picture, the multiple SCALING DATA APS with the same APS ID information (adaptation_parameter_set_id) must contain the same content.
[0318] As explained above in Tables 40 and 41, the APS ID can be signaled in the slice header, but this is just one example, and in this document, the APS ID can also be signaled in the picture header or tile group header, etc.
[0319] Furthermore, this document proposes a general approach to repositioning scaling list data within other header sets. As one embodiment, a general structure containing scaling list data within a header set is proposed. Currently, VVC uses APS as the appropriate header set, but it is also possible to encapsulate scaling list data, identified by the Nal Unit Type (NUT), within its own header set. This can be achieved as shown in Tables 42 and 43 below.
[0320] For example, Table 42 below shows an example of NAL unit types and their corresponding RBSP syntax structures. As mentioned above, the NAL unit type can be identified by the RBSP data structure contained within the NAL unit, and information regarding such NAL unit types can be stored and signaled in the NAL unit header.
[0321] [Table 42-1]
[0322] [Table 42-2]
[0323] As shown in Table 42 above, scaling list data can be defined as a single NAL unit type (e.g., SCALING_NUT), and a specific value (e.g., 21, or one of the reserved values not specified for the NAL unit type) can be specified for the NAL unit type SCALING_NUT. SCALING_NUT can be a type for a NAL unit that contains a scaling list data parameter set (e.g., Scaling_list_data_parameter set).
[0324] Furthermore, the availability of SCALING_NUT can be determined at a higher level than APS, PPS, and / or other appropriate headers, or at a lower level than other NAL unit types.
[0325] For example, Table 43 below shows the syntax for a scaling list data parameter set used to signal scaling list data.
[0326] [Table 43]
[0327] The semantics of the syntax elements included in the scaling list data parameter set syntax in Table 43 can be expressed as shown in Table 44.
[0328] [Table 44]
[0329] Referring to Tables 43 and 44 above, a scaling list data parameter set (e.g., scaling_list_data_parameter_set) can be a set of headers identified by a value of type NAL unit for SCALING_NUT (e.g., 21). The scaling_list_data_parameter_set_id syntax element can be parsed / signaled in the scaling list data parameter set. The scaling_list_data_parameter_set_id syntax element provides an identifier for the scaling list data for reference to other syntax elements. That is, a scaling list parameter set can be identified based on the scaling_list_data_parameter_set_id syntax element. The scaling_list_data_parameter_set_id syntax element can be called scaling list data parameter set ID information. A scaling list data parameter set can be shared between pictures and can be different in other slices within a picture.
[0330] The scaling list data (e.g., scaling_list_data syntax) can be parsed / signaled from the scaling list data parameter set identified by scaling_list_data_parameter_set_id.
[0331] Additionally, the `scaling_list_data_extension_flag` syntax element can be parsed / signaled in the scaling list data parameter set. The `scaling_list_data_extension_flag` syntax element can indicate whether or not the `scaling_list_data_extension_flag` syntax element exists in the scaling list data RBSP syntax structure. For example, a value of `scaling_list_data_extension_flag` of 1 indicates that the `scaling_list_data_extension_flag` syntax element exists in the scaling list data RBSP syntax structure. Alternatively, a value of `scaling_list_data_extension_flag` of 0 indicates that the `scaling_list_data_extension_flag` syntax element does not exist in the scaling list data RBSP syntax structure.
[0332] The `scaling_list_data_extension_data_flag` syntax element can be parsed / signaled based on the `scaling_list_data_extension_data_flag` syntax element. The `scaling_list_data_extension_data_flag` syntax element can be called an extension data flag for scaling list data. For example, if the value of `scaling_list_data_extension_flag` is 1, `scaling_list_data_extension_data_flag` can be parsed / signaled, in which case `scaling_list_data_extension_data_flag` can have any value.
[0333] As described above, according to one embodiment of this document, a single header set structure can be defined and used for scaling list data, and the header set for scaling list data can be specified for a NAL unit type (e.g., SCALING_NUT). In this case, the header set for scaling list data can be defined as a scaling list data parameter set (e.g., scaling_list_data_parameter_set), from which scaling list data can be obtained.
[0334] On the other hand, this document proposes a method for effectively coding the syntax element `scaling_lsit_pred_matrix_id_delta` included in scaling list data.
[0335] Currently, in the case of VVC, the `scaling_lsit_pred_matrix_id_delta` syntax element is coded using an unsigned integer 0th order Exp-Golomb-coded syntax element with the left bit first. However, to improve coding efficiency, one embodiment of this document proposes a method of coding `scaling_lsit_pred_matrix_id_delta` syntax elements having a range of 0 to 5 using a fixed-length code (e.g., u(3)), as shown in Table 45. In this case, it may be sufficient to use only 3 bits for coding to improve efficiency over the entire range.
[0336] For example, Table 45 below shows an example of a scaling list data syntax structure.
[0337] [Table 45]
[0338] The semantics of the syntax elements included in the scaling list data syntax of Table 45 can be represented as shown in Table 46.
[0339] [Table 46-1]
[0340] [Table 46-2]
[0341] As shown in Tables 45 and 46 above, the `scaling_list_pred_matrix_id_delta` syntax element can be parsed / signaled from the scaling list data syntax. The `scaling_list_pred_matrix_id_delta` syntax element can represent a reference scaling list used to derive the scaling list. In this case, `scaling_list_pred_matrix_id_delta` can be parsed using a fixed-length code (e.g., u(3)).
[0342] On the other hand, when signaling scaling list data with APS, it is possible to impose restrictions on the scaling list matrix. This document proposes a method for limiting the number of elements in the scaling list matrix. The method proposed in this document facilitates implementation and has the effect of limiting the worst-case memory requirement.
[0343] In one embodiment, the number of APS containing scaling list data (i.e., APS signaling scaling list data syntax) can be limited. To this end, the following constraints can be added. These constraints are intended to place holders, i.e., different values can be used.
[0344] For the sake of explanation, an APS containing scaling list data (i.e., an APS that signals scaling list data syntax) can be referred to as a SCALING LIST APS. In other words, as mentioned above, if the APS parameter type sent from the APS (e.g., aps_params_type) is of a type that represents scaling list data parameters (e.g., SCALING_APS), then the scaling list data sent from the APS can be referred to as a SCALING LIST APS.
[0345] For example, the total number of APS containing scaling list data (i.e., SCALING LIST APS) can be less than 3. Of course, other appropriate values can also be used. For example, an appropriate value within the range of 0 to 7 may be used. That is, the total number of APS containing scaling list data (i.e., SCALING LIST APS) can be determined to be within the range of 0 to 7.
[0346] Furthermore, for example, only one SCALING LIST APS (i.e., a SCALING LIST APS) per picture is allowed, containing scaling list data.
[0347] Table 47 below shows an example of syntax elements and their semantics that represent constraints for limiting APS that include scaling list data as described above.
[0348] [Table 47]
[0349] Referring to Table 47 above, the number of APS containing scaling list data (i.e., SCALING LIST APS) can be limited based on the syntax element (e.g., slice_scaling_list_aps_id) that represents the APS identification information (i.e., APS ID information) of the SCALING LIST APS.
[0350] For example, the syntax element slice_scaling_list_aps_id can represent the APS identification information (i.e., APS ID information) of the SCALING LIST APS referenced by the slice. In this case, the value of the syntax element slice_scaling_list_aps_id can be restricted to a specific value. For example, the value of the syntax element slice_scaling_list_aps_id can be restricted to a range of 0 to 3. This is just one example, and it can be restricted to other values. For example, the value of the syntax element slice_scaling_list_aps_id can be restricted to a range of 0 to 7.
[0351] Furthermore, for example, the TemporalId (i.e., Temporal ID) of a SCALING LIST APS NAL unit having an APS ID (adaptation_parameter_set_id) such as slice_lmcs_aps_id must be less than or equal to the TemporalId (i.e., Temporal ID) of the coded slice NAL unit.
[0352] Furthermore, for example, if multiple SCALING LIST APS with the same APS ID (adaptation_parameter_set_id) value are referenced by two or more slices on the same picture, then the multiple SCALING LIST APS with the same APS ID (adaptation_parameter_set_id) value must have the same content.
[0353] Furthermore, for example, only one SCALING LIST APS having the same APS ID (adaptation_parameter_set_id) and the same content should be referenced by one or more slices on the same picture. In other words, one or more slices within the same picture should reference the same APS containing scaling list data.
[0354] The following drawings were created to illustrate a specific example of this document. The names of specific devices, terms, and names (e.g., syntax / syntax element names) shown in the drawings are presented illustratively, and the technical features of this document are not limited to the specific names used in the following drawings.
[0355] Figures 7 and 8 schematically show an example of a video / image encoding method and related components according to the embodiments (etc.) described in this document.
[0356] The method disclosed in Figure 7 can be performed by the encoding device 200 disclosed in Figure 2. Specifically, step S700 in Figure 7 can be performed by the subtraction unit 231 disclosed in Figure 2, step S710 in Figure 7 can be performed by the conversion unit 232 disclosed in Figure 2, steps S720 to S730 in Figure 7 can be performed by the quantization unit 233 disclosed in Figure 2, and step S740 in Figure 7 can be performed by the entropy encoding unit 240 disclosed in Figure 2. Furthermore, the method disclosed in Figure 7 can be performed including the embodiments described above in this document. Therefore, in Figure 7, specific explanations of content that overlaps with the embodiments described above are omitted or simplified.
[0357] As shown in Figure 7, the encoding device can derive the current residual sample for the block (S700).
[0358] In one embodiment, the encoding device can first determine a prediction mode for the current block and derive prediction samples. For example, the encoding device can determine whether to perform interpretation or intrapretation on the current block, and can also determine a specific interpretation mode or a specific intrapretation mode based on the RD cost. The encoding device can perform predictions according to the determined prediction mode and derive prediction samples for the current block. At this time, various prediction methods disclosed in this document, such as interpretation or intrapretation, can be applied. The encoding device can also generate and encode information (e.g., prediction mode information) related to the prediction applied to the current block. The encoding device can then compare the original sample and the prediction sample for the current block to derive a residual sample.
[0359] The encoding device can derive conversion coefficients based on residual sampling (S710).
[0360] In one embodiment, the encoding device can derive conversion coefficients through a conversion process for residual samples. In this case, the encoding device can determine whether or not the conversion can be applied to the current block, taking coding efficiency into consideration. That is, the encoding device can determine whether or not the conversion is applied to the residual samples. For example, if the conversion is not applied to the residual samples, the encoding device can derive the residual samples as conversion coefficients. Alternatively, if the conversion is applied to the residual samples, the encoding device can perform the conversion on the residual samples and derive the conversion coefficients. In this case, the encoding device can generate conversion skip flag information based on whether or not the conversion is applied to the current block and encode it. The conversion skip flag information can be information indicating whether the conversion was applied to the current block or whether the conversion was skipped.
[0361] The encoding device can derive quantized conversion coefficients based on the conversion coefficients (S720).
[0362] In one embodiment, an encoding device can derive quantized conversion coefficients by applying a quantization process to the conversion coefficients. In this case, the encoding device can apply frequency-dependent weighting, which adjusts the quantization intensity with respect to frequency. In this case, the quantization process can be further performed based on frequency-dependent weighting scale values. The quantization scale values for frequency-dependent weighting can be derived using a scaling matrix. For example, an encoding / decoding device can use a predefined scaling matrix, and the encoding device can construct and encode frequency-dependent weighting scale information for the scaling matrix and signal this to the decoding device. The frequency-dependent weighting scale information may include scaling list data. A (modified) scaling matrix can be derived based on the scaling list data.
[0363] Furthermore, the encoding device, like the decoding device, can perform the inverse quantization process. In this case, the encoding device can derive a (corrected) scaling matrix based on the scaling list data, and then apply inverse quantization to the quantized transformation coefficients based on this matrix to derive the restored transformation coefficients. In this case, the restored transformation coefficients may differ from the original transformation coefficients due to losses in the transformation / quantization process.
[0364] Here, the scaling matrix can refer to the frequency-based quantization scaling matrix described above, and may be used interchangeably or interchangeably with quantization scaling matrix, quantization matrix, scaling matrix, scaling list, etc., for the sake of explanation, and is not limited to the specific names used in this embodiment.
[0365] In other words, the encoding device can further apply frequency-dependent weighting during the quantization process, and in this case, it can generate scaling list data as information about the scaling matrix. This process has been explained in detail using Tables 5 to 17 as examples, so in this embodiment, redundant content and detailed explanations will be omitted.
[0366] The encoding device can generate residual information based on the quantized conversion coefficients (S730).
[0367] Here, residual information is information generated through a transformation and / or quantization procedure, and can be information relating to quantized transformation coefficients, and may include, for example, information such as the values of the quantized transformation coefficients, positional information, transformation technique, transformation kernel, and quantization parameters.
[0368] Furthermore, if frequency-dependent weighting is applied to derive the quantized conversion coefficients during the quantization process, scaling list data for the quantized conversion coefficients may be generated. This scaling list data may include scaling list parameters used to derive the quantized conversion coefficients. In this case, the encoding device can generate scaling list data-related information, for example, an APS containing the scaling list data.
[0369] The encoding device can encode image information (or video information) (S740). Here, the image information may include the residual information. The image information may also include information related to the prediction used to derive the prediction sample (e.g., prediction mode information). The image information may also include information related to the scaling list data. In other words, the image information may include various information derived during the encoding process and may be encoded including such various information.
[0370] In one embodiment, the image information may include various information relating to the embodiments (etc.) described above in this document, and may include information disclosed in at least one of Tables 1 to 47 described above.
[0371] Furthermore, for example, image information can include an APS (adaptation parameter set). An APS can include APS ID information (APS identification information) and APS type information (type information of APS parameters). That is, an APS can be identified based on the APS ID information, and APS parameters corresponding to that type can be included in the APS based on the APS type information. For example, the APS type information can include the ALF type for the ALF (adaptive loop filter) parameter, the LMCS type for the LMCS (luma mapping with chroma scaling) parameter, and the scaling list type for the scaling list data parameter. This can be represented as shown in Table 35 above, and for example, if the value of the APS type information is 2, the APS type information can indicate that the APS includes a scaling list data parameter.
[0372] As an example, an APS can be configured as shown in Table 33 (or Table 18) above. The APS ID information (APS identification information) can be the adaptation_parameter_set_id described in Tables 33 and 34 (or Tables 18 and 19). The APS type information can be the aps_params_type described in Tables 33 to 35 (or Tables 18 to 20). For example, if the type information of the APS parameter (e.g., aps_params_type) is of type SCALING_APS, indicating that the APS contains scaling list data (or if the value of the type information of the APS parameter (e.g., aps_params_type) is 2, then the APS can contain scaling list data (e.g., scaling_list_data()). That is, based on the SCALING_APS type information indicating that the APS contains scaling list data, the encoding device can signal scaling list data (e.g., scaling_list_data()) via the APS. In other words, based on the APS type information (SCALING_APS type information), the APS may contain scaling list data. As mentioned above, the scaling list data may include scaling list parameters for deriving the scaling list / scaling matrix / scale factor used in the quantization / dequantization process. In other words, the scaling list data may include the syntax elements used to construct the scaling list.
[0373] Furthermore, as an example, APS ID information can have values within a specific range. For instance, the value of APS ID information can be within a specific range of 0 to 3, or 0 to 7. However, this is merely an example, and the range of values for APS ID information can be other. Also, the range of values for APS ID information can be determined based on APS type information (e.g., aps_params_type). For example, for an APS type information (e.g., SCALING_APS type) that indicates it is an APS related to scaling list data, the value of the APS ID information can be represented based on a syntax element (e.g., slice_scaling_list_aps_id) as shown in Table 47 above. For example, if the APS type information (e.g., aps_params_type) indicates that it is an APS containing scaling list data (e.g., SCALING_APS type), the value of the APS ID information can be within the range of 0 to 3, or 0 to 7. Here, slices within a single picture can refer to the same APS related to scaling list data. Alternatively, if the APS type information (e.g., aps_params_type) indicates that it is an APS related to ALF (e.g., ALF_APS type), the value of the APS ID information can be in the range of 0 to 7. Alternatively, if the APS type information (e.g., aps_params_type) indicates that it is an APS related to LMCS (e.g., LMCS_APS type), the value of the APS ID information can be in the range of 0 to 3.
[0374] Furthermore, image information can include, for example, an SPS (Sequence Parameter Set). The SPS can include a first availability flag information indicating whether the scaling list data is available or not. For example, the SPS can be configured as shown in Table 36 (or Table 21) above, and the first availability flag information can be the scaling_list_enabled_flag described in Tables 36 and 37 (or Tables 21 and 22). Additionally, the SPS can be configured so that this flag information is not parsed / signaled by removing the sps_scaling_list_data_present_flag syntax element. The sps_scaling_list_data_present_flag syntax element is flag information indicating whether or not the syntax structure of the scaling list data exists in the SPS, and the scaling list data specified by the SPS can be parsed / signaled according to this flag information. However, by removing the sps_scaling_list_data_present_flag syntax element, the SPS level can be configured to explicitly signal only the scaling list enabled flag (scaling_list_enabled_flag) without directly signaling the scaling list data.
[0375] Furthermore, image information can include a PPS (Picture ParameterSet), for example. As an example, a PPS can be configured as shown in Table 38 above, in which case the PPS can be configured so that it does not include availability flag information indicating the availability of scaling list data. That is, by removing the pps_scaling_list_data_present_flag syntax element from the PPS, this flag information can be configured not to be parsed / signaled. The pps_scaling_list_data_present_flag syntax element is flag information that indicates whether or not the syntax structure of the scaling list data exists in the PPS, and the scaling list data specified by the PPS can be parsed / signaled according to this flag information. However, by removing the pps_scaling_list_data_present_flag syntax element, the scaling list data can not be directly signaled at the PPS level.
[0376] In this case, a first availability flag (e.g., scaling_list_enabled_flag) indicating the availability of scaling list data may be signaled in SPS but not in PPS. Therefore, based on the first availability flag signaled in SPS (e.g., if the value of the first availability flag (e.g., scaling_list_enabled_flag) is 1 or true), the scaling list data included in APS can be retrieved.
[0377] Furthermore, for example, image information may include header information. The header information may be header information associated with the slice or picture containing the current block, and may include, for example, a picture header or a slice header. The header information may include scaling list data-related APS identification information. The scaling list data-related APS identification information included in the header information may represent APS ID information for an APS containing scaling list data. As an example, the scaling list data-related APS ID information included in the header information may be slice_scaling_list_aps_id as described in Tables 40 to 41 (or Tables 25 to 28), and may be identification information for an APS (containing scaling list data; i.e., SCALING LIST APS) referenced by the slice / picture containing the current block. In other words, an APS containing scaling list data can be identified based on the scaling list data-related APS ID information (e.g., slice_scaling_list_aps_id) in the header information.
[0378] In this case, whether or not the header information parses / signals APS ID information for APS containing scaling list data can be determined based on a first availability flag (scaling_list_enabled_flag) that is parsed / signaled in the SPS. For example, based on the first availability flag information indicating that the scaling list is available in the SPS (e.g., if the value of the first availability flag information (e.g., scaling_list_enabled_flag) is 1 or true), the header information may include APS ID information for APS containing scaling list data.
[0379] Furthermore, the header information may include, for example, a second availability flag indicating whether the scaling list data in the picture or slice is available. As an example, the second availability flag may be slice_scaling_list_present_flag (or slice_scaling_list_enabled_flag) as described in Tables 40 and 41 (or Tables 25 to 28).
[0380] In this case, whether the header information parses / signales the second availability flag information can be determined based on the first availability flag information (scaling_list_enabled_flag) that is parsed / signaled by the SPS. For example, based on the first availability flag information indicating that the scaling list is available in the SPS (e.g., if the value of the first availability flag information (e.g., scaling_list_enabled_flag) is 1 or true), the header information can include the second availability flag information. And based on the second availability flag information (e.g., if the value of the second availability flag information (e.g., slice_scaling_list_present_flag) is 1 or true), the header information can include the scaling list data-related APS ID information.
[0381] As an example, the encoding device can signal a second availability flag (e.g., slice_scaling_list_present_flag) via header information based on a first availability flag (e.g., scaling_list_enabled_flag) signaled by the SPS, as shown in Table 40 above. Subsequently, based on the second availability flag (e.g., slice_scaling_list_present_flag), the encoding device can signal APS ID information (e.g., slice_scaling_list_aps_id) for an APS containing scaling list data via header information. The encoding device can then signal scaling list data from the APS indicated by the signaled APS ID information (e.g., slice_scaling_list_aps_id).
[0382] As described above, the SPS and PPS syntaxes can be configured so that they do not directly signal the scaling list data syntax at the SPS or PPS level. For example, the SPS can explicitly signal only the scaling list enabled flag (scaling_list_enabled_flag), and thereafter, the scaling list (scaling_list_data()) can be individually parsed at lower-level syntax (e.g., APS) based on the enabled flag (scaling_list_enabled_flag) in the SPS. Therefore, according to one embodiment of this document, the scaling list data can be parsed / signaled by a hierarchical structure, thereby improving coding efficiency.
[0383] Image information containing the various types of information described above can be encoded and output in bitstream format. The bitstream can be transmitted to a decoding device via a network or (digital) storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.
[0384] Figures 9 and 10 schematically show an example of a video / image decoding method and related components according to the embodiments (etc.) of this document.
[0385] The method disclosed in Figure 9 can be performed by the decoding device 300 disclosed in Figure 3. Specifically, steps S900 to S910 in Figure 9 can be performed by the entropy decoding unit 310 disclosed in Figure 3, step S920 in Figure 9 can be performed by the inverse quantization unit 321 disclosed in Figure 3, step S930 in Figure 9 can be performed by the inverse transformation unit 321 disclosed in Figure 3, and step S940 in Figure 9 can be performed by the addition unit 340 disclosed in Figure 3. Furthermore, the method disclosed in Figure 9 can be performed including the embodiments described above in this document. Therefore, in Figure 9, specific explanations of content that overlaps with the embodiments described above are omitted or simplified.
[0386] As shown in Figure 9, the decoding device can receive image information (or video information) from the bitstream (S900).
[0387] In one embodiment, the decoding device can parse the bitstream to derive the information necessary for image restoration (or picture restoration) (e.g., video / image information). In this case, the image information may include residual information, which may include information such as the values of quantized transformation coefficients, position information, transformation technique, transformation kernel, and quantization parameters. The image information may also include information related to prediction (e.g., prediction mode information). Furthermore, the image information may include information related to scaling list data. In other words, the image information can include various information necessary in the decoding process and can be decoded based on coding methods such as exponential Golomb coding, CAVLC, or CABAC.
[0388] In one embodiment, the image information may include various information relating to the embodiments (etc.) described above in this document, and may include information disclosed in at least one of Tables 1 to 47 described above.
[0389] For example, image information can include an APS (adaptation parameter set). An APS can include APS ID information (APS identification information) and APS type information (type information of APS parameters). That is, an APS can be identified based on the APS ID information, and APS parameters corresponding to that type can be included in the APS based on the APS type information. For example, the APS type information can include the ALF type for the ALF (adaptive loop filter) parameter, the LMCS type for the LMCS (luma mapping with chroma scaling) parameter, and the scaling list type for the scaling list data parameter. This can be represented as shown in Table 35 above, and for example, if the value of the APS type information is 2, the APS type information can represent that the APS includes a scaling list data parameter.
[0390] As an example, an APS can be configured as shown in Table 33 (or Table 18) above. The APS ID information (APS identification information) can be the adaptation_parameter_set_id described in Tables 33 and 34 (or Tables 18 and 19). The APS type information can be the aps_params_type described in Tables 33 to 35 (or Tables 18 to 20). For example, if the type information of the APS parameter (e.g., aps_params_type) is of type SCALING_APS, indicating that it is an APS containing scaling list data (or if the value of the type information of the APS parameter (e.g., aps_params_type) is 2), then the APS can contain scaling list data (e.g., scaling_list_data()). That is, based on the SCALING_APS type information indicating that it is an APS containing scaling list data, the decoding device can obtain and parse the scaling list data (e.g., scaling_list_data()) via the APS. In other words, based on the APS type information (SCALING_APS type information), the APS may contain scaling list data. As mentioned above, the scaling list data may include scaling list parameters for deriving the scaling list / scaling matrix / scale factor used in the quantization / dequantization process. In other words, the scaling list data may include the syntax elements used to construct the scaling list.
[0391] Furthermore, as an example, APS ID information can have values within a specific range. For instance, the value of APS ID information can be within a specific range of 0 to 3, or 0 to 7. However, this is merely an example, and the range of values for APS ID information can be other. Also, the range of values for APS ID information can be determined based on APS type information (e.g., aps_params_type). For example, for an APS type information (e.g., SCALING_APS type) that indicates it is an APS related to scaling list data, the value of the APS ID information can be represented based on a syntax element (e.g., slice_scaling_list_aps_id) as shown in Table 47 above. For example, if the APS type information (e.g., aps_params_type) indicates that it is an APS containing scaling list data (e.g., SCALING_APS type), the value of the APS ID information can be within the range of 0 to 3, or 0 to 7. Here, slices within a single picture can refer to the same APS related to scaling list data. Alternatively, if the APS type information (e.g., aps_params_type) indicates that it is an APS related to ALF (e.g., ALF_APS type), the value of the APS ID information can be in the range of 0 to 7. Alternatively, if the APS type information (e.g., aps_params_type) indicates that it is an APS related to LMCS (e.g., LMCS_APS type), the value of the APS ID information can be in the range of 0 to 3.
[0392] Furthermore, image information can include, for example, an SPS (Sequence Parameter Set). The SPS can include a first availability flag information indicating whether the scaling list data is available or not. For example, the SPS can be configured as shown in Table 36 (or Table 21) above, and the first availability flag information can be the scaling_list_enabled_flag described in Tables 36 and 37 (or Tables 21 and 22). Additionally, the SPS can be configured so that this flag information is not parsed / signaled by removing the sps_scaling_list_data_present_flag syntax element. The sps_scaling_list_data_present_flag syntax element is flag information indicating whether or not the syntax structure of the scaling list data exists in the SPS, and the scaling list data specified by the SPS can be parsed / signaled according to this flag information. However, by removing the sps_scaling_list_data_present_flag syntax element, the SPS level can be configured to explicitly signal only the scaling list enabled flag (scaling_list_enabled_flag) without directly signaling the scaling list data.
[0393] Furthermore, image information can include a PPS (Picture ParameterSet), for example. As an example, a PPS can be configured as shown in Table 38 above, in which case the PPS can be configured so that it does not include availability flag information indicating the availability of scaling list data. That is, by removing the pps_scaling_list_data_present_flag syntax element from the PPS, this flag information can be configured not to be parsed / signaled. The pps_scaling_list_data_present_flag syntax element is flag information that indicates whether or not the syntax structure of the scaling list data exists in the PPS, and the scaling list data specified by the PPS can be parsed / signaled according to this flag information. However, by removing the pps_scaling_list_data_present_flag syntax element, the scaling list data can not be directly signaled at the PPS level.
[0394] In this case, a first availability flag (e.g., scaling_list_enabled_flag) indicating the availability of scaling list data may be signaled in SPS but not in PPS. Therefore, based on the first availability flag signaled in SPS (e.g., if the value of the first availability flag (e.g., scaling_list_enabled_flag) is 1 or true), the scaling list data included in APS can be retrieved.
[0395] Furthermore, for example, image information may include header information. The header information may be header information associated with the slice or picture containing the current block, and may include, for example, a picture header or a slice header. The header information may include scaling list data-related APS identification information. The scaling list data-related APS identification information included in the header information may represent APS ID information for an APS containing scaling list data. As an example, the scaling list data-related APS ID information included in the header information may be slice_scaling_list_aps_id as described in Tables 40 to 41 (or Tables 25 to 28), and may be identification information for an APS (containing scaling list data; i.e., SCALING LIST APS) referenced by the slice / picture containing the current block. In other words, an APS containing scaling list data can be identified based on the scaling list data-related APS ID information (e.g., slice_scaling_list_aps_id) in the header information. In other words, the decoding device can identify the APS based on the APS ID information (e.g., slice_scaling_list_aps_id) in the header information and obtain scaling list data from the APS.
[0396] In this case, whether or not the header information parses / signals APS ID information for APS containing scaling list data can be determined based on a first availability flag (scaling_list_enabled_flag) that is parsed / signaled in the SPS. For example, based on the first availability flag information indicating that the scaling list is available in the SPS (e.g., if the value of the first availability flag information (e.g., scaling_list_enabled_flag) is 1 or true), the header information may include APS ID information for APS containing scaling list data.
[0397] Furthermore, the header information may include, for example, a second availability flag indicating whether the scaling list data in the picture or slice is available. As an example, the second availability flag may be slice_scaling_list_present_flag (or slice_scaling_list_enabled_flag) as described in Tables 40 and 41 (or Tables 25 to 28).
[0398] In this case, whether the header information parses / signales the second availability flag information can be determined based on the first availability flag information (scaling_list_enabled_flag) that is parsed / signaled by the SPS. For example, based on the first availability flag information indicating that the scaling list is available in the SPS (e.g., if the value of the first availability flag information (e.g., scaling_list_enabled_flag) is 1 or true), the header information can include the second availability flag information. And based on the second availability flag information (e.g., if the value of the second availability flag information (e.g., slice_scaling_list_present_flag) is 1 or true), the header information can include APS ID information for the APS containing the scaling list data.
[0399] As an example, the decoding device can obtain second availability flag information (e.g., slice_scaling_list_present_flag) via header information based on first availability flag information (e.g., scaling_list_enabled_flag) signaled by the SPS, as shown in Table 40 (or Table 25) above. Then, based on the second availability flag information (e.g., slice_scaling_list_present_flag), it can obtain APS ID information (e.g., slice_scaling_list_aps_id) for an APS containing scaling list data via header information. The decoding device can then obtain scaling list data from the APS indicated by the APS ID information (e.g., slice_scaling_list_aps_id) obtained via header information.
[0400] As described above, the SPS and PPS syntaxes can be configured so that they do not directly signal the scaling list data syntax at the SPS or PPS level. For example, the SPS can explicitly signal only the scaling list enabled flag (scaling_list_enabled_flag), and thereafter, the scaling list (scaling_list_data()) can be individually parsed at lower-level syntax (e.g., APS) based on the enabled flag (scaling_list_enabled_flag) in the SPS. Therefore, according to one embodiment of this document, the scaling list data can be parsed / signaled by a hierarchical structure, thereby improving coding efficiency.
[0401] The decoding device can derive quantized conversion coefficients for the current block based on the residual information (S910).
[0402] In one embodiment, the decoding device can acquire residual information contained in the image information. As described above, the residual information may include information such as the values of quantized transformation coefficients, position information, transformation technique, transformation kernel, and quantization parameters. Based on the quantized transformation coefficient information contained in the residual information, the decoding device can derive the quantized transformation coefficients for the current block.
[0403] The decoding device can derive the transformation coefficients based on the quantized transformation coefficients (S920).
[0404] In one embodiment, a decoding device can derive conversion coefficients by applying an inverse quantization process to the quantized conversion coefficients. In this case, the decoding device can apply frequency-dependent weighting, which adjusts the quantization intensity with respect to frequency. In this case, the inverse quantization process can be further performed based on frequency-dependent weighting scale values. The quantization scale values for frequency-dependent weighting can be derived using a scaling matrix. For example, the decoding device can use a predefined scaling matrix, or it can use frequency-dependent weighting scale information for a scaling matrix signaled by an encoding device. The frequency-dependent weighting scale information may include scaling list data. A (modified) scaling matrix can be derived based on the scaling list data.
[0405] In other words, the decoding device can further apply frequency-dependent weighting when performing the inverse quantization process. In this case, the decoding device can derive the transformation coefficients by applying the inverse quantization process to the transformation coefficients quantized based on the scaling list data.
[0406] In one embodiment, the decoding device can acquire APS contained in image information and acquire scaling list data based on the APS type information contained in the APS. For example, the decoding device can acquire scaling list data contained in the APS based on the SCALING_APS type information, which indicates that the APS contains scaling list data. In this case, the decoding device can derive a scaling matrix based on the scaling list data, derive a scaling factor based on the scaling matrix, and derive conversion coefficients by applying inverse quantization based on the scaling factor. The process of performing scaling based on such scaling list data has been specifically explained with reference to Tables 5 to 17, so in this embodiment, redundant content and specific explanations will be omitted.
[0407] Furthermore, the decoding device can determine whether or not to apply frequency-based weighting during the inverse quantization process (i.e., whether or not to derive conversion coefficients using a (frequency-based quantization) scaling list during the inverse quantization process). For example, the decoding device can determine whether or not to use scaling list data based on a first availability flag obtained from the SPS included in the image information and / or a second availability flag obtained from the header information included in the image information. If it is decided to use scaling list data based on the first availability flag and / or the second availability flag information, the decoding device can obtain APS ID information for the APS containing the scaling list data via the header information, identify the APS based on the APS ID information, and obtain scaling list data from the identified APS.
[0408] The decoding device can derive residual samples based on the conversion coefficients (S930).
[0409] In one embodiment, the decoding device can perform an inverse transformation on the transformation coefficients for the current block to derive the residual sample of the current block. At this time, the decoding device can obtain information indicating whether or not to apply the inverse transformation to the current block (i.e., transformation skip flag information), and derive the residual sample based on this information (i.e., transformation skip flag information).
[0410] For example, if the inverse transformation is not applied to the transformation coefficients (i.e., the value of the transformation skip flag information for the current block is 1), the decoder can derive the transformation coefficients using the residual samples of the current block. Alternatively, if the inverse transformation is applied to the transformation coefficients (i.e., the value of the transformation skip flag information for the current block is 0), the decoder can perform the inverse transformation on the transformation coefficients to derive the residual samples of the current block.
[0411] The decoding device can generate a reconstructed sample based on the residual sample (S940).
[0412] In one embodiment, the decoding device can determine whether to perform interpretation or intrapretation for the current block based on prediction information (e.g., prediction mode information) included in the image information, and can perform prediction to derive a predicted sample for the current block based on the determination. The decoding device can then generate a reconstructed sample based on the predicted sample and the residual sample. At this time, the decoding device can immediately use the predicted sample as a reconstructed sample depending on the prediction mode, or it can generate a reconstructed sample by adding the residual sample to the predicted sample. Furthermore, a reconstructed block or reconstructed picture can be derived based on the reconstructed sample. Subsequently, as described above, the decoding device can apply deblocking filtering and / or in-loop filtering procedures such as the SAO procedure to the reconstructed picture as needed to improve subjective / objective image quality.
[0413] In the embodiments described above, the method is explained based on a flowchart in a series of steps or blocks, but the embodiments in this document are not limited to the order of the steps, and some steps may occur with other steps, in a different order, or simultaneously. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of this document.
[0414] The method described in this document can be implemented in software form, and the encoding and / or decoding devices described in this document can be included in devices that perform image processing, such as TVs, computers, smartphones, set-top boxes, and display devices.
[0415] When embodiments described in this document are implemented in software, the methods described above can be implemented by modules (processes, functions, etc.) that perform the functions described above. These modules are stored in memory and can be executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by a variety of well-known means. The processor may include an ASIC (application-specific integrated circuit), other chipsets, logic circuits, and / or data processing devices. The memory may include ROM (read-only memory), RAM (random access memory), flash memory, memory cards, storage media, and / or other storage devices. In other words, the embodiments described in this document can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored on a digital storage medium.
[0416] Furthermore, decoding and encoding devices to which this document applies may include multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video interaction devices, real-time communication devices such as video communications, mobile streaming devices, storage media, camcorders, video-on-demand (VoD) service providers, over-the-top (OTT) video equipment, internet streaming service providers, 3D video equipment, virtual reality (VR) equipment, argumentative reality (AR) equipment, image-phone video equipment, transportation terminals (e.g., vehicles (including autonomous vehicles), aircraft terminals, ship terminals, etc.), and medical video equipment, and can be used to process video signals or data signals. For example, over-the-top video equipment may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).
[0417] Furthermore, the processing methods to which this document applies can be produced in the form of programs executed on a computer and stored on a computer-readable recording medium. Multimedia data having the data structure described in this document can also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices that store data readable by a computer. The computer-readable recording medium can include, for example, Blu-ray discs (BDs), general-purpose serial buses (USB), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media implemented in the form of carrier waves (e.g., transmission over the Internet). Furthermore, bitstreams generated by encoding methods can be stored on a computer-readable recording medium or transmitted over a wireless network.
[0418] Furthermore, the embodiments described in this document can be implemented as a computer program product using program code, and the program code can be executed on a computer according to the embodiments described in this document. The program code can be stored on a computer-readable carrier.
[0419] Figure 11 shows an example of a content streaming system to which the embodiments disclosed in this document may be applied.
[0420] Referring to Figure 11, the content streaming system applicable to the embodiments of this document may broadly include an encoding server, a streaming server, a web server, a media storage facility, user equipment, and multimedia input devices.
[0421] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then transmitting this bitstream to the streaming server. In other cases, if a multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server can be omitted.
[0422] The bitstream can be generated by an encoding method or bitstream generation method applicable to the embodiments of this document, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0423] The streaming server transmits multimedia data to user devices based on user requests via a web server, and the web server acts as an intermediary to inform users about available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server controls the commands and responses between the devices within the content streaming system.
[0424] The streaming server can receive content from a media storage and / or encoding server. For example, if it starts receiving content from the encoding server, it can receive the content in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0425] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (such as smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, and digital signage.
[0426] Each server within the aforementioned content streaming system can be operated as a distributed server, in which case the data received by each server can be processed in a distributed manner.
[0427] The claims described in this document can be combined in various ways. For example, the technical features of the method claims in this document can be combined and realized in an apparatus, and the technical features of the apparatus claims in this document can be combined and realized in a method. Furthermore, the technical features of the method claims and the technical features of the apparatus claims in this document can be combined and realized in an apparatus, and the technical features of the method claims and the technical features of the apparatus claims in this document can be combined and realized in a method.
Claims
1. In an image decoding method performed by a decoding device, A step of obtaining image information including residual information and predictive information from a bitstream, The steps include: deriving quantized transformation coefficients for the current block based on the residual information; A step of deriving the transformation coefficients based on the quantized transformation coefficients, The steps include: deriving a residual sample based on the conversion coefficient; The steps include generating a reconstructed sample based on a predicted sample and the said residual sample, The aforementioned image information includes an APS (adaptation parameter set) containing scaling list data, The aforementioned APS includes APS ID information and APS type information, Based on the APS ID information, the APS is identified. The APS type information identifies whether the APS relates to an ALF (Adaptive loop filter), an LMCS (Luma mapping with chroma scaling), or the scaling list data. Based on the APS type information, the APS includes the scaling list data. Based on the APS type information which identifies the APS as an APS relating to the scaling list data, the APS ID information has a value within a specific range. The aforementioned specific range is predetermined with respect to the APS type information that identifies the APS as an APS relating to the scaling list data, A method wherein the specific range for the value of the APS ID information includes the range from 0 to 3.
2. In an image encoding method performed by an encoding device, The current step is to derive a residual sample for the block, The steps include: deriving conversion coefficients based on the aforementioned residual sample; The steps include: deriving quantized transformation coefficients by applying a quantization process to the aforementioned transformation coefficients; A step of generating residual information including information for the quantized conversion coefficients, The step of encoding image information including the said residual information, The aforementioned image information includes an APS (adaptation parameter set) containing scaling list data, The aforementioned APS includes APS ID information and APS type information, Based on the APS ID information, the APS is identified. The APS type information identifies whether the APS relates to an ALF (Adaptive loop filter), an LMCS (Luma mapping with chroma scaling), or the scaling list data. Based on the APS type information, the APS includes the scaling list data. Based on the APS type information which identifies the APS as an APS relating to the scaling list data, the APS ID information has a value within a specific range. The aforementioned specific range is predetermined with respect to the APS type information that identifies the APS as an APS relating to the scaling list data, A method wherein the specific range for the value of the APS ID information includes the range from 0 to 3.
3. Regarding methods for transmitting image-related data, A step of obtaining a bitstream relating to the image, wherein the bitstream is The current step is to derive a residual sample for the block, The steps include: deriving conversion coefficients based on the aforementioned residual sample; The steps include: deriving quantized transformation coefficients by applying a quantization process to the aforementioned transformation coefficients; A step of generating residual information including information for the quantized conversion coefficients, A step of encoding image information containing residual information, and a step of generating based on, The step of transmitting the data, which includes the bitstream, The aforementioned image information includes an APS (adaptation parameter set) containing scaling list data, The aforementioned APS includes APS ID information and APS type information, Based on the APS ID information, the APS is identified. The APS type information identifies whether the APS relates to an ALF (Adaptive loop filter), an LMCS (Luma mapping with chroma scaling), or the scaling list data. Based on the APS type information, the APS includes the scaling list data. Based on the APS type information which identifies the APS as an APS relating to the scaling list data, the APS ID information has a value within a specific range. The aforementioned specific range is predetermined with respect to the APS type information that identifies the APS as an APS relating to the scaling list data, A method wherein the specific range for the value of the APS ID information includes the range from 0 to 3.