Video or image coding based on luma mapping and chroma scaling

Through the video or image encoding method of brightness mapping and chromaticity scaling, the problem of high transmission and storage costs in high-resolution image/video compression is solved, the compression efficiency and visual quality are improved, and the encoding performance of resource consumption, adaptive parameter sets and dual-tree structure blocks are reduced.

CN114270823BActive Publication Date: 2025-08-15LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080058338.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-24
Filing Date
2020-06-24
Publication Date
2025-08-15
Estimated Expiration
2040-06-24

AI Technical Summary

Technical Problem

The prior art has problems with high transmission and storage costs in high resolution, high quality image/video compression, especially when dealing with special images such as virtual reality and holograms, the compression efficiency is insufficient.

Method used

Using video or image encoding methods based on luminance mapping and chromaticity scaling, by limiting the ID information range of LMCS APS and using linear mapping, the index derivation processing is simplified, a single chroma residual scaling factor, adaptive parameter set is applied, and the encoding of dual-tree structure blocks is supported, and resource consumption is reduced.

Benefits of technology

Improve image/video compression efficiency, improve subjective/objective visual quality, reduce resource requirements, support the encoding performance of dual-tree structure blocks, and reduce memory usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114270823B_ABST
    Figure CN114270823B_ABST
Patent Text Reader

Abstract

According to the disclosure of this document, image information is obtained from a bitstream, which includes prediction mode information and information associated with luma mapping and chroma scaling (LMCS), wherein the image information includes an LMCS adaptive parameter set (APS), and by limiting the range of APS ID information included in the LMCS APS, the memory used in the LMCS process can be reduced (limited).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of this document relates to video or image coding based on luma mapping and chroma scaling. Background Art

[0002] Recently, the demand for high-resolution, high-quality images / videos, such as 4K or 8K or higher ultra-high-definition (UHD) images / videos, has increased in various fields. As image / video data has high resolution and high quality, the amount of information or bit volume to be transmitted increases relative to existing image / video data. Therefore, transmitting image data using a medium such as an existing wired / wireless broadband line or an existing storage medium or storing image / video data using an existing storage medium increases transmission costs and storage costs.

[0003] In addition, interest in and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms have recently increased, and broadcasting of images / videos having characteristics different from real images (e.g., game images) has increased.

[0004] Therefore, very efficient image / video compression technology is required to effectively compress, transmit, store, and reproduce information of high-resolution, high-quality images / videos having various characteristics as described above.

[0005] In addition, luma mapping and chroma scaling (LMCS) processing is performed to improve compression efficiency and increase subjective / objective visual quality, and how to efficiently apply the LMCS process is discussed. Summary of the Invention

[0006] Technical Solution

[0007] According to an embodiment of this document, a method and apparatus for increasing image coding efficiency are provided.

[0008] According to the embodiments of this document, a high-efficiency filtering application method and device are provided.

[0009] According to the embodiments of this document, an efficient LMCS application method and apparatus are provided.

[0010] According to embodiments of this document, LMCS codewords (or ranges thereof) may be constrained.

[0011] According to embodiments of this document, a single chroma residual scaling factor that is directly signaled in the chroma scaling of the LMCS may be used.

[0012] According to an embodiment of the present document, a linear mapping (linear LMCS) may be used.

[0013] According to embodiments of the present document, information about the pivot point required for the linear mapping can be explicitly signaled.

[0014] According to embodiments of this document, a flexible number of bins may be used for brightness mapping.

[0015] According to the embodiments of this document, the index derivation process for inverse luma mapping and / or chroma residual scaling can be simplified.

[0016] According to an embodiment of this document, the LMCS process can be applied even when the luma block and the chroma block in one coding tree unit (CTU) have separate block tree structures (dual tree structure).

[0017] According to the implementation of this document, the number of LMCS APSs may be limited.

[0018] According to an embodiment of this document, the range of ID information included in the LMCS APS may be limited.

[0019] According to an embodiment of this document, a video / image decoding method performed by a decoding device is provided.

[0020] According to an embodiment of this document, a decoding device for performing video / image decoding is provided.

[0021] According to an embodiment of this document, a video / image encoding method performed by an encoding device is provided.

[0022] According to an embodiment of this document, there is provided an encoding device for performing video / image encoding.

[0023] According to an embodiment of this document, a computer-readable digital storage medium is provided, which stores encoded video / image information generated according to the video / image encoding method disclosed in at least one embodiment of this document.

[0024] According to an embodiment of this document, a computer-readable digital storage medium is provided, which stores encoded information or encoded video / image information that enables a decoding device to execute the video / image decoding method disclosed in at least one embodiment of this document.

[0025] Beneficial effects

[0026] According to the embodiments of this document, the overall image / video compression efficiency can be improved.

[0027] According to the embodiments of this document, subjective / objective visual quality can be improved through efficient filtering.

[0028] According to the embodiments of this document, LMCS processing for image / video encoding can be performed efficiently.

[0029] According to the embodiments of this document, the resources / costs (software or hardware) required for LMCS processing can be minimized.

[0030] According to the embodiments of this document, hardware implementation of LMCS processing can be facilitated.

[0031] According to embodiments of this document, the division operations required for the derivation of LMCS codewords in mapping (shaping) can be removed or minimized by constraining the LMCS codewords (or their range).

[0032] According to an embodiment of this document, a single chroma residual scaling factor can be used to remove the delay identified by the segment index.

[0033] According to the embodiments of this document, the chroma residual scaling process can be performed using linear mapping in LMCS without relying on (reconstruction of) the luma block, thus removing the delay in scaling.

[0034] According to the embodiments of this document, mapping efficiency in LMCS can be increased.

[0035] According to the embodiments of this document, by simplifying the index derivation process for inverse luma mapping and / or chroma residual scaling, the complexity of LMCS can be reduced, and thus the video / image coding efficiency can be increased.

[0036] According to the embodiments of this document, the LMCS process can be performed even on blocks with a dual tree structure, so the LMCS efficiency can be increased. In addition, the encoding performance (eg, objective / subjective picture quality) of blocks with a dual tree structure can be improved.

[0037] According to the embodiments of this document, since the number of LMCS APSs is limited, the complexity of the LMCS can be reduced, and thus fewer resources (eg, memory) can be consumed (used) in the LMCS.

[0038] According to the embodiments of this document, since the range of ID information included in the LMCS APS is limited, the memory used in the LMCS process can be reduced (limited). BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 An example of a video / image encoding system to which the embodiments of this document can be applied is shown.

[0040] Figure 2 This is a diagram schematically showing the configuration of a video / image encoding device to which an embodiment of this document can be applied.

[0041] Figure 3 This is a diagram schematically showing the configuration of a video / image decoding device to which an embodiment of this document can be applied.

[0042] Figure 4 An exemplary block tree structure is shown.

[0043] Figure 5 The layered structure of the encoded image / video is shown as an example.

[0044] Figure 6 The hierarchical structure of the CVS according to the embodiment of this document is exemplarily shown.

[0045] Figure 7 The hierarchical structure of the CVS according to the embodiment of this document is exemplarily shown.

[0046] Figure 8 The hierarchical structure of a CVS according to another embodiment of this document is exemplarily shown.

[0047] Figure 9 An exemplary LMCS structure according to an embodiment of this document is shown.

[0048] Figure 10 An LMCS structure according to another embodiment of this document is shown.

[0049] Figure 11 A graph representing an exemplary forward map is shown.

[0050] Figure 12 is a flowchart illustrating a method of deriving a chroma residual scaling index according to an embodiment of this document.

[0051] Figure 13 A linear fit of the pivot points according to an embodiment of this document is shown.

[0052] Figure 14 An example of linear shaping (or linear shaping, linear mapping) according to an embodiment of this document is shown.

[0053] Figure 15 An example of linear forward mapping in the embodiment of this document is shown.

[0054] Figure 16 An example of reverse forward mapping in the embodiment of this document is shown.

[0055] Figure 17 and Figure 18 An example of a video / image encoding method and related components according to an embodiment of this document is schematically shown.

[0056] Figure 19 and Figure 20 An example of an image / video decoding method and related components according to an embodiment of this document is schematically shown.

[0057] Figure 21 An example of a content streaming system to which the embodiments disclosed in this document can be applied is shown. DETAILED DESCRIPTION

[0058] This document can be modified in various forms, and specific embodiments thereof will be described and shown in the accompanying drawings. However, these embodiments are not intended to limit this document. The terms used in the following description are only used to describe specific embodiments and are not intended to limit this document. Singular expressions include plural expressions as long as they are clearly read differently. Terms such as "including" and "having" are intended to indicate the presence of features, quantities, steps, operations, elements, components, or combinations thereof used in the following description, so it should be understood that the possibility of the presence or addition of one or more different features, quantities, steps, operations, elements, components, or combinations thereof is not excluded.

[0059] In addition, the various configurations in the drawings described in this document are shown independently for the convenience of describing different features and functions, and do not mean that the various configurations are implemented as separate hardware or separate software. For example, two or more components among the various components can be combined to form a component, or a component can be divided into multiple components. Implementations in which the various components are integrated and / or separated are also included in the disclosure scope of this document.

[0060] Hereinafter, examples of the present embodiment will be described in detail with reference to the accompanying drawings. In addition, like reference numerals are used to indicate like elements throughout the drawings, and the same description about the like elements will be omitted.

[0061] Figure 1 An example of a video / image encoding system to which the embodiments of this document can be applied is shown.

[0062] Reference Figure 1 The video / image coding system may include a first device (source device) and a second device (receiver device). The source device may send the coded video / image information or data in the form of a file or stream to the receive device via a digital storage medium or a network.

[0063] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.

[0064] A video source may acquire video / images through a process of capturing, synthesizing, or generating video / images. A video source may include a video / image capture device and / or a video / image generation device. For example, a video / image capture device may include one or more cameras, a video / image archive including previously captured video / images, etc. A video / image generation device may include, for example, a computer, a tablet computer, and a smartphone, and may (electronically) generate video / images. For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capture process may be replaced by a process of generating relevant data.

[0065] An encoding device encodes input video / images. For compression and coding efficiency, the encoding device performs a series of processes such as prediction, transformation, and quantization. The encoded data (encoded video / image information) is output as a bitstream.

[0066] The transmitter can transmit the encoded image / image information or data, output as a bitstream, in the form of a file or stream to a receiver in a receiving device via a digital storage medium or network. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include components for generating a media file in a predetermined file format and can also include components for transmission via a broadcast / communication network. The receiver can receive / extract the bitstream and transmit the received bitstream to a decoding device.

[0067] The decoding device may decode a video / image by performing a series of processes such as dequantization, inverse transformation, and prediction corresponding to the operations of the encoding device.

[0068] The renderer may render the decoded video / image, and the rendered video / image may be displayed on a display.

[0069] This document relates to video / image coding. For example, the methods / implementations disclosed in this document may be applied to methods disclosed in the Versatile Video Coding (VVC) standard, the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the Audio Video Coding 2 standard (AVS2), or next-generation video / image coding standards (e.g., H.267, H.268, etc.).

[0070] This document proposes various embodiments of video / image coding, and unless otherwise specified, the above embodiments may also be performed in combination with each other.

[0071] In this document, video may refer to a series of images over time. A picture generally refers to a unit representing an image over a specific time range, and a slice / tile refers to a unit that constitutes a portion of a picture for coding purposes. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles. A tile may represent a rectangular area of a CTU row within a tile in a picture. A tile may be partitioned into multiple tiles, each of which may be composed of one or more CTU rows within the tile. A tile that is not partitioned into multiple tiles may also be referred to as a tile. Tile scanning may refer to a specific sequential ordering of CTUs within a partitioned picture, where CTUs may be ordered in a raster scan of CTUs within a tile, tiles within a tile may be ordered consecutively in a raster scan of tiles within a tile, and tiles within a picture may be ordered consecutively in a raster scan of tiles within a picture. A tile is a rectangular area of CTUs within a specific tile column and specific tile row in a picture. A tile column is a rectangular area of CTUs with a height equal to the height of the picture and a width specified by a syntax element in the picture parameter set. A tile row is a rectangular area of CTUs with a height specified by a syntax element in the picture parameter set and a width equal to the width of the picture. A tile scan is a specific sequential ordering of the CTUs that partition a picture, where the CTUs are ordered consecutively in a raster scan of the CTUs of the tiles, and the tiles in the picture are ordered consecutively in a raster scan of the tiles of the picture. A slice comprises an integer number of tiles of a picture that can be exclusively contained in a single NAL unit. A slice can consist of a consecutive sequence of multiple complete tiles or only complete tiles of one tile. In this document, tile group and slice may be used instead of each other. For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.

[0072] In addition, a picture can be divided into two or more sub-pictures. A sub-picture can be a rectangular area of one or more slices within a picture.

[0073] A pixel or picture element may refer to the smallest unit that constitutes a picture (or image). Furthermore, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component.

[0074] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to the area. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, a unit may be used interchangeably with terms such as a block or an area. In general, an M×N block may include M columns and N rows of samples (or sample arrays) or a set (or array) of transform coefficients. Alternatively, a sample may refer to a pixel value in the spatial domain, and when such a pixel value is transformed into the frequency domain, it may refer to a transform coefficient in the frequency domain.

[0075] In this document, "A or B" may mean "only A", "only B", or "both A and B". In other words, "A or B" in this document may be interpreted as "A and / or B". For example, in this document, "A, B or C (A, B or C)" means "only A", "only B", "only C", or "any combination of A, B, and C".

[0076] As used herein, a slash ( / ) or a comma (,) may mean "and / or." For example, "A / B" may mean "A and / or B." Thus, "A / B" may mean "only A," "only B," or "both A and B." For example, "A, B, C" may mean "A, B, or C."

[0077] In this document, “at least one of A and B” may mean “only A”, “only B”, or “both A and B”. In addition, in this document, the expression “at least one of A or B” or “at least one of A and / or B” may be interpreted the same as “at least one of A and B”.

[0078] In addition, in this document, "at least one of A, B, and C" means "only A", "only B", "only C", or "any combination of A, B, and C". In addition, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".

[0079] In addition, brackets used in this document may mean "for example." Specifically, when "prediction (intra-frame prediction)" is indicated, "intra-frame prediction" may be proposed as an example of "prediction." In other words, "prediction" in this document is not limited to "intra-frame prediction," and "intra-frame prediction" may be proposed as an example of "prediction." Furthermore, even when "prediction (i.e., intra-frame prediction)" is indicated, "intra-frame prediction" may be proposed as an example of "prediction."

[0080] Technical features described individually in one figure in this document can be implemented individually or simultaneously.

[0081] Figure 2Schematically illustrates the configuration of a video / image encoding device to which the embodiments of this document can be applied. Hereinafter, the so-called video encoding device may include an image encoding device.

[0082] Reference Figure 2 , the encoding device 200 includes an image segmenter 210, a predictor 220, a residual processor 230 and an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image segmenter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). In addition, the memory 270 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may also include the memory 270 as an internal / external component.

[0083] The image splitter 210 may split the input image (or picture or frame) input to the encoding device 200 into one or more processors. For example, a processor may be referred to as a coding unit (CU). In this case, the coding unit may be recursively split from a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree, binary tree, and ternary tree (QTBTTT) structure. For example, a coding unit may be split into multiple coding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quadtree structure may be applied first, followed by the binary tree structure and / or ternary structure. Alternatively, the binary tree structure may be applied first. The encoding process according to this document may be performed based on the final coding unit that is no longer split. In this case, the maximum coding unit may be used as the final coding unit based on coding efficiency according to image characteristics, or, if necessary, the coding unit may be recursively split into coding units of greater depth, and the coding unit of the optimal size may be used as the final coding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes (described later). As another example, the processor may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be split or divided from the final coding unit. The prediction unit may be a unit for sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0084] In some cases, the term "unit" may be used interchangeably with terms such as "block" or "region." In general, an M×N block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or pixel value, and may represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component. A sample may be used as a term corresponding to a picture (or image) of a pixel or picture element.

[0085] In the encoding device 200, a prediction signal (prediction block, prediction sample array) output from the inter predictor 221 or the intra predictor 222 is subtracted from an input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the transformer 232. In this case, as shown, the unit in the encoding device 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) may be referred to as the subtractor 231. The predictor may perform prediction on a block to be processed (hereinafter referred to as the current block) and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction based on the current block or CU. As described later in the description of each prediction mode, the predictor may generate various information related to the prediction (e.g., prediction mode information) and transmit the generated information to the entropy encoder 240. The information regarding the prediction may be encoded in the entropy encoder 240 and output in the form of a bitstream.

[0086] The intra-frame predictor 222 can predict the current block with reference to samples in the current picture. Depending on the prediction mode, the referenced samples may be located near the current block or may be spaced apart. In intra-frame prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. For example, the non-directional mode may include a DC mode and a planar mode. For example, depending on the level of detail of the prediction direction, the directional mode may include 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-frame predictor 222 may use the prediction mode applied to the neighboring blocks to determine the prediction mode applied to the current block.

[0087] The inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include information about the inter-frame prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. The reference picture containing the reference block and the reference picture containing the temporally neighboring block may be the same or different. Temporally neighboring blocks may be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and reference pictures containing temporally neighboring blocks may be referred to as collocated pictures (colPics). For example, the inter-frame predictor 221 may configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter-frame predictor 221 can use the motion information of the neighboring block as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of the neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be indicated by signaling the motion vector difference.

[0088] The predictor 220 may generate a prediction signal based on various prediction methods described below. For example, the predictor may apply not only intra prediction or inter prediction to predict a block, but also both intra prediction and inter prediction simultaneously. This may be referred to as combined inter and intra prediction (CIIP). In addition, the predictor may predict a block based on an intra block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode may be used for content image / video encoding, such as screen content coding (SCC), for games and the like. IBC essentially performs prediction within the current frame, but may be performed similarly to inter prediction, such that a reference block is derived within the current frame. That is, IBC may use at least one inter prediction technique described in this document. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the frame may be signaled based on information about the palette table and palette index.

[0089] The prediction signal generated by the predictor (including the inter-frame predictor 221 and / or the intra-frame predictor 222) can be used to generate a reconstructed signal or a residual signal. The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, when the relationship information between pixels is represented by a graph, GBT means a transform obtained from the graph. CNT refers to a transform generated based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process can be applied to square pixel blocks of the same size or to blocks of variable size other than square.

[0090] The quantizer 233 may quantize the transform coefficients and send them to the entropy encoder 240. The entropy encoder 240 may encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 233 may rearrange the block-type quantized transform coefficients into a one-dimensional vector form based on the coefficient scanning order and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. This information about the transform coefficients may be generated. The entropy encoder 240 may implement various encoding methods such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC). The entropy encoder 240 may encode information required for video / image reconstruction (e.g., syntax element values) in addition to the quantized transform coefficients, either together or separately. The encoded information (e.g., encoded video / image information) may be transmitted or stored in units of NALs (Network Abstraction Layers) in the form of a bitstream. The video / image information may also include information about various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Furthermore, the video / image information may also include general constraint information. In this document, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the above-described encoding process and included in a bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. The network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits the signal output from the entropy encoder 240 and / or a storage unit (not shown) that stores the signal may be included as internal / external components of the encoding device 200. Alternatively, the transmitter may be included in the entropy encoder 240.

[0091] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transformation to the quantized transform coefficients via the dequantizer 234 and the inverse transformer 235. The adder 250 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). If there is no residual for the block to be processed (for example, when skip mode is applied), the prediction block can be used as a reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. As described below, the generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture and can be used for inter-frame prediction of the next picture through filtering.

[0092] Furthermore, luma mapping with chroma scaling (LMCS) may be applied during picture encoding and / or reconstruction.

[0093] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 270 (specifically, the DPB of the memory 270). For example, the various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various types of information related to filtering and send the generated information to the entropy encoder 240, as described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoder 240 and output in the form of a bitstream.

[0094] The modified reconstructed picture transmitted to the memory 270 may be used as a reference picture in the inter predictor 221. When inter prediction is applied by the encoding apparatus, prediction mismatch between the encoding apparatus 200 and the decoding apparatus 300 may be avoided and encoding efficiency may be improved.

[0095] The DPB of the memory 270 may store a modified reconstructed picture used as a reference picture in the inter-frame predictor 221. The memory 270 may store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a reconstructed block in the picture. The stored motion information may be sent to the inter-frame predictor 221 and used as motion information of a spatially neighboring block or motion information of a temporally neighboring block. The memory 270 may store reconstructed samples of a reconstructed block in the current picture and may transmit the reconstructed samples to the intra-frame predictor 222.

[0096] Figure 3 is a schematic diagram showing the configuration of a video / image decoding device to which an embodiment of this document can be applied.

[0097] Reference Figure 3 , the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be configured by a hardware component (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware component may also include the memory 360 as an internal / external component.

[0098] When a bit stream including video / image information is input, the decoding apparatus 300 can reconstruct the bit stream corresponding to the bit stream in FIG. Figure 2 The video / image information is processed in the encoding device of the processing corresponding to the image. For example, the decoding device 300 can derive the unit / block based on the block segmentation related information obtained from the bit stream. The decoding device 300 can use the processor applied in the encoding device to perform decoding. Therefore, for example, the decoding processor can be a coding unit, and the coding unit can be divided from the coding tree unit or the maximum coding unit according to the quadtree structure, the binary tree structure and / or the ternary tree structure. One or more transform units can be derived from the coding unit. The reconstructed image signal decoded and output by the decoding device 300 can be reproduced by the reproduction device.

[0099] The decoding device 300 may receive Figure 2The received signal is output by the encoding device in the form of a bitstream, and the entropy decoder 310 can decode the received signal. For example, the entropy decoder 310 can parse the bitstream to derive information required for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may also include information about various parameter sets, such as the Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). In addition, the video / image information may also include general constraint information. The decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The signaled / received information and / or syntax elements described later in this document can be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 310 decodes the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and outputs the syntax elements required for image reconstruction and the quantized values of the residual transform coefficients. More specifically, the CABAC entropy decoding method receives bins corresponding to various syntax elements in a bitstream, determines a context model using information about the target syntax element to be decoded, information about the decoded target block, or information about symbols / cells decoded in a previous stage, and performs arithmetic decoding on the bins by predicting the probability of their occurrence based on the determined context model, generating symbols corresponding to the values of the respective syntax elements. In this case, after determining the context model, the CABAC entropy decoding method updates the context model by applying information from the decoded symbol / cell to the context model for the next symbol / cell. Information related to prediction, among the information decoded by the entropy decoder 310, can be provided to the predictor (inter-frame predictor 332 and intra-frame predictor 331), and residual values (i.e., quantized transform coefficients and related parameter information) entropy-decoded in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive a residual signal (residual block, residual sample, residual sample array). Furthermore, information related to filtering, among the information decoded by the entropy decoder 310, can be provided to the filter 350. In addition, a receiver (not shown) for receiving a signal output from the encoding device may also be configured as an internal / external element of the decoding device 300, or the receiver may be a component of the entropy decoder 310. In addition, the decoding device according to this document may be referred to as a video / image / picture decoding device, and the decoding device may be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.

[0100] The dequantizer 321 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 may rearrange the quantized transform coefficients in a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding device. The dequantizer 321 may dequantize the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.

[0101] The inverse transformer 322 inversely transforms the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0102] The predictor may perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the information on prediction output from the entropy decoder 310 and may determine a specific intra / inter prediction mode.

[0103] The predictor 330 can generate a prediction signal based on various prediction methods described below. For example, the predictor can apply not only intra-frame prediction or inter-frame prediction to predict a block, but also intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined inter-frame and intra-frame prediction (CIIP). In addition, the predictor can predict a block based on an intra-frame block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video encoding such as games, such as screen content coding (SCC). IBC basically performs prediction in the current picture, but can be performed similarly to inter-frame prediction so that a reference block is derived in the current picture. That is, IBC can use at least one inter-frame prediction technique described in this document. The palette mode can be considered an example of intra-frame coding or intra-frame prediction. When the palette mode is applied, the sample values within the picture can be signaled based on information about the palette table and palette index.

[0104] The intra-frame predictor 331 can predict the current block by referencing samples in the current picture. Depending on the prediction mode, the referenced samples may be located near the current block or spaced apart. In intra-frame prediction, the prediction modes may include multiple non-directional modes and multiple directional modes. The intra-frame predictor 331 can use the prediction modes applied to neighboring blocks to determine the prediction mode applied to the current block.

[0105] The inter-frame predictor 332 may derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include information on the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. For example, the inter-frame predictor 332 may configure a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index for the current block based on received candidate selection information. Inter-frame prediction may be performed based on various prediction modes, and information regarding the prediction may include information indicating the inter-frame prediction mode for the current block.

[0106] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter-frame predictor 332 and / or the intra-frame predictor 331). If there is no residual in the block to be processed, for example, when skip mode is applied, the prediction block can be used as the reconstructed block.

[0107] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next block to be processed in the current picture, may be output through filtering as described below, or may be used for inter prediction of the next picture.

[0108] In addition, luma mapping with chroma scaling (LMCS) can be applied in the picture decoding process.

[0109] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 360 (specifically, the DPB of the memory 360). For example, the various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0110] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store the motion information of the block from which the motion information in the current picture is derived (or decoded) and / or the motion information of the reconstructed block in the picture. The stored motion information can be sent to the inter-frame predictor 332 to be used as the motion information of the spatially adjacent block or the motion information of the temporally adjacent block. The memory 360 can store the reconstructed samples of the reconstructed block in the current picture and transmit the reconstructed samples to the intra-frame predictor 331.

[0111] In this document, the embodiments described in the filter 260, the inter-frame predictor 221, and the intra-frame predictor 222 of the encoding device 200 may be the same as or applied to the filter 350, the inter-frame predictor 332, and the intra-frame predictor 331 of the decoding device 300. This also applies to the inter-frame predictor 332 and the intra-frame predictor 331.

[0112] As described above, in video encoding, prediction is performed to increase compression efficiency. Thus, a prediction block including prediction samples of a current block (block to be encoded) can be generated. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived identically from the encoding device and the decoding device, and the encoding device decodes information about the residual between the original block and the prediction block (residual information) rather than the original sample values of the original block itself. By notifying the device with a signal, image coding efficiency can be increased. The decoding device can derive a residual block including residual samples based on the residual information, and generate a reconstructed block including reconstructed samples by summing the residual block and the prediction block, and generate a reconstructed picture including the reconstructed block.

[0113] Residual information can be generated through transformation and quantization. For example, the encoding device may derive a residual block between the original block and the prediction block, and perform transformation processing on the residual samples (residual sample array) included in the residual block to derive transform coefficients. Then, by performing quantization processing on the transform coefficients, the quantized transform coefficients are derived to signal the residual-related information to the decoding device (via the bitstream). Here, the residual information may include position information, transformation technology, transformation kernel and quantization parameters, value information of the quantized transform coefficients, etc. The decoding device may perform dequantization / inverse transformation processing based on the residual information and derive residual samples (or residual blocks). The decoding device may generate a reconstructed picture based on the prediction block and the residual block. The encoding device may also dequantize / inverse transform the quantized transform coefficients used for inter-frame prediction reference of a later picture to derive a residual block, and generate a reconstructed picture based on it.

[0114] In this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficient may be referred to as a transform coefficient. When transform / inverse transform is omitted, the transform coefficient may be referred to as a coefficient or a residual coefficient, or, for consistency of expression, may still be referred to as a transform coefficient.

[0115] In this document, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficients, and the information about the transform coefficients may be signaled via residual coding syntax. The transform coefficients may be derived based on the residual information (or information about the transform coefficients), and the scaled transform coefficients may be derived by inversely transforming (scaling) the transform coefficients. Residual samples may be derived based on the inverse transform (transform) of the scaled transform coefficients. This may also be applied / expressed in other parts of this document.

[0116] Intra-frame prediction may refer to the prediction of prediction samples of the current block based on reference samples in the picture to which the current block belongs (hereinafter referred to as the current picture). When intra-frame prediction is applied to the current block, neighboring reference samples to be used for intra-frame prediction of the current block may be derived. The neighboring reference samples of the current block may include samples adjacent to the left boundary of the current block of size nW×nH and a total of 2×nH samples adjacent to the lower left, samples adjacent to the upper boundary of the current block and a total of 2×nW samples adjacent to the upper right, and one sample adjacent to the upper left of the current block. Alternatively, the neighboring reference samples of the current block may include multiple upper neighboring samples and multiple left neighboring samples. In addition, the neighboring reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block of size nW×nH, a total of nW samples adjacent to the lower boundary of the current block, and one sample adjacent to the lower right of the current block.

[0117] However, some neighboring reference samples of the current block may not yet be decoded or available. In this case, the decoder can configure the neighboring reference samples to be used for prediction by replacing the unavailable samples with available samples. Alternatively, the neighboring reference samples to be used for prediction can be configured by interpolation of available samples.

[0118] When deriving neighboring reference samples, (i) the prediction sample may be derived based on an average or interpolation of neighboring reference samples of the current block, and (ii) the prediction sample may be derived based on reference samples existing in a specific (prediction) direction of the prediction sample among peripheral reference samples of the current block. The case of (i) may be referred to as a non-directional mode or a non-angular mode, and the case of (ii) may be referred to as a directional mode or an angular mode.

[0119] Furthermore, the prediction samples may be generated by interpolating between a second neighboring sample located in a direction opposite to the prediction direction of the intra prediction mode of the current block, among the neighboring reference samples. This may be referred to as linear interpolation intra prediction (LIP). Furthermore, a linear model may be used to generate chroma prediction samples based on luma samples. This may be referred to as LM mode.

[0120] In addition, a temporary prediction sample of the current block may be derived based on the filtered neighboring reference samples, and at least one reference sample derived according to the intra prediction mode among the existing neighboring reference samples (i.e., the unfiltered neighboring reference sample) and the temporary prediction sample may be weighted and summed to derive the prediction sample of the current block. This may be referred to as position-dependent intra prediction (PDPC).

[0121] In addition, the reference sample row with the highest prediction accuracy among the multiple reference sample rows adjacent to the current block can be selected to derive the prediction sample using the reference sample located in the prediction direction on the corresponding row. The reference sample row used herein can then be indicated (signaled) to the decoding device to perform intra-frame prediction encoding. The above situation can be referred to as multiple reference row (MRL) intra-frame prediction or MRL-based intra-frame prediction.

[0122] In addition, intra prediction can be performed by dividing the current block into vertical or horizontal sub-partitions based on the same intra prediction mode, and neighboring reference samples can be derived and used in sub-partition units. That is, in this case, the intra prediction mode of the current block is also applicable to the sub-partitions, and in some cases, intra prediction performance can be improved by deriving and using neighboring reference samples in sub-partition units. This prediction method may be referred to as intra sub-partition (ISP) or ISP-based intra prediction.

[0123] The above-mentioned intra-frame prediction methods may be separately referred to as intra-frame prediction types from the intra-frame prediction modes. Intra-frame prediction types may be referred to by various terms such as intra-frame prediction techniques or additional intra-frame prediction modes. For example, an intra-frame prediction type (or additional intra-frame prediction mode) may include at least one of the above-mentioned LIP, PDPC, MRL, and ISP. General intra-frame prediction methods other than specific intra-frame prediction types such as LIP, PDPC, MRL, or ISP may be referred to as normal intra-frame prediction types. Normal intra-frame prediction types may generally be applied when a specific intra-frame prediction type is not applied, and prediction may be performed based on the above-mentioned intra-frame prediction modes. In addition, post-filtering may be performed on the derived prediction samples as needed.

[0124] Specifically, the intra prediction process may include an intra prediction mode / type determination step, a neighboring reference sample derivation step, and a prediction sample derivation step based on the intra prediction mode / type. In addition, a post-filtering step may be performed on the derived prediction samples as needed.

[0125] When intra prediction is applied, the intra prediction mode applied to the current block may be determined using the intra prediction modes of neighboring blocks. For example, the decoding device may select one of the mpm candidates from an mpm list derived based on the intra prediction modes of neighboring blocks (e.g., left and / or above neighboring blocks) of the current block based on a received most probable mode (mpm) index, and select one of the remaining intra prediction modes not included in the mpm candidates (and planar mode) based on the remaining intra prediction mode information. The mpm list may be configured to include or exclude planar mode as a candidate. For example, if the mpm list includes planar mode as a candidate, the mpm list may have six candidates. If the mpm list does not include planar mode as a candidate, the mpm list may have three candidates. When the mpm list does not include planar mode as a candidate, a non-planar flag (e.g., intra_luma_not_planar_flag) may be signaled to indicate whether the intra prediction mode of the current block is not planar mode. For example, the mpm flag may be signaled first, and when the value of the mpm flag is 1, the mpm index and non-planar flag may be signaled. In addition, the mpm index may be signaled when the value of the non-planar flag is 1. Here, the mpm list is configured not to include the planar mode as a candidate without first signaling the non-planar flag to first check whether it is the planar mode, because the planar mode is always regarded as the mpm.

[0126] For example, whether the intra prediction mode applied to the current block is among the mpm candidates (and planar mode) or among the remaining modes may be indicated based on an mpm flag (e.g., intra_luma_mpm_flag). A value of 1 for the mpm flag may indicate that the intra prediction mode of the current block is among the mpm candidates (and planar mode), and a value of 0 for the mpm flag may indicate that the intra prediction mode of the current block is not among the mpm candidates (and planar mode). A value of 0 for a non-planar flag (e.g., intra_luma_not_planar_flag) may indicate that the intra prediction mode of the current block is planar mode, and a value of 1 for the non-planar flag value may indicate that the intra prediction mode of the current block is not planar mode. The mpm index may be signaled in the form of an mpm_idx or intra_luma_mpm_idx syntax element, and the remaining intra prediction mode information may be signaled in the form of a rem_intra_luma_pred_mode or intra_luma_mpm_remainder syntax element. For example, the remaining intra prediction mode information may index the remaining intra prediction modes that are not included in the mpm candidates (and planar modes) among all intra prediction modes in the order of prediction mode numbers to indicate one of them. The intra prediction mode may be an intra prediction mode of a luminance component (sample). Below, the intra prediction mode information may include at least one of an mpm flag (e.g., intra_luma_mpm_flag), a non-planar flag (e.g., intra_luma_not_planar_flag), an mpm index (e.g., mpm_idx or intra_luma_mpm_idx), and remaining intra prediction mode information (rem_intra_luma_pred_mode or intra_luma_mpm_remainder). In this document, the MPM list may be referred to by various terms such as MPM candidate list and candModeList. When MIP is applied to the current block, a separate mpm flag (eg, intra_mip_mpm_flag), an mpm index (eg, intra_mip_mpm_idx), and remaining intra prediction mode information (eg, intra_mip_mpm_remainder) for the MIP may be signaled, and the non-planar flag may not be signaled.

[0127] In other words, when block splitting is performed on an image, the current block to be encoded and the neighboring blocks typically have similar image characteristics. Therefore, there is a high probability that the current block and the neighboring blocks have the same or similar intra-prediction modes. Therefore, the encoder can use the intra-prediction modes of the neighboring blocks to encode the intra-prediction mode of the current block.

[0128] For example, the encoder / decoder may configure a list of most probable modes (MPMs) for the current block. The MPM list may also be referred to as an MPM candidate list. Herein, MPM may refer to a mode that improves encoding efficiency by considering the similarity between the current block and adjacent blocks in intra-frame prediction mode encoding. As described above, the MPM list may be configured to include a planar mode, or may be configured not to include a planar mode. For example, when the MPM list includes a planar mode, the number of candidates in the MPM list may be 6. Also, if the MPM list does not include a planar mode, the number of candidates in the MPM list may be 5.

[0129] The encoder / decoder can be configured with an MPM list consisting of 5 or 6 MPMs.

[0130] To configure the MPM list, three types of modes may be considered: default intra mode, neighbor intra mode, and derived intra mode.

[0131] For the neighboring intra mode, two neighboring blocks may be considered, ie, the left neighboring block and the above neighboring block.

[0132] As described above, if the MPM list is configured not to include the planar mode, the planar mode is excluded from the list and the number of MPM list candidates may be set to five.

[0133] In addition, the non-directional mode (or non-angular mode) among the intra prediction modes may include a DC mode based on an average of neighboring reference samples of the current block or a planar mode based on interpolation.

[0134] When inter-frame prediction is applied, the predictor of the encoding / decoding device can derive prediction samples by performing inter-frame prediction on a block-by-block basis. Inter-frame prediction can be a prediction derived in a manner that depends on data elements (e.g., sample values or motion information) from a picture other than the current picture. When inter-frame prediction is applied to the current block, the prediction block (prediction sample array) of the current block can be derived based on a reference block (reference sample array) specified by a motion vector in a reference picture indicated by a reference picture index. To reduce the amount of motion information transmitted in inter-frame prediction mode, the motion information of the current block can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information on the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in a reference picture. The reference picture comprising the reference block and the reference picture comprising the temporally neighboring blocks can be the same or different. Temporally neighboring blocks may be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and a reference picture including temporally neighboring blocks may be referred to as a collocated picture (colPic). For example, a motion information candidate list may be configured based on neighboring blocks of the current block, and a flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled. Inter-frame prediction may be performed based on various prediction modes. For example, in skip mode and merge mode, the motion information of the current block may be the same as that of the neighboring blocks. In skip mode, unlike merge mode, a residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be signaled. In this case, the motion vector of the current block may be derived using the sum of the motion vector predictor and the motion vector difference.

[0135] Depending on the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.), the motion information may include L0 motion information and / or L1 motion information. A motion vector in the L0 direction may be referred to as an L0 motion vector or MVL0, and a motion vector in the L1 direction may be referred to as an L1 motion vector or MVL1. Prediction based on the L0 motion vector may be referred to as L0 prediction, prediction based on the L1 motion vector may be referred to as L1 prediction, and prediction based on both the L0 motion vector and the L1 motion vector may be referred to as bi-prediction. Here, the L0 motion vector may indicate a motion vector associated with reference picture list L0 (L0), and the L1 motion vector may indicate a motion vector associated with reference picture list L1 (L1). Reference picture list L0 may include pictures that are earlier than the current picture in output order as reference pictures, and reference picture list L1 may include pictures that are later than the current picture in output order. The previous picture may be referred to as a forward (reference) picture, and the subsequent picture may be referred to as a backward (reference) picture. Reference picture list L0 may also include pictures that are later than the current picture in output order as reference pictures. In this case, the previous picture may be indexed first in reference picture list L0, and the subsequent picture may be indexed later. Reference picture list L1 may also include previous pictures that are earlier than the current picture in output order as reference pictures. In this case, the subsequent picture may be indexed first in reference picture list 1, and the previous picture may be indexed later. The output order may correspond to the picture order count (POC) order.

[0136] Figure 4 An exemplary block tree structure is shown. Figure 4 It is exemplarily shown that a CTU is divided into multiple CUs based on a quadtree and a nested multi-type tree structure.

[0137] The bold block edges represent quadtree partitioning, and the remaining edges represent multi-type tree partitioning. Quadtree partitioning of nested multi-type trees can provide a content-adaptive coding tree structure. A CU may correspond to a coding block CB. In addition, a CU may include a coding block of luma samples and two coding blocks of corresponding chroma samples. The size of a CU may be as large as the size of a CTU, or may be as small as 4×4 of a luma sample unit. For example, in the case of a 4:2:0 color format (or chroma format), the maximum chroma CB size may be 64×64, and the minimum chroma CB size may be 2×2.

[0138] In this document, for example, the maximum allowed luma TB size may be 64 × 64, and the maximum allowed chroma TB size may be 32 × 32. If the width or height of the CB split according to the tree structure is larger than the maximum transform width or height, the corresponding CB may be automatically (or implicitly) split until the TB size restrictions in the horizontal and vertical directions are met.

[0139] In this document, the coding tree scheme can support separate block tree structures for luma (component) blocks and chroma (component) blocks. Blocks with separate block tree structures can be blocks encoded using separate trees. The case where the luma block and chroma block in a CTU have the same block tree structure can be represented as a single tree structure (SINGLE_TREE). The case where the luma block and chroma block in a CTU have separate block tree structures can be represented as a dual tree structure (DUAL_TREE). In this case, the block tree type of the luma component can be called DUAL_TREE_LUMA, and the block tree type of the chroma component can be called DUAL_TREE_CHROMA. Blocks with a dual tree structure can be blocks encoded using a dual tree. For P and B slices / tile groups, the luma CTB and chroma CTB in a CTU can be restricted to have the same coding tree structure. However, for I slices / tile groups, the luma block and chroma block can have separate block tree structures. If a separate block tree mode is applied, the luma CTBs may be partitioned into CUs based on a specific coding tree structure, and the chroma CTBs may be partitioned into chroma CUs based on a different coding tree structure. This may mean that a CU in an I slice / patch group may consist of a coding block for the luma component and coding blocks for two chroma components, and a CU in a P or B slice / patch group may consist of blocks for three types of color components. In this document, a slice may be referred to as a patch / patch group, and a patch / patch group may be referred to as a slice.

[0140] Although a quadtree coding tree structure of nested multi-type trees has been described, the CU partition structure is not limited thereto. For example, the BT structure and the TT structure may be interpreted as concepts included in the multi-partition tree (MPT) structure, and the CU may be interpreted as being partitioned by the QT structure and the MPT structure. In the example of partitioning the CU by the QT structure and the MPT structure, the partition structure may be determined by signaling a syntax element (e.g., MPT_split_type) including information about how many blocks a leaf node of the QT structure is partitioned into and a syntax element (e.g., MPT_split_mode) including information about in which direction the leaf node of the QT structure is partitioned, either vertically or horizontally.

[0141] In another example, a CU may be partitioned using a method different from the QT structure, the BT structure, or the TT structure. That is, unlike partitioning a CU of a lower depth using a CU of 1 / 4 the size of a higher depth according to the QT structure, partitioning a CU of a lower depth using a CU of 1 / 2 the size of a higher depth according to the BT structure, or partitioning a CU of a lower depth using a CU of 1 / 4 or 1 / 2 the size of a higher depth according to the TT structure, a CU of a lower depth may be partitioned using a CU of 1 / 5, 1 / 3, 3 / 8, 3 / 5, 2 / 3, or 5 / 8 the size of a higher depth, depending on the situation, but the CU partitioning is not limited thereto.

[0142] Figure 5 The layered structure of the encoded image / video is shown as an example.

[0143] Reference Figure 5 , the coded image / video is divided into the Video Coding Layer (VCL) that processes the image / video and its own decoding processing, the subsystem that sends and stores the coded information, and the NAL (Network Abstraction Layer) that is responsible for the function and exists between the VCL and the subsystem.

[0144] In the VCL, VCL data including compressed image data (slice data) is generated, or a parameter set including a picture parameter set (PSP), a sequence parameter set (SPS) and a video parameter set (VPS) or a supplemental enhancement information (SEI) message additionally required for image decoding processing can be generated.

[0145] In NAL, a NAL unit can be generated by adding header information (NAL unit header) to the raw byte sequence payload (RBSP) generated in the VCL. In this case, RBSP refers to slice data, parameter sets, SEI messages, etc. generated in the VCL. The NAL unit header may include NAL unit type information specified according to the RBSP data included in the corresponding NAL unit.

[0146] As shown in the figure, NAL units can be classified into VCL NAL units and non-VCL NAL units according to the RBSP generated in the VCL. A VCL NAL unit may refer to a NAL unit including information about an image (slice data), and a non-VCL NAL unit may refer to a NAL unit including information required for decoding an image (parameter set or SEI message).

[0147] The VCL NAL units and non-VCL NAL units can be transmitted over a network by adding header information according to the data standard of the subsystem. For example, the NAL units can be converted into a predetermined standard data format such as the H.266 / VVC file format, the Real-time Transport Protocol (RTP), or the Transport Stream (TS) and transmitted over various networks.

[0148] As described above, the NAL unit may be specified with a NAL unit type according to the RBSP data structure included in the corresponding NAL unit, and information about the NAL unit type may be stored in a NAL unit header and signaled.

[0149] For example, NAL units can be classified into VCL NAL unit types and non-VCL NAL unit types according to whether the NAL unit includes information about the image (slice data). VCL NAL unit types can be classified according to the nature and type of the picture included in the VCL NAL unit, and non-VCL NAL unit types can be classified according to the type of parameter set.

[0150] The following are examples of NAL unit types specified according to the type of parameter sets included in non-VCL NAL unit types.

[0151] -APS (Adaptation Parameter Set) NAL unit: type of NAL unit including APS

[0152] -DPS (Decoding Parameter Set) NAL unit: the type of NAL unit that includes DPS

[0153] -VPS (Video Parameter Set) NAL unit: Type of NAL unit containing VPS

[0154] -SPS (Sequence Parameter Set) NAL unit: the type of NAL unit that includes the SPS

[0155] -PPS (Picture Parameter Set) NAL unit: type of NAL unit including PPS

[0156] - PH (Picture Header) NAL unit: the type of NAL unit containing PH

[0157] The above-mentioned NAL unit type may have syntax information of the NAL unit type, and the syntax information may be stored in the NAL unit header and notified by a signal. For example, the syntax information may be nal_unit_type, and the NAL unit type may be specified by the nal_unit_type value.

[0158] Furthermore, as described above, a picture may include multiple slices, and a slice may include a slice header and slice data. In this case, a picture header may be further added to the multiple slices (slice header and slice data set) in a picture. The picture header (picture header syntax) may include information / parameters generally applicable to a picture. In this document, a slice may be mixed with or replaced by a tile group. In addition, in this document, a slice header may be mixed with or replaced by a tile group header.

[0159] A slice header (slice header syntax) may include information / parameters generally applicable to a slice. An APS (APS syntax) or a PPS (PPS syntax) may include information / parameters generally applicable to one or more slices or pictures. An SPS (SPS syntax) may include information / parameters generally applicable to one or more sequences. A VPS (VPS syntax) may include information / parameters generally applicable to multiple layers. A DPS (DPS syntax) may include information / parameters generally applicable to the overall video. A DPS may include information / parameters related to the concatenation of coded video sequences (CVS). A high-level syntax (HLS) in this document may include at least one of an APS syntax, a PPS syntax, an SPS syntax, a VPS syntax, a DPS syntax, and a slice header syntax.

[0160] In this document, the image / image information encoded from the encoding device and notified to the decoding device in the form of a bit stream by a signal includes not only the segmentation-related information, intra-frame / inter-frame prediction information, residual information, loop filtering information, etc. in the picture, but also includes the information included in the slice header, the information included in the APS, the information included in the PPS, the information included in the SPS and / or the information included in the VPS.

[0161] Furthermore, to compensate for differences between the original image and the reconstructed image due to errors occurring during compression encoding processes such as quantization, loop filtering may be performed on the reconstructed samples or reconstructed pictures as described above. As described above, loop filtering may be performed by filters of an encoding device and a decoding device, and may include a deblocking filter, SAO, and / or an adaptive loop filter (ALF). For example, the ALF process may be performed after the deblocking filter process and / or the SAO process are completed. However, even in this case, the deblocking filter process and / or the SAO process may be omitted.

[0162] Furthermore, to increase coding efficiency, luma mapping and chroma scaling (LMCS) may be applied as described above. LMCS may be referred to as a loop shaper (shaping). To increase coding efficiency, LMCS control and / or LMCS-related information signaling may be performed hierarchically.

[0163] Figure 6 The hierarchical structure of the CVS according to the embodiment of this document is exemplarily shown.

[0164] Reference Figure 6 A coded video sequence (CVS) may include a sequence parameter set (SPS), one or more picture parameter sets (PPS), and one or more subsequent coded pictures. Each coded picture may be divided into rectangular regions. Rectangular regions may be referred to as tiles. One or more tiles may be collected to form a tile group or slice. In this case, the tile group header may be linked to the picture parameter set (PPS), and the PPS may be linked to the SPS.

[0165] Figure 7 The hierarchical structure of the CVS according to the embodiment of this document is exemplarily shown. Figure 8 The hierarchical structure of a CVS according to another embodiment of this document is exemplarily shown.

[0166] Reference Figure 8 A coded video sequence (CVS) may include an SPS, a PPS, a patch group header, patch data, and / or a CTU. Here, the patch group header and patch data may be referred to as a slice header and slice data, respectively.

[0167] The SPS may include a local flag to enable tools to be used in the CVS. In addition, the SPS may be referenced by a PPS that includes information about parameters that change for each picture. Each coded picture may include one or more coded rectangular domain tiles. Tiles may be grouped into raster scans that form tile groups. Each tile group is encapsulated with header information called a tile group header. Each tile is composed of a CTU that includes coded data. Here, the data may include original sample values, predicted sample values, and their luminance and chrominance components (luminance predicted sample values and chrominance predicted sample values).

[0168] According to existing methods, ALF data (ALF parameters) or LMCS data (LMCS parameters) is included in the tile group header. Considering that a video is composed of multiple pictures and a picture includes multiple tiles, frequently signaling ALF data (ALF parameters) or LMCS data (LMCS parameters) in tile group units leads to a problem of deteriorating coding efficiency.

[0169] According to the embodiments proposed in this document, ALF parameters or LMCS data (LMCS parameters) may be included in the APS to be signaled as follows.

[0170] Reference Figure 7 , an APS may be defined, and the APS may carry necessary ALF data (ALF parameters). In addition, the APS may have self-identification parameters, ALF data, and / or LMCS data. The self-identification parameters of the APS may include an APS ID. That is, the APS may include information indicating the APS ID. The patch group header or slice header may use the APS index information to refer to the APS. In other words, the patch group header or slice header may include the APS index information, and the ALF process may be performed for the target block based on the LMCS data (LMCS parameters) included in the APS having the APS ID indicated by the APS index information, or the ALF process of the target block may be performed based on the ALF data (ALF parameters) included in the APS having the APS ID indicated by the APS index information. Here, the APS index information may be referred to as APS ID information.

[0171] In an example, the SPS may include a flag that allows the use of ALF. For example, when the CVS starts, the SPS may be checked and the flag in the SPS may be checked. For example, the SPS may include the syntax of Table 1 below. The syntax of Table 1 may be part of the SPS.

[0172] [Table 1]

[0173]

[0174] For example, the semantics of the syntax elements included in the syntax of Table 1 above can be expressed as in the following table.

[0175] [Table 2]

[0176]

[0177] That is, the sps_alf_enabled_flag syntax element can indicate whether ALF is available based on whether its value is 0 or 1. The sps_alf_enabled_flag syntax element can be called an ALF available flag (first ALF available flag) and can be included in the SPS. That is, the ALF available flag can be signaled in the SPS (or SPS level). If the value of the ALF available flag signaled in the SPS is 1, the ALF can be basically determined to be available relative to the picture referenced by the SPS in the CVS. In addition, as described above, the ALF can be individually handled as on / off by signaling an additional available flag at a lower level than the SPS.

[0178] For example, if the ALF tool is available for CVS, an additional availability flag (which may be referred to as a second ALF availability flag) may be signaled in the patch group header or slice header. For example, if ALF is available at the SPS level, the second ALF availability flag may be parsed / signaled. If the second ALF availability flag value is 1, ALF data may be parsed through the patch group header or slice header. For example, the second ALF availability flag may specify ALF availability conditions for luma and chroma components. ALF data may be accessed through APSID information.

[0179] [Table 3]

[0180]

[0181] [Table 4]

[0182]

[0183] For example, the semantics of the syntax elements of the syntax of Table 3 or Table 4 above can be expressed as in the following table.

[0184] [Table 5]

[0185]

[0186]

[0187] [Table 6]

[0188]

[0189] The second ALF available flag may include a tile_group_alf_enabled_flag syntax element or a slice_alf_enabled_flag syntax element.

[0190] Based on the APS ID information (eg, tile_group_aps_id syntax element or slice_aps_id syntax element), the APS referenced by the corresponding tile group or the corresponding slice may be identified. The APS may include ALF data.

[0191] In addition, for example, the structure of the APS including the ALF data may be described based on the following syntax and semantics: The syntax of Table 7 may be a part of the APS.

[0192] [Table 7]

[0193]

[0194] [Table 8]

[0195]

[0196]

[0197] As described above, the adaptation_parameter_set_id syntax element may represent the identifier of the corresponding APS. That is, the APS may be identified based on the adaptation_parameter_set_id syntax element. The adaptation_parameter_set_id syntax element may be referred to as APS ID information. Additionally, the APS may include an ALF data field. The ALF data field may be parsed or signaled after the adaptation_parameter_set_id syntax element.

[0198] Additionally, for example, an APS extension flag (e.g., an aps_extension_flag syntax element) may be parsed or signaled in the APS. The APS extension flag may indicate whether an APS extension data flag (aps_extension_data_flag) syntax element is present. For example, the APS extension flag may be used to provide an extension point for later versions of the VVC standard.

[0199] Figure 9 An exemplary LMCS structure according to an embodiment of this document is shown. Figure 9The LMCS structure 900 includes a loop mapping section 910 for luma components and a luma-dependent chroma residual scaling section 920 for chroma components based on an adaptive piecewise linear (adaptive PWL) model. The dequantization and inverse transform 911, reconstruction 912, and intra prediction 913 blocks of the loop mapping section 910 represent the processing applied in the mapped (shaped) domain. The loop filter 915, motion compensation or inter prediction 917 blocks of the loop mapping section 910, and the reconstruction 922, intra prediction 923, motion compensation or inter prediction 924, and loop filter 925 blocks of the chroma residual scaling section 920 represent the processing applied in the original (non-mapped, non-shaped) domain.

[0200] like Figure 9 As shown, when LMCS is enabled, at least one of a reverse mapping (shaping) process 914, a forward mapping (shaping) process 918, and a chroma scaling process 921 may be applied. For example, the reverse mapping process may be applied to (reconstructed) luma samples (or luma samples or luma sample arrays) in a reconstructed picture. The reverse mapping process may be performed based on a piecewise function (reverse) index of the luma samples. The piecewise function (reverse) index may identify the segment to which the luma sample belongs. The output of the reverse mapping process is a modified (reconstructed) luma sample (or modified luma sample or modified luma sample array). LMCS may be enabled or disabled at the tile group (or slice), picture, or higher level.

[0201] A forward mapping process and / or a chroma scaling process may be applied to generate a reconstructed picture. The picture may include luma samples and chroma samples. A reconstructed picture using luma samples may be referred to as a reconstructed luma picture, and a reconstructed picture using chroma samples may be referred to as a reconstructed chroma picture. The combination of the reconstructed luma picture and the reconstructed chroma picture may be referred to as a reconstructed picture. The reconstructed luma picture may be generated based on a forward mapping process. For example, if inter-frame prediction is applied to the current block, forward mapping is applied to luma prediction samples derived based on (reconstructed) luma samples in a reference picture. Since the (reconstructed) luma samples in the reference picture are generated based on a reverse mapping process, forward mapping may be applied to the luma prediction samples, thereby deriving mapped (shaped) luma prediction samples. The forward mapping process may be performed based on a piecewise function index of the luma prediction samples. The piecewise function index may be derived based on the values of the luma prediction samples or the values of luma samples in a reference picture used for inter-frame prediction. If intra-frame prediction (or intra-block copy (IBC)) is applied to the current block, forward mapping is not required because the reverse mapping process has not yet been applied to the reconstructed samples in the current picture. The (reconstructed) luma samples in the reconstructed luma picture are generated based on the mapped luma prediction samples and the corresponding luma residual samples.

[0202] The reconstructed chroma picture may be generated based on a chroma scaling process. For example, the (reconstructed) chroma samples in the reconstructed chroma picture may be generated based on the chroma residual samples (cres ) and chroma prediction samples. Based on the (scaled) chroma residual samples (c resScale ) and the chroma residual scaling factor (cScaleInv can be called varScale) to derive the chroma residual samples (c res ). The chroma residual scaling factor can be calculated based on the shaped luma prediction sample value of the current block. For example, the chroma residual scaling factor can be calculated based on the shaped luma prediction sample value Y' pred The average brightness value ave(Y' pred ) to calculate the scaling factor. For reference, the (scaled) chroma residual samples derived based on inverse transform / dequantization can be referred to as c resScale , the chroma residual samples derived by performing the (inverse) scaling process on the (scaled) chroma residual samples may be referred to as c res .

[0203] Figure 10 An LMCS structure according to another embodiment of this document is shown. Figure 10 Reference Figure 9 Here, we mainly describe Figure 10 The LMCS structure and Figure 9 The differences between the LMCS structures 900. Figure 10 The loop mapping part and the luminance-related chrominance residual scaling part can be used with Figure 9 The loop mapping section 910 and the luma-dependent chroma residual scaling section 920 operate identically (similarly).

[0204] Reference Figure 10 , the chroma residual scaling factor can be derived based on the luma reconstructed samples. In this case, the average luma value (avgYr) can be obtained (derived) based on the neighboring luma reconstructed samples outside the reconstructed block, rather than the internal luma reconstructed samples of the reconstructed block, and the chroma residual scaling factor can be derived based on the average luma value (avgYr). Here, the neighboring luma reconstructed samples can be the neighboring luma reconstructed samples of the current block, or can be the neighboring luma reconstructed samples of the virtual pipeline data unit (VPDU) including the current block. For example, when intra prediction is applied to the target block, the reconstructed samples can be derived based on the prediction samples derived based on the intra prediction. In another example, when inter prediction is applied to the target block, forward mapping is applied to the prediction samples derived based on the inter prediction, and the reconstructed samples are generated (derived) based on the shaped (or forward mapped) luma prediction samples.

[0205] The video / image information notified by the bitstream signal may include LMCS parameters (information about LMCS). The LMCS parameters can be configured as high-level syntax (HLS, including slice header syntax) and the like. A detailed description and configuration of the LMCS parameters will be described later. As described above, the syntax table described in this document (and the following embodiments) can be configured or encoded at the encoder end and notified to the decoder end via the bitstream signal. The decoder can parse / decode the information about the LMCS in the syntax table (in the form of syntax components). One or more embodiments to be described below may be combined. The encoder can encode the current picture based on the information about the LMCS, and the decoder can decode the current picture based on the information about the LMCS.

[0206] The loop mapping of the luma component can adjust the dynamic range of the input signal to improve compression efficiency by redistributing codewords across the dynamic range. For luma mapping, a forward mapping (shaping) function (FwdMap) and an inverse mapping (shaping) function (InvMap) corresponding to the forward mapping function (FwdMap) can be used. The FwdMap function can be signaled using a piecewise linear model, for example, the piecewise linear model can have 16 segments or bins, and the segments can have equal lengths. In one example, the InvMap function does not need to be signaled, but is derived from the FwdMap function. That is, the inverse mapping can be a forward mapping function. For example, the inverse mapping function can be mathematically constructed as a symmetric function of the forward mapping, as reflected by the line y=x.

[0207] In-loop (luminance) shaping can be used to map input luma values (samples) to modified values in a shaped domain. The shaped values can be encoded and then mapped back to the original (unmapped, unshaped) domain after reconstruction. To compensate for interactions between luma and chroma signals, chroma residual scaling can be applied. In-loop shaping is accomplished using a high-level syntax that specifies a shaper model. The shaper model syntax can signal a piecewise linear model (PWL model). For example, the shaper model syntax can signal a PWL model with 16 bins or segments of equal length. A forward lookup table (FwdLUT) and / or an inverse lookup table (InvLUT) can be derived based on the piecewise linear model. For example, the PWL model precomputes 1024-entry forward (FwdLUT) and inverse (InvLUT) lookup tables (LUTs). As an example, when deriving the forward lookup table FwdLUT, the inverse lookup table InvLUT can be derived based on the forward lookup table FwdLUT. The forward lookup table FwdLUT can map the input luminance value Yi to the modified value Yr, and the reverse lookup table InvLUT can map the modified value Yr to the reconstructed value Y′i. The reconstructed value Y′i can be derived based on the input luminance value Yi.

[0208] In one example, the SPS may include the syntax of Table 9 below. The syntax of Table 9 may include sps_reshaper_enabled_flag as a tool enable flag. Here, sps_reshaper_enabled_flag may be used to specify whether a reshaper is used in a coded video sequence (CVS). That is, sps_reshaper_enabled_flag may be a flag that enables reshaping in the SPS. In one example, the syntax of Table 9 may be part of the SPS.

[0209] [Table 9]

[0210]

[0211] In one example, the semantics of the syntax elements sps_seq_parameter_set_id and sps_reshaper_enabled_flag may be as shown in Table 10 below.

[0212] [Table 10]

[0213]

[0214] In one example, the tile group header or the slice header may include the syntax of Table 11 or Table 12 below.

[0215] [Table 11]

[0216]

[0217] [Table 12]

[0218]

[0219] The semantics of the syntax elements included in the syntax of Table 11 or Table 12 may include, for example, matters disclosed in the following table.

[0220] [Table 13]

[0221]

[0222] [Table 14]

[0223]

[0224]

[0225] As an example, once the reshaping-enabled flag (i.e., sps_reshaper_enabled_flag) is parsed in the SPS, the tile group header can parse additional data (i.e., the information included in Table 13 or Table 14 above) used to construct the lookup table (FwdLUT and / or InvLUT). To do this, the state of the SPS reshaper flag (sps_reshaper_enabled_flag) can be first checked in the slice header or tile group header. When sps_reshaper_enabled_flag is true (or 1), an additional flag, tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag), can be parsed. The purpose of tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag) can be to indicate the presence of a reshaping model. For example, when tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag) is true (or 1), it may indicate that a reshaper is present for the current tile group (or current slice). When tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag) is false (or 0), it may indicate that a reshaper is not present for the current tile group (or current slice).

[0226] If a reshaper exists and is enabled in the current tile group (or current slice), the reshaper model (i.e., tile_group_reshaper_model() or slice_reshaper_model()) may be processed. Furthermore, an additional flag, tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag), may also be parsed. The tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag) may indicate whether the reshaping model is used for the current tile group (or slice). For example, if tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag) is 0 (or false), it may indicate that the reshaping model is not used for the current tile group (or current slice). If tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag) is 1 (or true), it may indicate that the reshaping model is used for the current tile group (or slice).

[0227] As an example, tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag) may be true (or 1) and tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag) may be false (or 0). This means that the reshaping model exists but is not used in the current tile group (or slice). In this case, the reshaping model can be used in future tiles (or slices). As another example, tile_group_reshaper_enable_flag may be true (or 1) and tile_group_reshaper_model_present_flag may be false (or 0). In this case, the decoder uses the reshaper from the previous initialization.

[0228] When parsing the reshaping model (i.e., tile_group_reshaper_model() or slice_reshaper_model()) and tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag), it may be determined (evaluated) whether conditions required for chroma scaling exist. The above conditions include condition 1 (the current tile group / slice has not been intra-coded) and / or condition 2 (the current tile group / slice has not been split into two separate coding quadtree structures for luma and chroma, i.e., the block structure of the current tile group / slice is not a dual-tree structure). If condition 1 and / or condition 2 are true and / or tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag) is true (or 1), tile_group_reshaper_chroma_residual_scale_flag (or slice_reshaper_chroma_residual_scale_flag) may be parsed. When tile_group_reshaper_chroma_residual_scale_flag (or slice_reshaper_chroma_residual_scale_flag) is enabled (if 1 or true), it may indicate that chroma residual scaling is enabled for the current tile group (or slice). When tile_group_reshaper_chroma_residual_scale_flag (or slice_reshaper_chroma_residual_scale_flag) is disabled (if 0 or false), it may indicate that chroma residual scaling is disabled for the current tile group (or slice).

[0229] The purpose of the patch group shaping model is to parse the data needed to construct a lookup table (LUT). These LUTs are constructed based on the idea that the distribution of the allowed range of luma values can be divided into a number of bins (e.g., 16 bins), which can be represented using a set of 16 PWL equations. Therefore, any luma value that falls within a given bin can be mapped to a modified luma value.

[0230] Figure 11 A diagram showing an exemplary forward map is shown. Figure 11 , five bins are shown as an example.

[0231] Reference Figure 11, the x-axis represents the input luma value, and the y-axis represents the modified output luma value. The x-axis is divided into five bins or slices, each of length L. That is, the five bins mapped to the modified luma value have the same length. The forward lookup table (FwdLUT) can be constructed using data available from the tile group header (i.e., shaper data), thus facilitating the mapping.

[0232] In one embodiment, an output pivot point associated with a bin index may be calculated. The output pivot point may set (mark) the minimum and maximum boundaries of an output range for luma codeword shaping. The calculation process of the output pivot point may be performed by calculating a piecewise cumulative distribution function (CDF) of the number of codewords. The output pivot range may be sliced based on the maximum number of bins to be used and the size of the lookup table (FwdLUT or InvLUT). As an example, the output pivot range may be sliced based on the product between the maximum number of bins and the size of the lookup table (size of the LUT * maximum number of bin indices). For example, if the product between the maximum number of bins and the size of the lookup table is 1024, the output pivot range may be sliced into 1024 entries. This sawtoothing of the output pivot range may be performed (applied or implemented) based on (using) a scaling factor. In one example, the scaling factor may be derived based on Equation 1 below.

[0233] [Formula 1]

[0234] SF=(y2-y1)*(1<<FP_PREC)+c

[0235] In Equation 1, SF represents a scaling factor, and y1 and y2 represent output pivot points corresponding to respective bins. In addition, FP_PREC and c may be predetermined constants. The scaling factor determined based on Equation 1 may be referred to as a scaling factor for forward shaping.

[0236] In another embodiment, for reverse shaping (reverse mapping), for the defined range of bins to be used (i.e., from reshaper_model_min_bin_idx to reshape_model_max_bin_idx), the input shaping pivot point and the mapping reverse output pivot point corresponding to the mapping pivot point of the forward LUT are obtained (given by the bin index considered * the number of initial codewords). In another example, the scaling factor SF can be derived based on the following equation 2.

[0237] [Formula 2]

[0238] SF=(y2-y1)*(1<<FP_PREC) / (x2-x1)

[0239] In Formula 2, SF represents a scaling factor, x1 and x2 represent input pivot points, and y1 and y2 represent output pivot points (output pivot points of reverse mapping) corresponding to each fragment (bin). Here, the input pivot point can be a pivot point mapped based on a forward lookup table (FwdLUT), and the output pivot point can be a pivot point reversely mapped based on an inverse lookup table (InvLUT). In addition, FP_PREC can be a predetermined constant value. The FP_PREC of Formula 2 can be the same as or different from the FP_PREC of Formula 1. The scaling factor determined based on Formula 2 can be referred to as a scaling factor for reverse shaping. During reverse shaping, the segmentation of the input pivot point can be performed based on the scaling factor of Formula 2. The scaling factor SF is used to slice the range of the input pivot point. Based on the segmented input pivot point, bin indices ranging from 0 to minimum bin index (reshaper_model_min_bin_idx) and / or from minimum bin index (reshaper_model_min_bin_idx) to maximum bin index (reshape_model_max_bin_idx) are assigned pivot values corresponding to minimum bin value and maximum bin value.

[0240] In one example, LMCS data (lmcs_data) may be included in an APS. For example, the semantics of the APS may be 32 APSs signaled for encoding.

[0241] The following table shows the syntax and semantics of an exemplary APS according to embodiments of this document.

[0242] [Table 15]

[0243]

[0244] [Table 16]

[0245]

[0246] Referring to Table 15 above, type information of the APS parameter (eg, aps_params_type) may be parsed or signaled in the APS. Type information of the APS parameter may be parsed or signaled after adaptation_parameter_set_id.

[0247] APS_params_type, ALF_APS, and LMCS_APS included in Table 15 above may be described according to Table 3.2 included in Table 16. That is, according to APS_params_type included in Table 15 above, the type of APS parameter applied to the APS may be configured as in Table 3.2 included in Table 16. Syntax elements included in Table 15 may be described with reference to Table 8. Descriptions related to the APS may be supported by the descriptions described above in conjunction with Tables 1 to 8.

[0248] Referring to Table 16, for example, aps_params_type may be a syntax element used to classify the type of the corresponding APS parameter. APS parameter types may include ALF parameters and LMCS parameters. Referring to Table 16, if the value of the type information aps_params_type is 0, the name of aps_params_type may be determined to be ALF_APS (or ALF APS), and the type of the APS parameter may be determined to be an ALF parameter (APS parameter may represent an ALF parameter). In this case, the ALF data field (i.e., alf_data()) may be parsed / signaled to the APS. If the value of the type information aps_params_type is 1, the name of aps_params_type may be determined to be LMCS_APS (or LMCS APS), and the type of the APS parameter may be determined to be an LMCS parameter (APS parameter may represent an LMCS parameter). In this case, the LMCS data field (i.e., lmcs_data()) may be parsed or signaled to the APS.

[0249] Table 17 and / or Table 18 below represent the syntax of the reshaper model according to an embodiment. The reshaper model may be referred to as an LMCS model. Although these reshaper models are exemplarily described herein as tile group reshapers, the present specification is not necessarily limited to this embodiment. For example, the reshaper model may be included in the APS, or the tile group reshaper model may be referred to as a slice reshaper model or LMCS data (LMCS data field). In addition, the prefix "reshaper_model" or "Rsp" may be used interchangeably with "lmcs". For example, in the following tables and the following description, reshaper_model_min_bin_idx, reshaper_model_delta_max_bin_idx, reshaper_model_max_bin_idx, RspCW, and RsepDeltaCW may be used interchangeably with lmcs_min_bin_idx, lmcs_delta_max_bin_idx, lmcs_max_bin_idx, lmcsCW, and lmcsDeltaCW, respectively.

[0250] The LMCS data (lmcs_data()) or the shaper model (tile group shaper or slice shaper) included in the above Table 15 can be expressed as the syntax included in the following table.

[0251] [Table 17]

[0252]

[0253] [Table 18]

[0254]

[0255] The semantics of the syntax elements included in the syntax of Table 17 and / or Table 18 may include, for example, matters disclosed in the following table.

[0256] [Table 19]

[0257]

[0258]

[0259] [Table 20]

[0260]

[0261]

[0262] The inverse mapping process of the luminance samples according to this document can be described in the form of standard documents as shown in the following table.

[0263] [Table 21]

[0264]

[0265]

[0266] The identification of the piecewise function index processing of luma samples according to this document can be described in the form of standard documents as shown in the following table: In Table 22, idxYInv can be called a reverse mapping index, and the reverse mapping index can be derived based on the reconstructed luma sample (lumaSample).

[0267] [Table 22]

[0268]

[0269] Luma mapping can be performed based on the above embodiments and examples, and the above syntax and components included therein may be merely exemplary representations, and the embodiments in this document are not limited to the above tables or formulas. Hereinafter, a method for performing chroma residual scaling (scaling of the chroma components of the residual samples) based on luma mapping will be described.

[0270] (Lumence-dependent) chroma residual scaling is designed to compensate for the interaction between the luma signal and its corresponding chroma signal. For example, chroma residual scaling is also signaled at the tile group level whether it is enabled. In one example, if luma mapping is enabled and if dual-tree partitioning (also known as separate chroma trees) is not applied for the current tile group, an additional flag is signaled to indicate whether luma-dependent chroma residual scaling is enabled. In other examples, luma-dependent chroma residual scaling is disabled when luma mapping is not used, or when dual-tree partitioning is used in the current tile group. In another example, luma-dependent chroma residual scaling is always disabled for chroma blocks with an area less than or equal to 4.

[0271] The chroma residual scaling may be based on the average value of the corresponding luma prediction block (the luma component of the prediction block to which intra prediction mode and / or inter prediction mode is applied). The scaling operation at the encoder and / or decoder side may be implemented using fixed-point integer arithmetic based on Equation 3 below.

[0272] [Formula 3]

[0273] c'=sign(c)*((abs(c)*s+2CSCALE_FP_PREC-1)>>CSCALE_FP_PREC)

[0274] In Equation 3, c' represents the scaled chroma residual sample (the scaled chroma component of the residual sample), c represents the chroma residual sample (the chroma residual sample, the chroma component of the residual sample), s represents the chroma residual scaling factor, and CSCALE_FP_PREC represents a (predefined) constant value to specify the precision. For example, CSCALE_FP_PREC can be 11.

[0275] Figure 12 is a flowchart illustrating a method for deriving a chroma residual scaling index according to an embodiment of this document. Figure 12 The method in can be based on Figure 9 and included in Figure 9 The tables, formulas, variables, arrays and functions in the relevant description are executed.

[0276] In step S1210, it is determined whether the prediction mode of the current block is an intra prediction mode or an inter prediction mode based on the prediction mode information. If the prediction mode is an intra prediction mode, the current block or the prediction samples of the current block are considered to be in the shaped (mapped) area. If the prediction mode is an inter prediction mode, the current block or the prediction samples of the current block are considered to be in the original (unmapped, unshaped) area.

[0277] In step S1220, when the prediction mode is intra prediction mode, the average of the current block (or the luma prediction samples of the current block) may be calculated (derived). That is, the average of the current block in the shaped region is directly calculated. The average may also be referred to as a mean value or average.

[0278] In step S1221, when the prediction mode is an inter prediction mode, forward shaping (forward mapping) may be performed (applied) on the luma prediction samples of the current block. Through forward shaping, the luma prediction samples based on the inter prediction mode may be mapped from the original region to the shaped region. In one example, forward shaping of the luma prediction samples may be performed based on the shaping model described in Table 17 and / or Table 18 above.

[0279] In step S1222, the average of the forward-shaped (forward-mapped) luma prediction samples may be calculated (derived). That is, an average process of the forward-shaped results may be performed.

[0280] In step S1230, a chroma residual scaling index may be calculated. When the prediction mode is intra prediction mode, the chroma residual scaling index may be calculated based on the average of luma prediction samples. When the prediction mode is inter prediction mode, the chroma residual scaling index may be calculated based on the average of forward-shaped luma prediction samples.

[0281] In an embodiment, the chroma residual scaling index may be calculated based on a for loop syntax. The following table shows an exemplary for loop syntax for deriving (calculating) the chroma residual scaling index.

[0282] [Table 23]

[0283]

[0284] In Table 23, idxS represents the chroma residual scaling index, idxFound represents the index indicating whether the chroma residual scaling index that satisfies the condition of the if statement is obtained, S represents a predetermined constant value, and MaxBinIdx represents the maximum allowable bin index. ReshapPivot[idxS+1] (in other words, LmcsPivot[idxS+1]) can be derived based on Table 19 and / or Table 20 above.

[0285] In an embodiment, a chroma residual scaling factor may be derived based on the chroma residual scaling index. Equation 4 is an example for deriving the chroma residual scaling factor.

[0286] [Formula 4]

[0287] s=ChromaScaleCoef[idxS]

[0288] In Equation 4, s represents the chroma residual scaling factor, and ChromaScaleCoef can be a variable (or array) derived based on Table 19 and / or Table 20 above.

[0289] As described above, the average luma value of the reference sample may be obtained, and a chroma residual scaling factor may be derived based on the average luma value. As described above, the chroma component residual samples may be scaled based on the chroma residual scaling factor, and the chroma component reconstructed samples may be generated based on the scaled chroma component residual samples.

[0290] In one embodiment of this document, a signaling structure for efficiently applying the above-mentioned LMCS is proposed. According to this embodiment of this document, for example, LMCS data can be included in the HLS (i.e., APS), and through the lower-level header information of the APS (i.e., picture header, slice header), the LMCS model (shaper model) can be adaptively derived by signaling the ID of the APS (referred to as header information). The LMCS model can be derived based on LMCS parameters. In addition, for example, multiple APS IDs can be signaled through the header information, and thus different LMCS models can be applied on a block-by-block basis within the same picture / slice.

[0291] In one embodiment according to the present document, a method is proposed to efficiently perform the operations required for LMCS. According to the semantics described above in Table 19 and / or Table 20, a division operation by the segment length lmcsCW[i] (also referred to as RspCW[i] in this document) is required to derive InvScaleCoeff[i]. The segment length of the reverse mapping may not be a power of 2, which means that the division cannot be performed by bit shifting.

[0292] For example, calculating InvScaleCoeff may require up to 16 divisions per slice. According to Table 19 and / or Table 20 above, for 10-bit encoding, lmcsCW[i] ranges from 8 to 511, so to implement division operations by lmcsCW[i] using the LUT, the LUT size must be 504. Furthermore, for 12-bit encoding, lmcsCW[i] ranges from 32 to 2047, so the LUT size must be 2016 to implement division operations by lmcsCW[i] using the LUT. That is, division is expensive in hardware implementation, so it is desirable to avoid it whenever possible.

[0293] In one aspect of this embodiment, lmcsCW[i] can be constrained to be a multiple of a fixed number (or a predetermined number). Consequently, the lookup table (LUT) for division (the capacity or size of the LUT) can be reduced. For example, if lmcsCW[i] becomes a multiple of 2, the size of the LUT for the division process can be reduced by half.

[0294] In another aspect of this embodiment, it is proposed that for encoding at higher internal bit depths, on top of the existing constraint that "the value of lmcsCW[i] should be in the range of (OrgCW>>3) to (OrgCW<<3-1)", if the coding bit depth is higher than 10, lmcsCW[i] is further constrained to be a multiple of 1<<(BitDepthY-10). Here, BitDepthY can be the luma bit depth. Therefore, the possible number of lmcsCW[i] does not change with the coding bit depth, and the size of the LUT required to calculate InvScaleCoeff does not increase due to the higher coding bit depth. For example, for a 12-bit internal coding bit depth, the limit value of lmcsCW[i] is a multiple of 4, then the LUT for the replacement division process will be the same as for 10-bit encoding. This aspect can be implemented alone, but can also be implemented in combination with the above aspects.

[0295] In another aspect of this embodiment, lmcsCW[i] can be constrained to a narrower range. For example, lmcsCW[i] can be constrained to be in the range from (OrgCW>>1) to (OrgCW<<1)-1. Then, for 10-bit encoding, the range of lmcsCW[i] can be [32, 127], so only a LUT of size 96 is required to calculate InvScaleCoeff.

[0296] In another aspect of this embodiment, lmcsCW[i] can be approximated to the nearest power of 2 and used in the shaper design. Thus, the division in the reverse mapping can be performed (and replaced) by bit shifting.

[0297] In one embodiment of this document, a constraint on the LMCS codeword range is proposed. According to Table 8, the LMCS codeword values range from (OrgCW>>3) to (OrgCW<<3)-1. This codeword range is too wide. Large discrepancies between RspCW[i] and OrgCW can lead to visual artifacts.

[0298] According to one embodiment of the present document, it is proposed to constrain the codewords of the LMCS PWL mapping to a narrow range. For example, the range of lmcsCW[i] can be in the range of (OrgCW>>1) to (OrgCW<<1)-1.

[0299] In one embodiment according to the present document, a single chroma residual scaling factor is proposed for chroma residual scaling in LMCS. Existing methods for deriving chroma residual scaling factors use the average value of the corresponding luminance block and derive the slope of each segment of the inverse luminance mapping as the corresponding scaling factor. In addition, the process of identifying the segment index requires the availability of the corresponding luminance block, which leads to latency issues. This is undesirable for hardware implementation. According to this embodiment of the present document, scaling in chroma blocks may not depend on the luminance block value, and identification of the segment index may not be required. Therefore, the chroma residual scaling process in LMCS can be performed without latency issues.

[0300] In one embodiment according to the present document, a single chroma scaling factor can be derived from both the encoder and decoder based on luma LMCS information. When the LMCS luma model is received, the chroma residual scaling factor can be updated. For example, when the LMCS model is updated, the single chroma residual scaling factor can be updated.

[0301] The following table shows an example of obtaining a single chroma scaling factor according to this embodiment.

[0302] [Table 24]

[0303]

[0304] Referring to Table 24, a single chroma scaling factor (eg, ChromaScaleCoeff or ChromaScaleCoeffSingle) may be obtained by averaging the inverse luma mapping slopes of all segments within LMCS_min_bin_idx and lmcs_max_bin_idx.

[0305] Figure 13 A linear fit of the pivot points according to an embodiment of the present document is shown. Figure 13 In FIG, pivot points P1, Ps and P2 are shown. The following embodiments or examples thereof will be described with reference to FIG. Figure 13 To describe.

[0306] In this embodiment example, a single chroma scaling factor may be obtained based on a linear approximation of the luma PWL mapping between the pivot points lmcs_min_bin_idx and lmcs_max_bin_idx+1 (LmcsMaxBinIdx+1). That is, the inverse slope of the linear mapping may be used as the chroma residual scaling factor. For example, Figure 13 The linear line 1 may be a straight line connecting the pivot points P1 and P2. Figure 13, in P1, the input value is x1 and the mapped value is 0, and in P2, the input value is x2 and the mapped value is y2. The inverse slope (inverse scale) of Linear Line 1 is (x2-x1) / y2, and a single chroma scaling factor ChromaScaleCoeffSingle can be calculated based on the input and mapped values of pivot points P1 and P2 and the following formula.

[0307] [Formula 5]

[0308] ChromaScaleCoeffSingle=(x2-x1)*(1<<CSCALE_FP_PREC) / y2

[0309] In Equation 5, CSCALE_FP_PREC represents a shift factor. For example, CSCALE_FP_PREC may be a predetermined constant value. In one example, CSCALE_FP_PREC may be 11.

[0310] In another example according to this embodiment, referring to Figure 13 , the input value at the pivot point Ps is min_bin_idx+1, and the mapped value at the pivot point Ps is ys. Therefore, the inverse slope (inverse scale) of Linear Line 1 can be calculated as (xs-x1) / ys, and a single chroma scaling factor ChromaScaleCoeffSingle can be calculated based on the input and mapped values of the pivot points P1 and Ps and the following formula.

[0311] [Formula 6]

[0312] ChromaScaleCoeffSingle=(xs-x1)*(1<<CSCALE_FP_PREC) / ys

[0313] In Equation 6, CSCALE_FP_PREC represents a shift factor (bit shift factor), and for example, CSCALE_FP_PREC may be a predetermined constant value. In one example, CSCALE_FP_PREC may be 11, and descaling bit shifting may be performed based on CSCALE_FP_PREC.

[0314] In another example according to this embodiment, a single chroma residual scaling factor can be derived based on a linear approximation line. Examples of derivation of a linear approximation line may include a linear connection of pivot points (i.e., lmcs_min_bin_idx, lmcs_max_bin_idx+1). For example, the linear approximation result may be represented by a codeword mapped by PWL. The mapped value y2 at P2 may be the sum of the codewords of all bins (shards), and the difference between the input value at P2 and the input value at P1 (x2-x1) is OrgCW*(lmcs_max_bin_idx-lmcs_min_bin_idx+1) (for OrgCW, see Table 19 and / or Table 20 above). The following table shows an example of obtaining a single chroma scaling factor according to the above embodiment.

[0315] [Table 25]

[0316]

[0317] Referring to Table 25, a single chroma scaling factor (eg, ChromaScaleCoeffSingle) may be obtained from two pivot points (ie, lmcs_min_bin_idx, lmcs_max_bin_idx). For example, the inverse slope of a linear mapping may be used as the chroma scaling factor.

[0318] In another example of this embodiment, a single chroma scaling factor can be obtained by linear fitting of the pivot points to minimize the error (or mean square error) between the linear fit and the existing PWL mapping. This example can be more accurate than simply connecting the two pivot points at lmcs_min_bin_idx and lmcs_max_bin_idx. There are many ways to find the optimal linear mapping, and one example is described below.

[0319] In one example, parameters b1 and b0 of the linear fitting formula y=b1*x+b0 for minimizing the sum of the least squares errors may be calculated based on the following Formula 7 and / or Formula 8.

[0320] [Formula 7]

[0321]

[0322] [Formula 8]

[0323]

[0324] In equations 7 and 8, x is the original brightness value, y is the shaped brightness value, and is the mean of x and y, xi and y i Represents the value of the i-th pivot point.

[0325] Reference Figure 13 , another simple approximation to the recognition linear map is given by:

[0326] - Linear line 1 is obtained by connecting the pivot points of the PWL map at lmcs_min_bin_idx and lmcs_max_bin_idx+1. This linear line is calculated with input values that are multiples of OrgW, lmcs_pivots_linear[i]

[0327] -Use Linear Line 1 and sum the differences of the pivot point mapping values using PWL mapping.

[0328] -Get the average difference avgDiff.

[0329] - Adjust the last pivot point of the Linear Line by the average difference, e.g. 2*avgDiff

[0330] - Use the inverse slope of the adjusted linear line as the chroma residual scale.

[0331] According to the above linear fitting, the chroma scaling factor (ie, the inverse slope of the forward mapping) can be derived (obtained) based on the following Equation 9 or Equation 10.

[0332] [Formula 9]

[0333] ChromaScaleCoeffSingle=OrgCW*(1< <CSCALE_FP_PREC) / lmcs_pivots_linear[lmcs_min_bin_idx+1]

[0334] [Equation 10]

[0335] ChromaScaleCoeffSingle=OrgCW*(lmcs_max_bin_idx-lmcs_max_bin_idx+1)*(1< <CSCALE_FP_PREC) / lmcs_pivots_linear[lmcs_max_bin_idx+1]

[0336] In the above formula, lmcs_pivots_linear[i] may be a mapping value for linear mapping. For linear mapping, all segments of the PWL mapping between the minimum bin index and the maximum bin index may have the same LMCS codeword (lmcsCW). That is, lmcs_pivots_linear[lmcs_min_bin_idx+1] may be the same as lmcsCW[lmcs_min_bin_idx].

[0337] In addition, in Equations 9 and 10, CSCALE_FP_PREC represents a shift factor (bit shift factor). For example, CSCALE_FP_PREC may be a predetermined constant value. In one example, CSCALE_FP_PREC may be 11.

[0338] By using a single chroma residual scaling factor (ChromaScaleCoeffSingle), there is no longer a need to calculate the average of the corresponding luma blocks and search for indices in the PWL linear map to obtain the chroma residual scaling factor. Therefore, coding efficiency using chroma residual scaling can be increased. This not only eliminates the dependency on the corresponding luma block, solving the latency issue, but also reduces complexity.

[0339] The LMCS data-related semantics and / or chroma sample-related luminance-dependent chroma residual scaling process according to the above embodiment can be described in the form of standard documents as shown in the following table.

[0340] [Table 26]

[0341]

[0342] [Table 27]

[0343]

[0344]

[0345] In another embodiment of this document, the encoder can determine parameters related to a single chroma scaling factor and signal these parameters to the decoder. Using signaling, the encoder can use other information available at the encoder to derive the chroma scaling factor. This embodiment aims to eliminate the chroma residual scaling delay problem.

[0346] For example, another example of identifying a linear mapping to be used to determine the chroma residual scaling factor is given as follows:

[0347] - lmcs_pivots_linear[i] is calculated for this linear line with input values that are multiples of OrgW by connecting the pivot points of the PWL map at lmcs_min_bin_idx and lmcs_max_bin_idx+1

[0348] - Get a weighted sum of the differences of the mapped values of the pivot point using those of the linear line 1 and the luma PWL mappings. The weights may be based on encoder statistics (e.g. a histogram of the bins).

[0349] -Get the weighted average difference avgDiff.

[0350] - Adjust the last pivot point of Linear 1 by the weighted average difference, e.g. 2*avgDiff

[0351] - Use the inverse slope of the adjusted linear line to calculate the chromaticity residual scale.

[0352] The following table shows an example of the syntax for signaling the y value for chroma scaling factor derivation.

[0353] [Table 28]

[0354]

[0355] In Table 28, the syntax element lmcs_chroma_scale may specify a single chroma (residual) scaling factor (ChromaScaleCoeffSingle=lmcs_chroma_scale) for LMCS chroma residual scaling. That is, information regarding the chroma residual scaling factor may be directly signaled, and the signaled information may be derived as the chroma residual scaling factor. In other words, the value of the signaled information regarding the chroma residual scaling factor may be (directly) derived as the value of the single chroma residual scaling factor. Here, the syntax element lmcs_chroma_scale may be signaled together with other LMCS data (i.e., syntax elements related to the absolute value and sign of the codeword, etc.).

[0356] Alternatively, the encoder may signal only the necessary parameters to derive the chroma residual scaling factors at the decoder. To derive the chroma residual scaling factors at the decoder, the input value x and the mapped value y are needed. Since the x value is the bin length and is a known number, it does not need to be signaled. After all, only the y value needs to be signaled in order to derive the chroma residual scaling factors. Here, the y value can be the mapped value of any pivot point in the linear mapping (i.e., Figure 13 The mapping value of P2 or Ps in ).

[0357] The following table shows an example of signaling mapping values used to derive the chroma residual scaling factor.

[0358] [Table 29]

[0359]

[0360] [Table 30]

[0361]

[0362] One of the syntaxes of Tables 29 and 30 above can be used to signal the y value at any linear pivot point specified by the encoder and decoder. That is, the encoder and decoder can use the same syntax to derive the y value.

[0363] First, an embodiment according to Table 29 will be described. In Table 29, lmcs_cw_linear may represent a mapping value at Ps or P2. That is, in the embodiment according to Table 29, a fixed number may be signaled through lmcs_cw_linear.

[0364] In an example according to this embodiment, if lmcs_cw_linear represents the mapping value of one bin (ie, Figure 13 lmcs_pivots_linear[lmcs_min_bin_idx+1]) in Ps, the chroma scaling factor can be derived based on the following formula.

[0365] [Equation 11]

[0366] ChromaScaleCoeffSingle=OrgCW*(1< <CSCALE_FP_PREC) / Imcs_cw_linear

[0367] In another example according to this embodiment, if lmcs_cw_linear represents lmcs_max_bin_idx+1 (ie, Figure 13 lmcs_pivots_linear[lmcs_max_bin_idx+1]) in P2, the chroma scaling factor can be derived based on the following formula.

[0368] [Equation 12]

[0369] ChromaScaleCoeffSingle=OrgCW*(lmcs_max_bin_idx-lmcs_max_bin_idx+1)*(1< <CSCALE_FP_PREC) / Imcs_cw_linear

[0370] In the above formula, CSCALE_FP_PREC represents a shift factor (bit shift factor), for example, CSCALE_FP_PREC may be a predetermined constant value. In one example, CSCALE_FP_PREC may be 11.

[0371] Next, an embodiment according to Table 30 is described. In this embodiment, lmcs_cw_linear may be signaled as a delta value relative to a fixed number (i.e., lmcs_delta_abs_cw_linear, lmcs_delta_sign_cw_linear_flag). In this example embodiment, when lmcs_cw_linear represents lmcs_pivots_linear[lmcs_min_bin_idx+1] (i.e., Figure 13When the mapping value in Ps is used, lmcs_cw_linear_delta and lmcs_cw_linear can be derived based on the following formula.

[0372] [Equation 13]

[0373] lmcs_cw_linear_delta=(1-2*lmcs_delta_sign_cw_linear_flag)*lmcs_delta_abs_linear_cw

[0374] [Equation 14]

[0375] lmcs_cw_linear=lmcs_cw_linear_delta+OrgCW

[0376] In another example of this embodiment, when lmcs_cw_linear represents lmcs_pivots_linear[lmcs_max_bin_idx+1] (ie, Figure 13 When the mapping values in P2) are used, lmcs_cw_linear_delta and lmcs_cw_linear can be derived based on the following formula.

[0377] [Equation 15]

[0378] lmcs_cw_linear_delta=(1-2*lmcs_delta_sign_cw_linear_flag)*lmcs_delta_abs_linear_cw

[0379] [Equation 16]

[0380] lmcs_cw_linear=lmcs_cw_linear_delta+OrgCW*(lmcs_max_bin_idx-lmcs_max_bin_idx+1)

[0381] In the above formula, OrgCW can be a value derived based on Table 19 and / or Table 20 above.

[0382] The LMCS data-related semantics and / or chroma sample-related luminance-dependent chroma residual scaling process according to the above embodiment can be described in the form of standard documents as shown in the following table.

[0383] [Table 31]

[0384]

[0385]

[0386] [Table 32]

[0387]

[0388]

[0389] Figure 14 An example of linear shaping (or linear shaping, linear mapping) according to an embodiment of this document is shown. That is, in this embodiment, it is proposed to use a linear shaper in LMCS. For example, Figure 14 This example in may involve forward linear shaping (mapping).

[0390] Reference Figure 14 The linear shaper may include two pivot points, P1 and P2. P1 and P2 may represent input values and mapped values. For example, P1 may be (min_input, 0) and P2 may be (max_input, max_mapped). Here, min_input represents the minimum input value, and max_input represents the maximum input value. Any input value less than or equal to min_input is mapped to 0, and any input value greater than max_input is mapped to max_mapped. Any input luminance value within min_input and max_input is linearly mapped to other values. Figure 14 An example of a mapping is shown. The pivot points P1 , P2 can be determined at the encoders, and a linear fit can be used to approximate a piecewise linear mapping.

[0391] In another embodiment according to this document, another example of a method for signaling a linear shaper can be provided. The pivot points P1 and P2 of the linear shaper model can be explicitly signaled. The following table shows an example of the syntax and semantics for explicitly signaling the linear shaper model according to this example.

[0392] [Table 33]

[0393]

[0394] [Table 34]

[0395]

[0396] Referring to Table 33 and Table 34, the input value of the first pivot point can be derived based on the syntactic element lmcs_min_input, and the input value of the second pivot point can be derived based on the syntactic element lmcs_max_input. The mapped value of the first pivot point can be a predetermined value (a value known to both the encoder and the decoder). For example, the mapped value of the first pivot point is 0. The mapped value of the second pivot point can be derived based on the syntactic element lmcs_max_mapped. That is, the linear shaping model can be explicitly (directly) signaled based on the information signaled in the syntax of Table 33.

[0397] Alternatively, lmcs_max_input and lmcs_max_mapped can be signaled as Δ values. The following table shows an example of the syntax and semantics of signaling the linear shaping model as Δ values.

[0398] [Table 35]

[0399]

[0400] [Table 36]

[0401]

[0402] Referring to Table 36, the input value of the first pivot point can be derived based on the syntactic element lmcs_min_input. For example, lmcs_min_input can have a mapped value of 0. lmcs_max_input_delta can specify the difference between the input value of the second pivot point and the maximum luminance value (i.e., (1 << bitdepthY)-1). lmcs_max_mapped_delta can specify the difference between the mapped value of the second pivot point and the maximum luminance value (i.e., (1 << bitdepthY)-1).

[0403] According to an embodiment of this document, forward mapping of luminance prediction samples, inverse mapping of luminance reconstruction samples, and chroma residual scaling can be performed based on the above examples of the linear shaper. In one example, for inverse scaling of luminance (reconstruction) samples (pixels) in the inverse mapping based on the linear shaper, only one inverse scaling factor may be required. The same applies to forward mapping and chroma residual scaling. That is, the steps of determining ScaleCoeff[i], InvScaleCoeff[i], and ChromaScaleCoeff[i] (where i is the bin index) can be replaced by only one single factor. Here, one single factor is the fixed-point representation of the (positive) slope or inverse slope of the linear mapping. In one example, the inverse luminance mapping scaling factor (the inverse scaling factor in the inverse mapping of luminance reconstruction samples) can be derived based on at least one of the following equations.

[0404] [Equation 17]

[0405] InvScaleCoeffSingle=OrgCW / lmcsCWLinear

[0406] [Equation 18]

[0407] InvScaleCoeffSingle=OrgCW*(lmcs_max_bin_idx-lmcs_max_bin_idx+1) / lmcsCWLinearAll

[0408] [Equation 19]

[0409] InvScaleCoeffSingle=(lmcs_max_input-lmcs_min_input) / lmcsCWLinearAll

[0410] lmcsCWLinear in Equation 17 can be derived from Table 31. lmcsCWLinearALL in Equations 18 and 19 can be derived from at least one of Tables 33 to 36. In Equation 17 or 18, OrgCW can be derived from Table 19 and / or Table 20.

[0411] The following table describes the formula and syntax (conditional statement) indicating the forward mapping processing of luma samples (i.e., luma prediction samples) in picture reconstruction. In the following table and formula, FP_PREC is a constant value of the bit shift and can be a predetermined value. For example, FP_PREC can be 11 or 15.

[0412] [Table 37]

[0413]

[0414] [Table 38]

[0415]

[0416] Table 37 can be used to derive luma samples for forward mapping in the luma mapping process based on Tables 17 to 20 described above. That is, Table 37 can be described together with Tables 19 and 20. In Table 37, luma (prediction) samples PredMapPSamples[i][j] for forward mapping as output can be derived from luma (prediction) samples predSamples[i][j] as input. idxY of Table 37 can be referred to as a (forward) mapping index, and the mapping index can be derived based on the predicted luma sample.

[0417] Table 38 can be used to derive the luminance samples of the forward mapping in the linear shaper-based luminance mapping. For example, lmcs_min_input, lmcs_max_input, lmcs_max_mapped, and ScaleCoeffSingle in Table 38 can be derived from at least one of Tables 33 to 36. In Table 38, when "lmcs_min_input < predSamples[i][j] < lmcs_max_input", the luminance (predicted) sample PredMapSamples[i][j] of the forward mapping can be derived from the input luminance (predicted) sample predSamples[i][j] as the output. By comparing between Table 37 and Table 38, the change with respect to the existing LMCS according to the application of the linear shaper can be seen from the perspective of the forward mapping.

[0418] The following formula and table describe the inverse mapping process of the luminance samples (i.e., luminance reconstruction samples). In the following formula and table, the "lumaSample" as the input can be the luminance reconstruction sample before the inverse mapping (before modification). The "invSample" as the output can be the inverse mapped (modified) luminance reconstruction sample. In other cases, the clipped invSample can be referred to as the modified luminance reconstruction sample.

[0419] [Equation 20]

[0420] invSample = InputPivot[idxYInv] + (InvScaleCoeff[idxYInv] * (lumaSample - LmcsPivot[idxYInv]) + (1 << (FP_PREC - 1))) >> FP_PREC

[0421] [Equation 21]

[0422] invSample = lmcs_min_input + (lnvScaleCoeffSingle * (lumaSample - lmcs_min_input) + (1 << (FP_PREC - 1))) >> FP_PREC

[0423] [Table 39]

[0424]

[0425] [Table 40]

[0426]

[0427] Equation 21 may be used to derive inverse-mapped luma samples in the luma mapping according to this document. In Equation 20, the index idxInv may be derived based on Table 50, Table 51, or Table 52 described later.

[0428] Equation 21 can be used to derive luma samples for reverse mapping from luma mapping based on the application of a linear shaper. For example, lmcs_min_input in Equation 21 can be derived from at least one of Tables 33 to 36. Comparing Equations 20 and 21 shows the changes from the conventional LMCS based on the application of a linear shaper from the perspective of forward mapping.

[0429] Table 39 may include an example of an equation for deriving luma samples for reverse mapping in luma mapping. For example, the index idxInv may be derived based on Table 50, Table 51, or Table 52 described later.

[0430] Table 40 may include other examples of equations for deriving inverse-mapped luma samples in luma mapping. For example, lmcs_min_input and / or lmcs_max_mapped in Table 40 may be derived from at least one of Tables 33 to 36, and / or InvScaleCoeffSingle in Table 40 may be derived from at least one of Tables 33 to 36, and / or Equations 17 to 19 may be derived from one.

[0431] Based on the above example of a linear shaper, the segment index identification process can be omitted. That is, in this example, since there is only one segment with valid shaped luma pixels, the segment index identification process for inverse luma mapping and chroma residual scaling can be eliminated. Consequently, the complexity of inverse luma mapping can be reduced. Furthermore, the delay caused by relying on luma segment index identification during chroma residual scaling can be eliminated.

[0432] According to an embodiment using the above-described linear shaper, the following advantages can be provided for LMCS: i) the encoder shaper design can be simplified, thereby preventing possible artifacts caused by abrupt changes between piecewise linear segments; ii) the decoder reverse mapping process can be simplified by eliminating the segment index identification process; iii) by eliminating the segment index identification process, the delay problem caused by the dependence on the corresponding luma block in the chroma residual scaling can be eliminated; iv) the signaling overhead can be reduced, making frequent updates of the shaper more feasible; v) for many places where a loop of 16 segments was previously required, the loop can be eliminated. For example, to derive InvScaleCoeff[i], the number of division operations according to lmcsCW[i] can be reduced to 1.

[0433] In another embodiment according to this document, an LMCS based on a flexible bin is proposed. Here, a flexible bin may refer to a bin whose quantity is not fixed to a predetermined (predefined, specific) quantity. In the existing embodiment, the number of bins in the LMCS is fixed to 16, and for the input sample values, these 16 bins are equally distributed. In this embodiment, a flexible number of bins is proposed, and these segments (bins) may not be equally distributed in terms of the original pixel values.

[0434] The following table exemplarily shows the syntax of the LMCS data (data field) according to this embodiment and the semantics of the syntax elements included therein.

[0435] [Table 41]

[0436]

[0437] [Table 42]

[0438]

[0439] Referring to Table 41, the information lmcs_num_bins_minus1 regarding the number of bins can be signaled. Referring to Table 42, lmcs_num_bins_minus1 + 1 can be equal to the number of bins, and the number of bins can be in the range from 1 to (1 << BitDepthY) - 1. For example, lmcs_num_bins_minus1 or lmcs_num_bins_minus1 + 1 can be a multiple of a power of 2.

[0440] In the embodiment described together with Table 41 and Table 42, regardless of whether the shaper is linear (the signaling of lmcs_num_bins_minus1), the number of pivot points can be derived based on lmcs_num_bins_minus1 (the information regarding the number of bins), and the input values and mapped values (LmcsPivot_input[i], LmcsPivot_mapped[i]) of the pivot points can be derived based on the sum of the signaled codeword values (lmcs_delta_input_cw[i], lmcs_delta_mapped_cw[i]) (here, the initial input value LmcsPivot_input[0] and the initial output value LmcsPivot_mapped[0] are 0).

[0441] Figure 15 An example of linear forward mapping in the embodiment of this document is shown. Figure 16 An example of inverse forward mapping in the embodiment of this document is shown.

[0442] In accordance with Figure 15 and Figure 16In an embodiment, a method is proposed to support both conventional LMCS and linear LMCS. In an example according to this embodiment, conventional LMCS and / or linear LMCS can be indicated based on the syntax element lmcs_is_linear. In the encoder, after determining the linear LMCS line, the mapping value (i.e., Figure 15 and Figure 16 The codeword in bin LmcsMaxBinIdx may be signaled using the syntax of the LMCS data or shaper mode described above.

[0443] The following table exemplarily shows the syntax of LMCS data (data field) and the semantics of syntax elements included therein according to an example of this embodiment.

[0444] [Table 43]

[0445]

[0446] [Table 44]

[0447]

[0448]

[0449] The following table exemplarily shows the syntax of LMCS data (data field) and the semantics of syntax elements included therein according to another example of this embodiment.

[0450] [Table 45]

[0451]

[0452] [Table 46]

[0453]

[0454]

[0455] Referring to Tables 43 to 46, when lmcs_is_linear_flag is true, all LMCSDeltaCW[i] between lmcs_min_bin_idx and LmcsMaxBinIdx may have the same value. That is, lmcsCW[i] may have the same value for all segments between lmcs_min_bin_idx and LmcsMaxBinIdx. The scale, descaling, and chroma scale may be the same for all segments between lmcs_min_bin_idx and lmcsMaxBinIdx. However, if linear_reshape is true, there is no need to derive the segment index; the scale and descaling from just one segment may be used.

[0456] The following table exemplarily shows the identification process of the segment index according to this embodiment.

[0457] [Table 47]

[0458]

[0459] According to another embodiment of this document, the application of conventional 16-segment PWL LMCS and linear LMCS may depend on a higher level syntax (ie, sequence level).

[0460] The following table exemplarily shows the syntax of the SPS according to the present embodiment and the semantics of the syntax elements included therein.

[0461] [Table 48]

[0462]

[0463] [Table 49]

[0464]

[0465] Referring to Tables 48 and 49, enabling of normal LMCS and / or linear LMCS may be determined (signaled) by syntax elements included in the SPS. Referring to Table 48, based on the syntax element sps_linear_lmcs_enabled_flag, one of normal LMCS or linear LMCS may be used in units of sequences.

[0466] In addition, whether only linear LMCS, conventional LMCS, or both are enabled may also depend on the profile level. In one example, for a particular profile (i.e., an SDR profile), only linear LMCS may be allowed, for another profile (i.e., an HDR profile), only conventional LMCS may be allowed, and for another profile, both conventional LMCS and / or linear LMCS may be allowed.

[0467] According to another embodiment of the present document, the LMCS segment index identification process can be used in inverse luma mapping and chroma residual scaling. In this embodiment, the segment index identification process can be used for blocks where chroma residual scaling is enabled, and is also called for all luma samples in the shaping (mapping) domain. This embodiment aims to keep its complexity low.

[0468] The following shows the identification process (derivation process) of the existing piecewise function index.

[0469] [Table 50]

[0470]

[0471] In an example, in the segment index identification process, the input samples may be classified into at least two categories. For example, the input samples may be classified into three categories, namely, first, second, and third categories. For example, the first category may indicate samples (values) less than LmcsPivot[lmcs_min_bin_idx+1], the second category may indicate samples (values) greater than or equal to LmcsPivot[LmcsMaxBinIdx] (values), and the third category may indicate samples (values) between LmcsPivot[lmcs_min_bin_idx+1] and LmcsPivot[LmcsMaxBinIdx].

[0472] In this embodiment, the recognition process is optimized by eliminating class classification. This is because the input to the segment index recognition process is the luminance value in the shaped (mapped) domain, and no value should exceed the mapped value at the pivot points lmcs_min_bin_idx and LmcsMaxBinIdx+1. Therefore, the conditional processing used in the existing segment index recognition process to classify samples into classes is unnecessary. For more details, a specific example is described in the table below.

[0473] In an example according to this embodiment, the identification process included in Table 50 can be replaced by one of the following Tables 51 or 52. Referring to Tables 51 and 52, the first two categories in Table 50 can be removed, and for the last category, the boundary value (second boundary value or end point) in the iterative for loop is changed from LmcsMaxBinIdx to LmcsMaxBinIdx+1. That is, the identification process can be simplified and the complexity of segment index derivation can be reduced. Therefore, according to this embodiment, LMCS-related encoding can be performed efficiently.

[0474] [Table 51]

[0475]

[0476] [Table 52]

[0477]

[0478] Referring to Table 51, the comparison process corresponding to the condition of the if statement (the expression corresponding to the condition of the if statement) may be iteratively performed on all bin indices from the minimum bin index to the maximum bin index. When the expression corresponding to the condition of the if statement is true, the bin index may be derived as a reverse mapping index for reverse luma mapping (or a reverse scaling index for chroma residual scaling). Based on the reverse mapping index, a modified reconstructed luma sample (or a scaled chroma residual sample) may be derived.

[0479] According to the embodiment of Table 51, problems that may occur during the process of identifying (reverse) piecewise function indices can be solved. Due to the embodiment of Table 51, arithmetic loopholes can be eliminated and / or repeated operation processing due to overlapping boundary conditions can be omitted. If this document does not follow the embodiment of Table 51, the mapping value used in the LMCS of the current block in the for syntax (loop syntax) for identifying the segment index (reverse mapping index) may exceed (deviate from) LmcsPivot[idxYInv+1]. According to this embodiment, a mapping value within the appropriate range used in the LMCS of the current block during the index identification process can be used.

[0480] The following table shows the impact of improved coding performance through the implementation of Table 51.

[0481] [Table 53]

[0482]

[0483]

[0484] [Table 54]

[0485]

[0486] [Table 55]

[0487]

[0488] Referring to Tables 53 to 55, encoding performance can be improved according to the embodiment of Table 51. In addition, arithmetic loopholes can be removed while maintaining encoding performance according to the embodiment of Table 51, and / or overlapping arithmetic processing can be omitted.

[0489] In an embodiment according to the present document, relative to slices coded via a separate block tree (e.g., intra slices) in existing embodiments, the delay caused by the chroma residual scaling dependency in the corresponding luma block is higher than for slices coded via dual tree. Therefore, in existing embodiments, LMCS chroma residual scaling is not applied to slices coded via a separate block tree.

[0490] In this embodiment, chroma residual scaling can be applied even for slices coded by separate trees. When a single chroma residual scaling factor is used as described above, there is no dependency between the chroma residual scaling in the corresponding luminance blocks, so there may be no delay in the application of chroma residual scaling.

[0491] The following table shows the syntax and semantics of the slice header according to this embodiment.

[0492] [Table 56]

[0493]

[0494] [Table 57]

[0495]

[0496] Referring to Table 57, regardless of the conditional clause (or its flag) indicating whether the current block has a dual-tree structure or a single-tree structure, a chroma residual scaling flag may be signaled.

[0497] In embodiments according to this document, ALF data and / or LMCS data can be signaled in an APS. For example, 32 APSs can be used. In an example, if all APSs are used for ALF and / or LMCS, the APS buffering may require approximately 10KB of on-chip memory. To limit (reduce) the memory required to store ALF / LMCS parameters and the computational complexity required for LMCS, this embodiment proposes limiting the number of ALF and / or LMCS APSs.

[0498] In the example according to this embodiment, regardless of the number of slices or tiles included in a picture, one LMCS model per picture may be used (allowed). The number of APSs used for LMCS may be less than 32. For example, the number of APSs used for LMCS may be 4.

[0499] The following table shows the semantics related to APS according to this embodiment.

[0500] [Table 58]

[0501]

[0502] Referring to Table 58, the maximum number of LMCS APSs may be predetermined. For example, the maximum number of LMCS APSs may be 4. Referring to Tables 15 and 58, a plurality of APSs may include an LMCS APS. The (maximum) number of LMCS APSs may be 4. In this example, the LMCS data field included in one of the LMCS APSs may be used in the LMCS process of the current block in the current picture.

[0503] The following table represents the semantics of syntax elements included in a slice header (or picture header).

[0504] [Table 59]

[0505]

[0506] The syntax elements described in Table 59 may be described with reference to Table 56. In an example, the syntax element slice_lmcs_aps_id may be included in a slice header. In another example, the syntax element slice_lmcs_aps_id of Table 59 may be included in a picture header, and in this case, slice_lmcs_aps_id may be modified to ph_lmcs_aps_id.

[0507] According to the embodiments of this document, the total number of LMCS APSs can be equal to or less than four. Furthermore, in the example of this embodiment, only one LMCS model can be used for each screen. In this case, regardless of the image / video resolution, only one LMCS model can be used for each screen. According to this embodiment, due to the limitation of LMCS APSs, implementation is facilitated and resource (memory) overrun issues can be resolved.

[0508] The following table represents the semantics of examples of syntax elements disclosed in this document.

[0509] [Table 60]

[0510]

[0511] The syntax elements described in Table 60 may be described with reference to Table 56. In an example, the syntax element slice_lmcs_aps_id may be included in a slice header. In another example, the syntax element slice_lmcs_aps_id of Table 60 may be included in a picture header, and in this case, slice_lmcs_aps_id may be modified to ph_lmcs_aps_id.

[0512] Referring to Table 60, the adaptation_parameter_set_id of the LMCS APS referenced by the current slice (or current picture) may be identified by the syntax element slice_lmcs_aps_id. That is, slice_lmcs_aps_id may represent the adaptation_parameter_set_id of the LMCS APS referenced by the current slice (or current picture). The value of slice_lmcs_aps_id may be in the range of 0 to 3. That is, the value of slice_lmcs_aps_id may be 0, 1, 2, or 3. For example, the 0th LMCS APS (or the first LMCS APS) indicated by slice_lmcs_aps_id among the 0th to third LMCS APSs (or the first to fourth LMCS APSs) can be referenced by the current slice (or current picture), and the LMCS data (LMCS data field or shaper model) included in the 0th LMCS APS (or the first LMCS APS) can be used for the LMCS process of the current slice (or current picture) (the LMCS process of the current slice (or current picture) can be performed based on the LMCS data (LMCS data field or shaper model) included in the 0th LMCS APS (or the first LMCS APS)).

[0513] According to at least one embodiment disclosed in this document, the LMCS-related encoding process can be cleaned up or simplified.

[0514] The following figures are created to illustrate specific examples of this specification. Since the names of specific devices or the names of specific signals / messages / fields described in the figures are presented as examples, the technical features of this specification are not limited to the specific names used in the following figures.

[0515] Figure 17 and Figure 18 An example of a video / image encoding method and related components according to an embodiment of this document is schematically shown. Figure 17 The method disclosed in Figure 2 Specifically, for example, Figure 17 S1700 and S1710 of the encoding device may be performed by the predictor 220 of the encoding device, S1720 may be performed by the residual processor 230 of the encoding device, and S1730 may be performed by the predictor 220 or the residual processor 230 of the encoding device. S1740 may be performed by the residual processor 230 or the adder 250 of the encoding device, S1750 may be performed by the residual processor 230 of the encoding device, and S1760 may be performed by the entropy encoder 240 of the encoding device. Figure 17 The methods disclosed in may include the embodiments described in detail above in this document.

[0516] Reference Figure 17 , the encoding device may derive an inter-frame prediction mode of a current block in a current picture (S1700). The encoding device may derive at least one of the various modes disclosed in this document among the inter-frame prediction modes.

[0517] The encoding apparatus may generate predicted luma samples based on the inter prediction mode (S1710). The encoding apparatus may generate the predicted luma samples by performing prediction on original samples included in the current block.

[0518] The encoding device may derive the predicted chroma samples. The encoding device may derive the residual chroma samples based on the original chroma samples and the predicted chroma samples of the current block. For example, the encoding device may derive the residual chroma samples based on the difference between the predicted chroma samples and the original chroma samples.

[0519] The encoding device may derive bins and LMCS codewords for luma mapping. The encoding device may derive bins and / or LMCS codewords based on SDR or HDR.

[0520] The encoding apparatus may derive LMCS-related information (S1720). The LMCS-related information may include a bin for luma mapping and an LMCS codeword.

[0521] The encoding device may generate a mapped predicted luma sample based on the mapping process of the luma sample (S1730). The encoding device may generate the mapped predicted luma sample based on the bin and / or LMCS codeword for luma mapping. For example, the encoding device may derive an input value and a mapping value (output value) of a pivot point for luma mapping, and may generate the mapped predicted luma sample based on the input value and the mapping value. In an example, the encoding device may derive a mapping index idxY based on a first predicted luma sample, and may generate a first mapped predicted luma sample based on the input value and the mapping value of the pivot point corresponding to the mapping index. In another example, a linear mapping (linear shaping or linear LMCS) may be used, and the mapped predicted luma sample may be generated based on a forward mapping scaling factor derived from two pivot points in the linear mapping. Therefore, due to the linear mapping, the index derivation process may be omitted.

[0522] The encoding device may generate scaled residual chroma samples. Specifically, the encoding device may derive a chroma residual scaling factor and generate scaled residual chroma samples based on the chroma residual scaling factor. Here, chroma residual scaling at the encoding level may be referred to as forward chroma residual scaling. Therefore, the chroma residual scaling factor derived by the encoding device may be referred to as a forward chroma residual scaling factor, and forward scaled residual chroma samples may be generated.

[0523] The encoding device may generate a reconstructed luma sample. The encoding device may generate the reconstructed luma sample based on the mapped predicted luma sample. Specifically, the encoding device may add the residual luma sample to the mapped predicted luma sample, and may generate the reconstructed luma sample based on the added result.

[0524] The encoding device may generate modified reconstructed luma samples based on the reverse mapping process of the luma samples (S1740). The encoding device may generate the modified reconstructed luma samples based on the bins used for luma mapping, the LMCS codewords, and the reconstructed luma samples. The encoding device may generate the modified reconstructed luma samples by performing a reverse mapping process on the reconstructed luma samples. For example, the encoding device may derive a reverse mapping index (e.g., invYIdx) based on the reconstructed luma samples and / or the mapping values assigned to the respective bin indices (e.g., LmcsPivot[i], i=lmcs_min_bin_idx...LmcsMaxBinIdx+1) in the reverse mapping process. The encoding device may generate the modified reconstructed luma samples based on the mapping value LmcsPivot[invYIdx] assigned to the reverse mapping index.

[0525] The encoding device may generate residual luma samples based on the mapped predicted luma samples. For example, the encoding device may derive residual luma samples based on the difference between the mapped predicted luma samples and the original luma samples.

[0526] The encoding device may derive residual information (S1750). In an example, the encoding device may generate residual information based on the mapped predicted luma samples and the modified reconstructed luma samples. For example, the encoding device may derive residual information based on scaled residual chroma samples and / or residual luma samples. The encoding device may derive transform coefficients based on a transform process of the scaled residual chroma samples and / or residual luma samples. For example, the transform process may include at least one of DCT, DST, GBT, or CNT. The encoding device may derive quantized transform coefficients based on a quantization process of the transform coefficients. The quantized transform coefficients may have a one-dimensional vector form based on a coefficient scanning order. The encoding device may generate residual information for a specified quantized transform coefficient. The residual information may be generated by various encoding methods such as exponential Golomb, CAVLC, CABAC, etc.

[0527] The encoding device may encode the image / video information (S1760). The image information may include information about LMCS data and / or residual information. For example, the LMCS-related information may include information about linear LMCS. In one example, at least one LMCS codeword may be derived based on the information about linear LMCS. The encoded video / image information may be output in the form of a bitstream. The bitstream may be transmitted to a decoding device via a network or storage medium.

[0528] According to an embodiment of this document, the image / video information may include various types of information. For example, the image / video information may include the information disclosed in at least one of Tables 1 to 60 above.

[0529] In an embodiment, the image information may include an LMCS APS. The LMCS APS may include an LMCS data field lmcs_data() and / or ID information aps_adaptation_parameter_set_id. The value of the ID information may be within a predetermined range. For example, the value of the ID information may be within a range of 0 to 3. That is, the value of the ID information may be 0, 1, 2, or 3. The LMCS data field may include LMCS-related information. Based on the LMCS data field, the LMCS codeword used in the mapping process and the reverse mapping process may be derived. The maximum number of LMCS APSs may be a predetermined value. For example, the maximum number of LMCS APSs (predetermined value) may be 4.

[0530] In an embodiment, in a mapping process for luma samples, a mapping index may be derived based on the predicted luma sample, and a mapped predicted luma sample may be generated using a first mapping value based on the mapping index.

[0531] In an embodiment, the LMCS-related information may include information about bins used for reverse mapping. A minimum bin index and a maximum bin index may be derived based on the information about the bins. In the reverse mapping process for luma samples, a reverse mapping index may be derived based on the bin indices from the minimum bin index to the maximum bin index based on the mapping values, and a modified reconstructed luma sample may be generated using the second mapping value based on the reverse mapping index.

[0532] In an embodiment, the encoding device may generate a segment index for chroma residual scaling, derive a chroma residual scaling factor based on the segment index, and generate scaled residual chroma samples based on the residual chroma samples and the chroma residual scaling factor.

[0533] In an embodiment, when the current block has a single tree structure or a dual tree structure (when the current block has a separate tree structure, or when the current block is encoded using a separate tree), the encoding device may generate a chroma residual scaling available flag indicating whether chroma residual scaling is applied to the current block. Furthermore, regardless of the block tree structure of the current block, the chroma residual scaling available flag indicating whether chroma residual scaling is applied to the current block may be generated. When chroma residual scaling is applied to the current picture, current slice, and / or current block, the value of the chroma residual scaling available flag may be 1.

[0534] In an embodiment, the chroma residual scaling factor may be a single chroma residual scaling factor.

[0535] In an embodiment, based on the value of the type information being 1, the APS may include an LMCS data field including LMCS parameters.

[0536] In an embodiment, the image information may include header information. Here, the header information may be a picture header (or slice header). The header information may include LMCS-related APS ID information. In an example, the LMCS-related APS ID information may indicate the ID of the LMCS APS of the current picture or current block. In another example, the value of the LMCS-related APS ID information may be equal to the value of the ID information. The value of the LMCS-related APS ID information may be in the range of 0 to 3.

[0537] In an embodiment, the image information may include a sequence parameter set (SPS). The SPS may include a linear LMCS available flag indicating whether the linear LMCS is available.

[0538] In an embodiment, a minimum bin index (e.g., lmcs_min_bin_idx) and / or a maximum bin index (e.g., LmcsMaxBinIdx) may be derived based on information about the LMCS data. Based on the minimum bin index, a first mapping value LmcsPivot[lmcs_min_bin_idx] may be derived. A second mapping value LmcsPivot[LmcsMaxBinIdx] or LmcsPivot[LmcsMaxBinIdx+1] may be derived based on the maximum bin index. The value of the reconstructed luma sample (e.g., lumaSample of Table 36 or 37) may be in the range of a first mapping value to a second mapping value. In an example, the values of all reconstructed luma samples may be in the range of a first mapping value to a second mapping value. In another example, some sample values of the reconstructed luma samples may be in the range of a first mapping value to a second mapping value.

[0539] In an embodiment, the information regarding LMCS data may include an LMCS data field and information regarding linear LMCS. The information regarding linear LMCS may be referred to as information regarding linear mapping. The LMCS data field may include a linear LMCS flag indicating whether linear LMCS is applied. If the linear LMCS flag has a value of 1, a mapped predicted luma sample may be generated based on the information regarding linear LMCS.

[0540] In embodiments, information about the linear LMCS may include information about the first pivot point (e.g., Figure 11 of P1) and information about the second pivot point (e.g., Figure 11For example, the input value and mapped value of the first pivot point may be the minimum input value and the minimum mapped value, respectively. The input value and mapped value of the second pivot point may be the maximum input value and the maximum mapped value, respectively. Input values between the minimum input value and the maximum input value may be linearly mapped.

[0541] In one embodiment, the image information includes information about a maximum input value and information about a maximum mapping value. The maximum input value is equal to the value of the information about the maximum input value (i.e., lmcs_max_input in Table 33). The maximum mapping value is equal to the value of the information about the maximum mapping value (i.e., lmcs_max_mapped in Table 33).

[0542] In one embodiment, the information about the linear mapping includes information about the input delta value of the second pivot point (i.e., lmcs_max_input_delta in Table 35) and information about the mapped delta value of the second pivot point (i.e., lmcs_max_mapped_delta in Table 35). The maximum input value may be derived based on the input delta value of the second pivot point, and the maximum mapped value may be derived based on the mapped delta value of the second pivot point.

[0543] In one embodiment, the maximum input value and the maximum mapping value may be derived based on at least one equation included in Table 36 above.

[0544] In one embodiment, generating the mapped predicted luma samples includes: deriving a forward mapping scaling factor (i.e., ScaleCoeffSingle) for the predicted luma samples; and generating the mapped predicted luma samples based on the forward mapping scaling factor. The forward mapping scaling factor may be a single factor for the predicted luma samples.

[0545] In one embodiment, the forward mapping scaling factor may be derived based on at least one equation included in Table 36 and / or Table 38 above.

[0546] In one embodiment, the mapped predicted luma samples may be derived based on at least one equation included in Table 38 above.

[0547] In one embodiment, the encoding device may derive an inverse mapping scaling factor (i.e., InvScaleCoeffSingle) for the reconstructed luma sample (i.e., lumaSample). Additionally, the encoding device may generate a modified reconstructed luma sample (i.e., invSample) based on the reconstructed luma sample and the inverse mapping scaling factor. The inverse mapping scaling factor may be a single factor for the reconstructed luma sample.

[0548] In one embodiment, the segment index derived based on the reconstructed luma samples may be used to derive the reverse mapping scaling factor.

[0549] In one embodiment, the segment index may be derived based on Table 51 above. That is, the comparison process (lumaSample < LmcsPivot[idxYInv + 1]) included in Table 51 may be iteratively performed from the segment index as the minimum bin index to the segment index as the maximum bin index.

[0550] In one embodiment, the inverse mapping scaling factor may be derived based on at least one of the formulas included in Table 33, Table 34, Table 35, and Table 36 or Formula 11 or Formula 12 above.

[0551] In one embodiment, the modified reconstructed luma sample may be derived based on Formula 20, Formula 21, Table 39, and / or Table 40 above.

[0552] In one embodiment, the LMCS-related information may include information about the number of bins for the predicted luma samples used for derivation of the mapping (i.e., lmcs_num_bins_minus1 in Table 41). For example, the number of pivot points for the luma mapping may be set to be equal to the number of bins. In one example, the encoding device may generate the Δ input value and Δ mapping value of the pivot points respectively according to the number of bins. In one example, the input value and mapping value of the pivot points are derived based on the Δ input value (i.e., lmcs_delta_input_cw[i] in Table 41) and the Δ mapping value (i.e., lmcs_delta_mapped_cw[i] in Table 41), and the mapped predicted luma samples may be generated based on the input value (i.e., LmcsPivot_input[i] in Table 42) and the mapping value (i.e., LmcsPivot_mapped[i] in Table 42).

[0553] In one embodiment, the encoding device may derive the LMCS Δ codeword based on at least one LMCS codeword and the original codeword (OrgCW) included in the LMCS-related information, and may derive the mapped luma prediction samples based on at least one LMCS codeword and the original codeword. In one example, the information about the linear mapping may include the information about the LMCS Δ codeword.

[0554] In one embodiment, at least one LMCS codeword may be derived based on the sum of the LMCS Δ codeword and OrgCW. For example, OrgCW is (1 << BitDepthY) / 16, where BitDepthY represents the luma bit depth. This embodiment may be based on Formula 12.

[0555] In one embodiment, at least one LMCS codeword may be derived based on the sum of the LMCSΔ codeword and OrgCW*(lmcs_max_bin_idx - lmcs_min_bin_idx + 1), where, for example, lmcs_max_bin_idx and lmcs_min_bin_idx are the maximum bin index and the minimum bin index respectively, and OrgCW may be (1<<BitDepthY) / 16. This embodiment may be based on Equation 15 and Equation 16.

[0556] In one embodiment, at least one LMCS codeword may be a multiple of 2.

[0557] In one embodiment, when the luminance bit depth (BitDepthY) of the reconstructed luminance samples is higher than 10, at least one LMCS codeword may be a multiple of 1<<(BitDepthY - 10).

[0558] In one embodiment, at least one LMCS codeword may be in the range from (OrgCW>>1) to (OrgCW<<1)-1.

[0559] In the above paragraphs, the information about the LMCS data may be the same as the information about the LMCS.

[0560] Figure 19 and Figure 20 FIG. schematically shows an example of an image / video decoding method and related components according to an embodiment of this document. Figure 19 The method disclosed in Figure 3 may be performed by the decoding device disclosed in Figure 19 Specifically, for example, S1900 of Figure 19 may be performed by the entropy decoder 310 of the decoding device, S1910 and S1920 may be performed by the predictor 330 of the decoding device, S1930 may be performed by the residual processor 320 or the predictor 330 of the decoding device, and S1940 may be performed by the residual processor 320, the predictor 330 and / or the adder 340 of the decoding device.

[0561] Referring to Figure 19 , the decoding device may receive / acquire video / image information (S1900). The video / image information may include information about the LMCS data and / or residual information. For example, the information about the LMCS data may include information about luminance mapping (i.e., forward mapping, inverse mapping, linear mapping), information about chrominance residual scaling, and / or indices related to the LMCS (or shaping, shaper) (i.e., maximum bin index, minimum bin index, mapping index). The decoding device may receive / acquire the image / video information through the bitstream.

[0562] According to an embodiment of this document, the image / video information may include various types of information. For example, the image / video information may include the information disclosed in at least one of Tables 1 to 60 above.

[0563] The decoding apparatus may derive a prediction mode of a current block in a current picture based on the prediction mode information (S1910). The decoding apparatus may derive at least one of various modes disclosed in this document among inter-prediction modes.

[0564] The decoding device may generate predicted luma samples. The decoding device may derive predicted luma samples of the current block based on the prediction mode. In this case, various prediction methods disclosed in this document, such as inter-frame prediction or intra-frame prediction, may be applied.

[0565] The decoding apparatus may generate a predicted luma sample (S1920). The decoding apparatus may derive the predicted luma sample of the current block based on the prediction mode. The decoding apparatus may generate the predicted luma sample by performing prediction of the original sample included in the current block.

[0566] The image information may include residual information. The decoding device may generate residual chroma samples based on the residual information. Specifically, the decoding device may derive quantized transform coefficients based on the residual information. The quantized transform coefficients may have a one-dimensional vector form based on a coefficient scanning order. The decoding device may derive the transform coefficients based on a dequantization process of the quantized transform coefficients. The decoding device may derive residual chroma samples and / or residual luma samples based on the transform coefficients.

[0567] The decoding device may generate a mapped predicted luma sample (S1930). The decoding device may generate the mapped predicted luma sample based on a mapping process of the luma sample. For example, the decoding device may derive an input value and a mapping value (output value) of a pivot point of the luma mapping, and may generate the mapped predicted luma sample based on the input value and the mapping value. In one example, the decoding device may derive a (forward) mapping index (idxY) based on the first predicted luma sample, and may generate the first mapped predicted luma sample based on the input value and the mapping value of the pivot point corresponding to the mapping index. In other examples, a linear mapping (linear shaping, linear LMCS) may be used, and the mapped predicted luma sample may be generated based on a forward mapping scaling factor derived from two pivot points in the linear mapping, so that the index derivation process may be omitted due to the linear mapping.

[0568] The decoding device may generate residual luma samples based on the residual information. For example, the decoding device may derive quantized transform coefficients based on the residual information. The quantized transform coefficients may have a one-dimensional vector form based on a coefficient scanning order. The decoding device may derive the transform coefficients based on dequantization processing of the quantized transform coefficients. The decoding device may derive residual samples based on the dequantization processing of the transform coefficients. The residual samples may include residual luma samples and / or residual chroma samples.

[0569] The decoding device may generate a reconstructed luma sample. The decoding device may generate a reconstructed luma sample of the current block based on the mapped predicted luma sample. Specifically, the decoding device may sum the residual luma sample with the mapped predicted luma sample, and may generate the reconstructed luma sample based on the summation result.

[0570] The decoding device may generate modified reconstructed luma samples. The decoding device may generate the modified reconstructed luma samples based on an inverse mapping process of the luma samples. The decoding device may generate the modified reconstructed luma samples based on the LMCS-related information and the reconstructed luma samples (S1940). The decoding device may generate the modified reconstructed luma samples by an inverse mapping process of the reconstructed luma samples.

[0571] The decoding device may generate scaled residual chroma samples. Specifically, the decoding device may derive a chroma residual scaling factor and generate scaled residual chroma samples based on the chroma residual scaling factor. Here, chroma residual scaling on the decoding side, as opposed to the encoding side, may be referred to as inverse chroma residual scaling. Therefore, the chroma residual scaling factor derived by the decoding device may be referred to as an inverse chroma residual scaling factor, and inversely scaled residual chroma samples may be generated.

[0572] The decoding device may generate reconstructed chroma samples. The decoding device may generate the reconstructed chroma samples based on the scaled residual chroma samples. Specifically, the decoding device may perform a prediction process on the chroma components and may generate predicted chroma samples. The decoding device may generate the reconstructed chroma samples based on the sum of the predicted chroma samples and the scaled residual chroma samples.

[0573] In an embodiment, the image information may include an LMCS APS. The LMCS APS may include an LMCS data field lmcs_data() and / or ID information aps_adaptation_parameter_set_id. The value of the ID information may be within a predetermined range. For example, the value of the ID information may be within a range of 0 to 3. That is, the value of the ID information may be 0, 1, 2, or 3. The LMCS data field may include LMCS-related information. Based on the LMCS data field, the LMCS codeword used in the mapping process and the reverse mapping process may be derived. The maximum number of LMCS APSs may be a predetermined value. For example, the maximum number of LMCS APSs (predetermined value) may be 4.

[0574] In an embodiment, in a mapping process for luma samples, a mapping index may be derived based on the predicted luma sample, and a mapped predicted luma sample may be generated using a first mapping value based on the mapping index.

[0575] In an embodiment, the LMCS-related information may include information about bins used for reverse mapping. A minimum bin index and a maximum bin index may be derived based on the information about the bins. In the reverse mapping process for luma samples, a reverse mapping index may be derived based on the bin indices from the minimum bin index to the maximum bin index based on the mapping values, and a modified reconstructed luma sample may be generated using the second mapping value based on the reverse mapping index.

[0576] In an embodiment, a segment index may be identified based on information about the LMCS data (e.g., idxYInv in Table 35, Table 36, or Table 37). The decoding device may derive a chroma residual scaling factor based on the segment index. The decoding device may generate scaled residual chroma samples based on the residual chroma samples and the chroma residual scaling factor.

[0577] In an embodiment, when the current block has a single tree structure or a dual tree structure (when the current block has a separate tree structure or when the current block is encoded using a separate tree), a chroma residual scaling available flag indicating whether chroma residual scaling is applied to the current block may be signaled. Furthermore, regardless of the block tree structure of the current block, a chroma residual scaling available flag indicating whether chroma residual scaling is applied to the current block may be signaled. When chroma residual scaling is applied to the current picture, current slice, and / or current block, the value of the chroma residual scaling available flag may be 1.

[0578] In an embodiment, the chroma residual scaling factor may be a single chroma residual scaling factor.

[0579] In an embodiment, based on the value of the type information being 1, the APS may include an LMCS data field including LMCS parameters.

[0580] In an embodiment, the image information may include header information. Here, the header information may be a picture header (or slice header). The header information may include LMCS-related APS ID information. In an example, the LMCS-related APS ID information may indicate the ID of the LMCS APS of the current picture or current block. In another example, the value of the LMCS-related APS ID information may be equal to the value of the ID information. The value of the LMCS-related APS ID information may be in the range of 0 to 3.

[0581] In an embodiment, a minimum bin index (e.g., lmcs_min_bin_idx) and / or a maximum bin index (e.g., LmcsMaxBinIdx) may be derived based on information about the LMCS data. Based on the minimum bin index, a first mapping value LmcsPivot[lmcs_min_bin_idx] may be derived. A second mapping value LmcsPivot[LmcsMaxBinIdx] or LmcsPivot[LmcsMaxBinIdx+1] may be derived based on the maximum bin index. The value of the reconstructed luma sample (e.g., lumaSample of Table 51 or Table 52) may be within a range from a first mapping value to a second mapping value. In an example, the values of all reconstructed luma samples may be within a range from a first mapping value to a second mapping value. In another example, some sample values of the reconstructed luma samples may be within a range from a first mapping value to a second mapping value.

[0582] In an embodiment, the image information may include a sequence parameter set (SPS). The SPS may include a linear LMCS available flag indicating whether the linear LMCS is available.

[0583] In an embodiment, the chroma residual scaling factor may be a single chroma residual scaling factor.

[0584] In an embodiment, the information regarding LMCS data may include an LMCS data field and information regarding linear LMCS. The information regarding linear LMCS may be referred to as information regarding linear mapping. The LMCS data field may include a linear LMCS flag indicating whether linear LMCS is applied. If the linear LMCS flag has a value of 1, a mapped predicted luma sample may be generated based on the information regarding linear LMCS.

[0585] In embodiments, information about the linear LMCS may include information about the first pivot point (e.g., Figure 11 of P1) and information about the second pivot point (e.g., Figure 11 For example, the input value and mapped value of the first pivot point may be the minimum input value and the minimum mapped value, respectively. The input value and mapped value of the second pivot point may be the maximum input value and the maximum mapped value, respectively. Input values between the minimum input value and the maximum input value may be linearly mapped.

[0586] In an embodiment, the image information may include information about a maximum input value and information about a maximum mapping value. The maximum input value may be equal to the value of the information about the maximum input value (e.g., lmcs_max_input in Table 33). The maximum mapping value may be equal to the value of the information about the maximum mapping value (e.g., lmcs_max_mapped in Table 33).

[0587] In an embodiment, the information about the linear mapping may include information about the input delta value of the second pivot point (e.g., lmcs_max_input_delta of Table 35) and information about the mapped delta value of the second pivot point (e.g., lmcs_max_mapped_delta of Table 35). The maximum input value may be derived based on the input delta value of the second pivot point, and the maximum mapped value may be derived based on the mapped delta value of the second pivot point.

[0588] In an embodiment, the maximum input value and the maximum mapping value may be derived based on at least one equation included in the above-mentioned Table 36.

[0589] In one embodiment, generating the mapped predicted luma samples includes: deriving a forward mapping scaling factor (i.e., ScaleCoeffSingle) for the predicted luma samples; and generating the mapped predicted luma samples based on the forward mapping scaling factor. The forward mapping scaling factor may be a single factor for the predicted luma samples.

[0590] In one embodiment, the segment index derived based on the reconstructed luma samples may be used to derive the reverse mapping scaling factor.

[0591] In one embodiment, the segment index may be derived based on the above-mentioned Table 51. That is, the comparison process (lumaSample <LmcsPivot[idxYInv+1])。

[0592] In one embodiment, the forward mapping scaling factor may be derived based on at least one equation included in Table 36 and / or Table 38 above.

[0593] In one embodiment, the mapped predicted luma samples may be derived based on at least one equation included in Table 38 above.

[0594] In one embodiment, the decoding device may derive an inverse mapping scaling factor (i.e., InvScaleCoeffSingle) for the reconstructed luma sample (i.e., lumaSample). Additionally, the decoding device may generate a modified reconstructed luma sample (i.e., invSample) based on the reconstructed luma sample and the inverse mapping scaling factor. The inverse mapping scaling factor may be a single factor for the reconstructed luma sample.

[0595] In one embodiment, the reverse mapping scaling factor may be derived based on at least one equation included in Table 33, Table 34, Table 35, and Table 36 above, or Equation 11, or Equation 12.

[0596] In one embodiment, the modified reconstructed luminance samples may be derived based on Equation 20, Equation 21, Table 39, and / or Table 40 described above.

[0597] In one embodiment, the LMCS related information may include information about the number of bins for the predicted luminance samples used to derive the mapping (i.e., lmcs_num_bins_minus1 in Table 41). For example, the number of pivot points of the luminance mapping may be set to be equal to the number of bins. In one example, the decoding device may generate the Δ input values and Δ mapping values of the pivot points according to the number of bins, respectively. In one example, the input values and mapping values of the pivot points are derived based on the Δ input values (i.e., lmcs_delta_input_cw[i] in Table 41) and the Δ mapping values (i.e., lmcs_delta_mapped_cw[i] in Table 41), and the predicted luminance samples of the mapping may be generated based on the input values (i.e., LmcsPivot_input[i] in Table 42) and the mapping values (i.e., LmcsPivot_mapped[i] in Table 42).

[0598] In one embodiment, the decoding device may derive the LMCS Δ codeword based on at least one LMCS codeword and the original codeword (OrgCW) included in the LMCS related information, and may derive the mapped luminance prediction samples based on at least one LMCS codeword and the original codeword. In one example, the information about the linear mapping may include the information about the LMCS Δ codeword.

[0599] In one embodiment, at least one LMCS codeword may be derived based on the sum of the LMCS Δ codeword and OrgCW. For example, OrgCW is (1<<BitDepthY) / 16, where BitDepthY represents the luminance bit depth. This embodiment may be based on Equation 14.

[0600] In one embodiment, at least one LMCS codeword may be derived based on the sum of the LMCS Δ codeword and OrgCW*(lmcs_max_bin_idx - lmcs_min_bin_idx + 1). For example, lmcs_max_bin_idx and lmcs_min_bin_idx are the maximum bin index and the minimum bin index, respectively, and OrgCW may be (1<<BitDepthY) / 16. This embodiment may be based on Equation 15 and Equation 16.

[0601] In one embodiment, at least one LMCS codeword may be a multiple of 2.

[0602] In one embodiment, when the luma bit depth (BitDepthY) of the reconstructed luma samples is higher than 10, at least one LMCS codeword may be a multiple of 1<<(BitDepthY-10).

[0603] In one embodiment, the at least one LMCS codeword may be in the range from (OrgCW>>1) to (OrgCW<<1)-1.

[0604] In the above-described embodiment, the method is described based on a flow chart having a series of steps or blocks. This document is not limited to the order of the above-described steps or blocks. Some steps or blocks may occur simultaneously or in a different order than other steps or blocks as described above. In addition, it will be understood by those skilled in the art that the steps shown in the above-described flow chart are not exclusive and may include additional steps, or one or more steps in the flow chart may be deleted without affecting the scope of this document.

[0605] The method according to the above-mentioned embodiment of this document can be implemented in the form of software, and the encoding device and / or decoding device according to this document can be included in a device that performs image processing of a TV, computer, smart phone, set-top box, display device, etc.

[0606] When the embodiments in this document are implemented in software, the above methods can be implemented as modules (processes, functions, etc.) that perform the above functions. The modules can be stored in a memory and executed by a processor. The memory can be located inside or outside the processor and can be connected to the processor by various well-known means. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits and / or data processing devices. The memory may include a read-only memory (ROM), a random access memory (RAM), flash memory, a memory card, a storage medium and / or other storage devices. That is, the embodiments described in this document can be implemented and executed on a processor, a microprocessor, a controller or a chip. For example, the functional units shown in the various figures can be implemented and executed on a computer, a processor, a microprocessor, a controller or a chip. In this case, information about the instructions or algorithms used to implement can be stored in a digital storage medium.

[0607] In addition, the decoding device and the encoding device to which this document is applied can be included in multimedia broadcast transmission / reception devices, mobile communication terminals, home theater video devices, digital theater video devices, surveillance cameras, video chat devices, real-time communication devices (e.g., video communication), mobile streaming devices, storage media, cameras, VoD service providing devices, over-the-top (OTT) video devices, Internet streaming service providing devices, three-dimensional (3D) video devices, teleconferencing video devices, transportation user devices (i.e., vehicle user devices, aircraft user devices, ship user devices, etc.), and medical video devices, and can be used to process video signals and data signals. For example, over-the-top (OTT) video devices may include game consoles, Blu-ray players, Internet access TVs, home theater systems, smart phones, tablet PCs, digital video recorders (DVRs), etc.

[0608] In addition, the processing method to which this document is applied can be generated in the form of a program to be executed by a computer and can be stored in a computer-readable recording medium. Multimedia data with a data structure according to this document can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices that store data readable by a computer system. For example, computer-readable recording media may include BD, universal serial bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage device. In addition, computer-readable recording media include media implemented in the form of carrier waves (i.e., transmission over the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium or can be sent via a wired / wireless communication network.

[0609] In addition, the embodiments of this document can be implemented as a computer program product according to a program code, and the program code can be executed in a computer through the embodiments of this document. The program code can be stored on a computer-readable carrier.

[0610] Figure 21 An example of a content streaming system to which the embodiments disclosed in this document are applicable is shown.

[0611] Reference Figure 21 A content streaming system to which the embodiments of this document are applied may mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.

[0612] The encoding server compresses the content input from the multimedia input device (e.g., a smartphone, a camera, a camcorder, etc.) into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, when the multimedia input device (e.g., a smartphone, a camera, a camcorder, etc.) directly generates the bitstream, the encoding server may be omitted.

[0613] A bitstream may be generated by an encoding method or a bitstream generation method to which an embodiment of this document is applied, and a streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0614] The streaming server transmits multimedia data to a user device via a network server based on a user's request. The network server serves as a medium for informing the user of services. When a user requests a desired service from the network server, the network server transmits the requested service to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server is used to control commands and responses between devices in the content streaming system.

[0615] The streaming server can receive content from a media storage device and / or an encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined time.

[0616] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, ultrabooks, wearable devices (e.g., smart watches, smart glasses, head-mounted displays), digital TVs, desktop computers, digital signage, etc. Each server in the content streaming system may operate as a distributed server, in which case data received from each server may be distributed.

[0617] Each server in the content streaming system may operate as a distributed server, and in this case, data received from each server may be distributed and processed.

[0618] The claims described herein may be combined in various ways. For example, the technical features of the method claims of this document may be combined and implemented as a device, and the technical features of the device claims of this document may be combined and implemented as a method. Furthermore, the technical features of the method claims of this document and the technical features of the device claims of this document may be combined to implement a device, and the technical features of the method claims of this document and the technical features of the device claims of this document may be combined and implemented as a method.

Claims

1. An image decoding method performed by a decoding device, the image decoding method comprising the following steps: Obtain image information from a bitstream, the image information including prediction mode information and luminance mapping and chrominance scaling LMCS related information; deriving an inter-frame prediction mode for a current block in a current picture based on the prediction mode information; generating predicted luma samples based on the inter prediction mode; generating mapped predicted luma samples based on the mapping process on the luma samples; and generating modified reconstructed luma samples based on an inverse mapping process of the luma samples, The image information includes LMCS adaptive parameter set APS, Among them, LMCS APS includes LMCS data field and ID information, Wherein, the LMCS data field includes the LMCS related information, The LMCS codewords used in the mapping process and the reverse mapping process are derived based on the LMCS data field. Wherein, the value of the ID information is within a predetermined range, wherein the predetermined range is in the range of 0 to 3, and One or more slices of the current picture refer to only one LMCS APS.

2. An image encoding method performed by an encoding device, the image encoding method comprising the following steps: Derivation of inter-frame prediction mode; generating predicted luma samples based on the inter prediction mode; Generate luminance mapping and chroma scaling LMCS related information; generating mapped predicted luma samples based on the mapping process of the luma samples; generating modified reconstructed luma samples based on an inverse mapping process of the luma samples; generating residual information based on the mapped predicted luma samples and the modified reconstructed luma samples; as well as Encoding the image information including the LMCS related information and the residual information, The image information includes LMCS adaptive parameter set APS, Among them, LMCS APS includes LMCS data field and ID information, Wherein, the LMCS data field includes the LMCS related information, wherein the LMCS codeword used in the mapping process and the reverse mapping process is derived based on the LMCS data field, and Wherein, the value of the ID information is within a predetermined range, wherein the predetermined range is in the range of 0 to 3, and One of the LMCS APSs is referenced by one or more slices of a picture.

3. A method for transmitting data for image information, the method comprising the following steps: Obtaining a bitstream of the image information including luma mapping and chroma scaling LMCS related information and residual information, generating the LMCS related information by deriving an inter-frame prediction mode, generating predicted luma samples based on the inter-frame prediction mode, and generating luma mapping and chroma scaling LMCS related information, and generating the mapped predicted luma samples by performing a mapping process on the luma samples, generating modified reconstructed luma samples by performing a reverse mapping process on the luma samples, and generating residual information based on the mapped predicted luma samples and the modified reconstructed luma samples to generate the residual information; as well as sending the data of the bitstream including the image information, wherein the image information includes the LMCS related information and the residual information, The image information includes LMCS adaptive parameter set APS, Among them, LMCS APS includes LMCS data field and ID information. Wherein, the LMCS data field includes the LMCS related information, wherein the LMCS codeword used in the mapping process and the reverse mapping process is derived based on the LMCS data field, and Wherein, the value of the ID information is within a predetermined range, wherein the predetermined range is in the range of 0 to 3, and One of the LMCS APSs is referenced by one or more slices of a picture.