Video or image encoding based on luminance mapping

By using Luminance Mapping and Chromatography Scaling (LMCS) processing, the problem of low compression efficiency for high-resolution images/videos is solved, which improves coding efficiency and visual quality, reduces resource consumption, and is suitable for blocks with dual-tree structures.

CN114270851BActive Publication Date: 2026-07-21NOKIA TECHNOLOGIES OY
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NOKIA TECHNOLOGIES OY
Filing Date
2020-06-24
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies are inefficient in high-resolution, high-quality image/video compression, and have high transmission and storage costs, making it difficult to meet the needs of virtual reality and immersive media.

Method used

Employing Luminance Mapping and Chroma Scaling (LMCS) processing, this method utilizes linear mapping and signal-informed monochromatic residual scaling factors, combined with flexible bin usage and index derivation processing. It is suitable for dual-tree structures within coding tree units, limiting the number of LMCS applications and the range of ID values, and simplifying inverse luminance mapping and chroma residual scaling.

Benefits of technology

It improves image/video compression efficiency, enhances subjective/objective visual quality, reduces resources and complexity, is suitable for blocks with dual-tree structures, reduces memory usage, and improves encoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114270851B_ABST
    Figure CN114270851B_ABST
Patent Text Reader

Abstract

According to the disclosure of the present document, the ID value of the LMCS APS can be within a predetermined range. Thus, the LMCS process can be efficiently performed, and the complexity of the LMCS is reduced. Due to the performance improvement of the LMCS, the efficiency of video / image encoding can be increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to video or image coding based on luminance mapping. Background Technology

[0002] Recently, the demand for high-resolution, high-quality images / videos, such as 4K, 8K, or even higher Ultra High Definition (UHD) images / videos, has increased across various fields. As image / video data becomes more high-resolution and higher-quality, the amount of information or bits to be transmitted increases relative to existing image / video data. Therefore, transmitting image data using media such as existing wired / wireless broadband lines or existing storage media, or storing image / video data using existing storage media, increases both transmission and storage costs.

[0003] In addition, interest and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms have recently increased, and the broadcasting of images / videos with characteristics different from real images (e.g., game images) has increased.

[0004] Therefore, highly efficient image / video compression technologies are needed to effectively compress, transmit, store, and reproduce high-resolution, high-quality image / video information with the various characteristics described above.

[0005] In addition, luminance mapping and chroma scaling (LMCS) processing is performed to improve compression efficiency and increase subjective / objective visual quality, and how to apply the LMCS process efficiently is discussed. Summary of the Invention

[0006] Technical solution

[0007] According to the embodiments of this document, a method and apparatus for increasing image coding efficiency are provided.

[0008] According to the implementation method of this document, a high-efficiency filtering application method and device are provided.

[0009] According to the implementation method of this document, an efficient LMCS application method and device are provided.

[0010] According to the implementation method described in this document, LMCS codewords (or their range) can be constrained.

[0011] According to the implementation of this document, a single chroma residual scaling factor can be used that is directly signaled in the chroma scaling of LMCS.

[0012] According to the implementation method described in this document, a linear mapping (linear LMCS) can be used.

[0013] According to the implementation method described in this document, information about the pivot point required for a linear mapping can be explicitly communicated using signals.

[0014] According to the implementation method described in this document, a flexible number of bins can be used for luminance mapping.

[0015] According to the implementation of this document, the index derivation process for inverse luminance mapping and / or chroma residual scaling can be simplified.

[0016] According to the implementation of this document, the LMCS process can be applied even when the luma block and chroma block in a coding tree unit (CTU) have separate block tree structures (dual tree structures).

[0017] According to the implementation method described in this document, the number of LMCS APS can be limited.

[0018] According to the implementation method described in this document, the ID value of LMCS APS can be within a specific range.

[0019] According to the embodiments of this document, a video / image decoding method performed by a decoding device is provided.

[0020] According to embodiments of this document, a decoding device for performing video / image decoding is provided.

[0021] According to the embodiments of this document, a video / image encoding method performed by an encoding device is provided.

[0022] According to the embodiments of this document, an encoding device for performing video / image encoding is provided.

[0023] According to one embodiment of this document, a computer-readable digital storage medium is provided, wherein encoded video / image information generated according to the video / image encoding method disclosed in at least one embodiment of this document is stored.

[0024] According to embodiments of this document, a computer-readable digital storage medium is provided, which stores encoded information or encoded video / image information that causes a decoding device to perform the video / image decoding method disclosed in at least one embodiment of this document.

[0025] Beneficial effects

[0026] According to the implementation method described in this document, the overall image / video compression efficiency can be improved.

[0027] According to the implementation method described in this document, subjective / objective visual quality can be improved through efficient filtering.

[0028] According to the implementation method described in this document, LMCS processing for image / video coding can be performed efficiently.

[0029] The implementation method described in this document can minimize the (software or hardware) resources / costs required for LMCS processing.

[0030] The implementation method described in this document facilitates the hardware implementation of LMCS processing.

[0031] According to the implementation of this document, the division operations required for the derivation of LMCS codewords in the mapping (shaping) can be eliminated or minimized by constraining the LMCS codewords (or their range).

[0032] According to the implementation method described in this document, a single chroma residual scaling factor can be used to remove the delay identified by the segment index.

[0033] According to the implementation of this document, chroma residual scaling can be performed using linear mapping in LMCS without relying on the reconstruction of the luma block, thus eliminating the delay in scaling.

[0034] According to the implementation method described in this document, the mapping efficiency in LMCS can be increased.

[0035] According to the implementation of this document, the complexity of LMCS can be reduced by simplifying the index derivation process used for inverse luminance mapping and / or chroma residual scaling, thereby increasing video / image coding efficiency.

[0036] According to the implementation method described in this document, the LMCS process can be performed even on blocks with a dual-tree structure, thus increasing LMCS efficiency. Furthermore, the coding performance (e.g., objective / subjective picture quality) of blocks with a dual-tree structure can be improved.

[0037] According to the implementation of this document, since the number of LMCS APS is limited, the complexity of LMCS can be reduced, and therefore less resources (e.g., memory) can be consumed (used) in LMCS.

[0038] According to the implementation method described in this document, the LMCS process can be executed efficiently due to the LMCS APS ID values ​​within a specific range, thus reducing the complexity of LMCS. Attached Figure Description

[0039] Figure 1 Examples of video / image coding systems to which embodiments of this document can be applied are shown.

[0040] Figure 2 This is a diagram schematically illustrating the configuration of a video / image encoding device to which the embodiments of this document can be applied.

[0041] Figure 3 This is a diagram schematically illustrating the configuration of a video / image decoding device to which the embodiments of this document can be applied.

[0042] Figure 4 An example of a layered structure for encoded images / videos is shown.

[0043] Figure 5 A hierarchical structure of CVS according to an embodiment of this document is illustrated by way of example.

[0044] Figure 6 A hierarchical structure of CVS according to an embodiment of this document is illustrated by way of example.

[0045] Figure 7 A hierarchical structure of CVS according to another embodiment of this document is illustrated by way of example.

[0046] Figure 8 An exemplary LMCS structure according to an embodiment of this document is shown.

[0047] Figure 9 An LMCS structure according to another embodiment of this document is shown.

[0048] Figure 10 A graph representing an exemplary forward mapping is shown.

[0049] Figure 11 This is a flowchart illustrating a method for deriving a chromaticity residual scaling index according to an embodiment of this document.

[0050] Figure 12 Linear fitting of the pivot point is shown according to the embodiment of this document.

[0051] Figure 13 An example of linear shaping (or linear mapping) according to the implementation of this document is shown.

[0052] Figure 14 An example of a linear forward mapping in the implementation of this document is shown.

[0053] Figure 15 An example of reverse forward mapping in the implementation of this document is shown.

[0054] Figure 16 and Figure 17 Examples of video / image coding methods and related components according to embodiments of this document are illustrated schematically.

[0055] Figure 18 and Figure 19 Examples of image / video decoding methods and related components according to embodiments of this document are illustrated schematically.

[0056] Figure 20 Examples of content streaming systems to which the embodiments disclosed in this document can be applied are shown. Detailed Implementation

[0057] This document may be modified in various forms, and its particular embodiments will be described and illustrated in the accompanying drawings. However, these embodiments are not intended to limit this document. The terminology used in the following description is used only to describe particular embodiments and is not intended to limit this document. Singular expressions include plural expressions, provided that they are clearly not read differently. Terms such as “comprising” and “having” are intended to indicate the presence of the features, quantities, steps, operations, elements, components, or combinations thereof used in the following description, and therefore it should be understood that the possibility of having or adding one or more different features, quantities, steps, operations, elements, components, or combinations thereof is not excluded.

[0058] Furthermore, the various configurations shown in the accompanying drawings for ease of description of different features are shown independently and do not imply that each configuration is implemented as separate hardware or separate software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Implementations in which components are integrated and / or separated are also included within the scope of this disclosure.

[0059] Hereinafter, examples of this embodiment will be described in detail with reference to the accompanying drawings. Furthermore, similar reference numerals are used throughout the drawings to indicate similar elements, and identical descriptions of similar elements will be omitted.

[0060] Figure 1 Examples of video / image coding systems to which embodiments of this document can be applied are shown.

[0061] Reference Figure 1 A video / image encoding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data to the receiving device in the form of a file or stream via a digital storage medium or network.

[0062] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.

[0063] Video sources can acquire video / images through processes that capture, synthesize, or generate video / images. Video sources may include video / image capture devices and / or video / image generation devices. For example, a video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. For example, a video / image generation device may include a computer, tablet computer, and smartphone, and may generate video / images (electronically). For example, virtual video / images may be generated via a computer, etc. In this case, the video / image capture process may be replaced by a process that generates related data.

[0064] Encoding devices can encode input video / images. For compression and encoding efficiency, encoding devices can perform a series of processes such as prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output as a bitstream.

[0065] The transmitter can send encoded image / image information or data, output as a bitstream, to a receiver of a receiving device in the form of a file or stream via a digital storage medium or network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include elements for generating media files according to a predetermined file format and may include elements for transmission over a broadcast / communication network. The receiver can receive / extract the bitstream and send the received bitstream to a decoding device.

[0066] Decoding devices can decode video / images by performing a series of processes such as dequantization, inverse transform, and prediction, which correspond to the operations of encoding devices.

[0067] The renderer can render decoded video / images. The rendered video / images can be displayed on a monitor.

[0068] This document relates to video / image coding. For example, the methods / implementations disclosed in this document can be applied to methods disclosed in the Universal Video Coding (VVC) standard, the Basic Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the Audio Video Coding 2 (AVS2) standard, or next-generation video / image coding standards (e.g., H.267, H.268, etc.).

[0069] This document presents various implementations of video / image coding, and unless otherwise specified, the above implementations may also be combined with each other.

[0070] In this document, video can refer to a series of images over time. A frame typically refers to a unit representing an image within a specific time range, and a slice / tile refers to a unit that constitutes part of a frame in terms of encoding. A slice / tile may include one or more Code Tree Units (CTUs). A frame may consist of one or more slices / tiles. A frame may consist of one or more tile groups. A tile group may include one or more tiles. A tile can represent a rectangular area of ​​a CTU row within a tile in a frame. A tile can be divided into multiple tiles, each tile consisting of one or more CTU rows within the tile. A tile that is not divided into multiple tiles can also be referred to as a tile. A tile scan can represent a specific order of CTUs that divide a frame, wherein CTUs can be ordered in a raster scan of CTUs within a tile, and tiles within a tile can be ordered consecutively in a raster scan of tiles in a tile, and tiles in a frame can be ordered consecutively in a raster scan of tiles in a frame. A tile is a rectangular area of ​​CTUs within a specific tile column and a specific tile row in a frame. A tile column is a rectangular region of CTUs whose height is equal to the height of the frame and whose width is specified by a syntax element in the frame parameter set. A tile row is a rectangular region of CTUs whose height is specified by a syntax element in the frame parameter set and whose width is equal to the width of the frame. A tile scan is a specific ordering of CTUs that divide the frame, wherein CTUs are ordered consecutively in a tile raster scan, and tiles in the frame are ordered consecutively in a frame raster scan. A slice comprises an integer number of tiles in the frame that can be exclusively contained in a single NAL unit. A slice can consist of a consecutive sequence of multiple complete tiles or a single complete tile. In this document, tile groups and slices may be used interchangeably. For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.

[0071] Furthermore, a frame can be divided into two or more sub-frames. A sub-frame can be a rectangular area of ​​one or more slices within the frame.

[0072] A pixel or image unit can refer to the smallest unit that makes up a picture (or image). Additionally, the term "sample" can be used as the counterpart to a pixel. A sample can typically represent a pixel or pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.

[0073] A unit can represent a basic unit of image processing. A unit may include a specific region of an image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "region." In general, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients in M ​​columns and N rows. Alternatively, a sample may refer to a pixel value in the spatial domain, and when such pixel value is transformed to the frequency domain, it may refer to a transform coefficient in the frequency domain.

[0074] In this document, “A or B” may mean “A only”, “B only”, or “both A and B”. In other words, “A or B” in this document can be interpreted as “A and / or B”. For example, in this document, “A, B or C (A, B or C)” means “A only”, “B only”, “C only”, or “any combination of A, B and C”.

[0075] The forward slash ( / ) or comma (、) used in this document can mean "and / or". For example, "A / B" can mean "A and / or B". Therefore, "A / B" can mean "A only", "B only", or "both A and B". For example, "A, B, C" can mean "A, B or C".

[0076] In this document, "at least one of A and B" may mean "only A", "only B" or "both A and B". Furthermore, in this document, the expressions "at least one of A or B" or "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".

[0077] Additionally, in this document, "at least one of A, B, and C" means "A only", "B only", "C only" or "any combination of A, B, and C". Furthermore, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".

[0078] Additionally, the parentheses used in this document may mean "for example". Specifically, when indicating "prediction (intra-frame prediction)", "intra-frame prediction" may be suggested as an example of "prediction". In other words, "prediction" in this document is not limited to "intra-frame prediction", and "intra-frame prediction" may be suggested as an example of "prediction". Furthermore, even when indicating "prediction (i.e., intra-frame prediction)", "intra-frame prediction" may be suggested as an example of "prediction".

[0079] The technical features described individually in a single figure in this document can be implemented individually or simultaneously.

[0080] Figure 2This diagram schematically illustrates the configuration of a video / image encoding apparatus to which embodiments of this document can be applied. Hereinafter, the term "video encoding apparatus" may include an image encoding apparatus.

[0081] Reference Figure 2 The encoding device 200 includes an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transform 232, a quantizer 233, a dequantizer 234, and an inverse transform 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image segmenter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). Additionally, the memory 270 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware component may also include the memory 270 as an internal / external component.

[0082] Image segmenter 210 can segment an input image (or picture or frame) input to encoding device 200 into one or more processors. For example, a processor may be referred to as a coding unit (CU). In this case, the coding unit may be recursively segmented from coding tree unit (CTU) or maximum coding unit (LCU) according to a quadtree-binary-trinary tree (QTBTTT) structure. For example, a coding unit may be segmented into multiple deeper coding units based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, a quadtree structure may be applied first, followed by a binary tree structure and / or a ternary structure. Alternatively, a binary tree structure may be applied first. The encoding process according to this document may be performed based on the final coding unit that is no longer segmented. In this case, based on the image characteristics and encoding efficiency, the maximum coding unit may be used as the final coding unit, or if necessary, the coding unit may be recursively segmented into deeper coding units, and the coding unit with the optimal size may be used as the final coding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes (described later). As another example, the processor may also include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be split or separated from the final encoding unit described above. The prediction unit may be a unit for predicting samples, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving residual signals from transform coefficients.

[0083] In some cases, a unit can be used interchangeably with terms such as a block or region. Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. Samples can typically represent pixels or pixel values, and may represent pixel / pixel values ​​of only the luminance component or only the chrominance component. A sample can be used as a term corresponding to a frame (or image) of pixels or picometers.

[0084] In the encoding device 200, a residual signal (residual block, residual sample array) is generated by subtracting the prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 from the input image signal (original block, original sample array), and the generated residual signal is sent to the converter 232. In this case, as shown, the unit in the encoding device 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) may be called the subtractor 231. The predictor can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block including the prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. As described later in the description of the various prediction modes, the predictor can generate various information related to the prediction (e.g., prediction mode information) and send the generated information to the entropy encoder 240. The information about the prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.

[0085] Intra-predictor 222 can refer to samples in the current frame to predict the current block. Depending on the prediction mode, the referenced samples may be located near or separated from the current block. In intra-prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. For example, non-directional modes may include DC mode and planar mode. For example, depending on the level of detail in the prediction direction, the directional modes may include 33 or 65 directional prediction modes. However, this is just an example, and more or fewer directional prediction modes may be used depending on the settings. Intra-predictor 222 can use the prediction modes applied to neighboring blocks to determine the prediction mode applied to the current block.

[0086] Inter-frame predictor 221 can deduce the prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference frame. Here, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. The reference frame including the reference block and the reference frame including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a juxtaposed reference block, a juxtaposed CU (colCU), etc., and the reference frame including the temporally neighboring block may be referred to as a juxtaposed frame (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to deduce the motion vector and / or reference frame index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 221 can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, motion vectors from neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0087] Predictor 220 can generate a prediction signal based on various prediction methods described below. For example, the predictor can apply not only intra-frame prediction or inter-frame prediction to predict a block, but also both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Alternatively, the predictor can predict blocks based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as screen content coding (SCC) in games, etc. IBC essentially performs prediction in the current frame, but can be performed similarly to inter-frame prediction, such that a reference block is derived in the current frame. That is, IBC can use at least one inter-frame prediction technique described in this document. The palette mode can be considered as an example of intra-frame coding or intra-frame prediction. When the palette mode is applied, the sample values ​​in the frame can be signaled based on information about the palette table and palette index.

[0088] The predicted signal generated by the predictor (including inter-frame predictor 221 and / or intra-frame predictor 222) can be used to generate a reconstructed signal or a residual signal. Transformer 232 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform technique may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graphical Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graphic when the relationship information between pixels is represented graphically. CNT refers to a transform generated based on the predicted signal generated using all previously reconstructed pixels. Furthermore, the transform processing can be applied to square pixel blocks of the same size or to blocks of variable size other than squares.

[0089] Quantizer 233 quantizes the transform coefficients and sends them to entropy encoder 240, which encodes the quantized signal (information about the quantized transform coefficients) and outputs a bitstream. This information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the block-type quantized transform coefficients into a one-dimensional vector based on the coefficient scan order, and generate information about the quantized transform coefficients based on this one-dimensional vector. Entropy encoder 240 can perform various encoding methods such as exponential Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). Entropy encoder 240 can encode information required for video / image reconstruction other than the quantized transform coefficients (e.g., values ​​of syntax elements, etc.) together or separately. The encoded information (e.g., encoded video / image information) can be sent or stored in NAL (Network Abstraction Layer) units as a bitstream. The video / image information may also include information about various parameter sets, such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Additionally, the video / image information may include general constraint information. In this document, information and / or syntactic elements transmitted from the encoding device / notified by signal to the decoding device may be included in the video / picture information. The video / image information may be encoded by the above-described encoding process and included in the bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. The network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits the signal output from the entropy encoder 240 and / or a storage unit (not shown) that stores the signal may be included as an internal / external element of the encoding device 200, and alternatively, the transmitter may be included in the entropy encoder 240.

[0090] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients via dequantizer 234 and inverse transformer 235. Adder 250 adds the reconstructed residual signal to the prediction signal output from inter-frame predictor 221 or intra-frame predictor 222 to generate a reconstructed signal (reconstructed frame, reconstructed block, reconstructed sample array). If the block to be processed has no residual (e.g., in the case of applying skip mode), the prediction block can be used as a reconstructed block. Adder 250 may be referred to as a reconstructor or reconstructed block generator. As described below, the generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current frame and can be filtered for inter-frame prediction of the next frame.

[0091] In addition, luminance mapping with chroma scaling (LMCS) can be applied during screen encoding and / or reconstruction.

[0092] Filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 270 (specifically, the DPB of memory 270). Various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. Filter 260 can generate various filtering-related information and send the generated information to entropy encoder 240, as described later in the description of each filtering method. The filtering-related information can be encoded by entropy encoder 240 and output as a bitstream.

[0093] The modified reconstructed frame sent to memory 270 can be used as a reference frame in inter-frame predictor 221. When inter-frame prediction is applied by the encoding device, prediction mismatch between encoding device 200 and decoding device 300 can be avoided and encoding efficiency can be improved.

[0094] The DPB of memory 270 can store reconstructed frames modified for use as reference frames in inter-frame predictor 221. Memory 270 can store motion information of blocks in the current frame that derive (or encode) motion information and / or motion information of already reconstructed blocks in the frame. The stored motion information can be sent to inter-frame predictor 221 and used as motion information for spatially or temporally neighboring blocks. Memory 270 can store reconstructed samples of reconstructed blocks in the current frame and can transmit the reconstructed samples to intra-frame predictor 222.

[0095] Figure 3 This is a schematic diagram illustrating the configuration of a video / image decoding device to which the embodiments of this document can be applied.

[0096] Reference Figure 3 The decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an intra-frame predictor 331 and an inter-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to embodiments, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 may be configured by hardware components (e.g., a decoder chipset or processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.

[0097] When the input includes a bitstream containing video / image information, the decoding device 300 can reconstruct and... Figure 2 The encoding device processes video / image information corresponding to the image. For example, the decoding device 300 can deduce units / blocks based on block segmentation information obtained from the bitstream. The decoding device 300 can use a processor applied in the encoding device to perform decoding. Therefore, for example, the processor for decoding can be an encoding unit, and the encoding unit can be segmented from encoding tree units or maximum encoding units according to a quadtree structure, binary tree structure, and / or ternary tree structure. One or more transform units can be derived from the encoding units. The reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproduction device.

[0098] Decoding device 300 can receive from Figure 2The encoding device outputs a signal in the form of a bitstream, and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets, such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). In addition, the video / image information may also include general constraint information. The decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The information and / or syntax elements notified / received by signals, as described later in this document, can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 310 decodes the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, and outputs the quantized values ​​of the syntax elements and transform coefficients of the residuals required for image reconstruction. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntactic element in the bitstream, determines a context model using information about the target syntactic element, decoding information about the target block, or information about symbols / bins decoded in a previous stage, and performs arithmetic decoding on the bins by predicting the probability of bin occurrence based on the determined context model, generating symbols corresponding to the values ​​of each syntactic element. In this case, the CABAC entropy decoding method can update the context model after determining the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin. Information related to prediction from the information decoded by the entropy decoder 310 can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values ​​(i.e., quantized transform coefficients and related parameter information) from the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). Additionally, information about filtering from the information decoded by the entropy decoder 310 can be provided to the filter 350. Furthermore, a receiver (not shown) for receiving signals output from the encoding device may be configured as an internal / external element of the decoding device 300, or the receiver may be a component of the entropy decoder 310. Additionally, the decoding device according to this document may be referred to as a video / image / picture decoding device, and the decoding device may be classified as an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.

[0099] Dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. Dequantizer 321 can rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device. Dequantizer 321 can use quantization parameters (e.g., quantization step size information) to perform dequantization on the quantized transform coefficients and obtain the transform coefficients.

[0100] The inverse transformer 322 performs inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).

[0101] The predictor can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on information about the prediction output from the entropy decoder 310, and can determine a specific intra-frame / inter-frame prediction mode.

[0102] Predictor 330 can generate prediction signals based on various prediction methods. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Alternatively, the predictor can predict blocks based on an intra-block copy (IBC) prediction mode or a palette mode. IBC prediction mode or palette mode can be used for content image / video coding in games, such as screen content coding (SCC). IBC essentially performs prediction in the current frame, but can be performed similarly to inter-frame prediction, such that a reference block is derived in the current frame. That is, IBC can use at least one inter-frame prediction technique described in this document. Palette mode can be considered as an example of intra-frame coding or intra-frame prediction. When applying palette mode, sample values ​​within the frame can be signaled based on information about the palette table and palette index.

[0103] Intra-predictor 331 can predict the current block by referencing samples in the current frame. Depending on the prediction mode, the referenced samples may be located near or separated from the current block. In intra-prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. Intra-predictor 331 can use the prediction modes applied to neighboring blocks to determine the prediction mode applied to the current block.

[0104] Inter-frame predictor 332 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference frame. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. For example, inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference frame index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the inter-frame prediction mode of the current block.

[0105] Adder 340 generates a reconstruction signal (reconstructed frame, reconstruction block, reconstruction sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including inter-frame predictor 332 and / or intra-frame predictor 331). If the block to be processed has no residual, such as when a skip mode is applied, the prediction block can be used as a reconstruction block.

[0106] Adder 340 may be referred to as a reconstructor or reconstruction block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current frame, and can be output through filtering as described below, or it can be used for inter-frame prediction of the next frame.

[0107] In addition, Luminance Mapping with Chroma Scaling (LMCS) can be applied in the image decoding process.

[0108] Filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 360 (specifically, the DPB of memory 360). For example, various filtering methods may include deblocking filtering, sample adaptive shifting, adaptive loop filtering, bilateral filtering, etc.

[0109] The (modified) reconstructed frame stored in the DPB of memory 360 can be used as a reference frame in inter-frame predictor 332. Memory 360 can store motion information of blocks in the current frame from which motion information is derived (or decoded) and / or motion information of already reconstructed blocks in the frame. The stored motion information can be sent to inter-frame predictor 332 as motion information of spatially or temporally neighboring blocks. Memory 360 can store reconstructed samples of reconstructed blocks in the current frame and transmit the reconstructed samples to intra-frame predictor 331.

[0110] In this document, the implementations described in the filter 260, inter-frame predictor 221, and intra-frame predictor 222 of the encoding device 200 can be applied in the same way as or respectively corresponding to the filter 350, inter-frame predictor 332, and intra-frame predictor 331 of the decoding device 300. This can also be applied to the inter-frame predictor 332 and the intra-frame predictor 331.

[0111] As described above, in video coding, prediction is performed to increase compression efficiency. This generates a prediction block that includes prediction samples of the current block (the block to be encoded). Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived identically from both the encoding and decoding devices, and the encoding device decodes information about the residual between the original block and the prediction block (residual information) rather than the original sample values ​​of the original block themselves. By signaling the device, image coding efficiency can be increased. The decoding device can derive a residual block including residual samples based on the residual information, and generate a reconstruction block including reconstructed samples by summing the residual block and the prediction block, and generate a reconstructed image including the reconstruction block.

[0112] Residual information can be generated through transform and quantization processes. For example, the encoding device can derive a residual block between the original block and the prediction block, and perform transform processing on the residual samples (residual sample array) included in the residual block to derive transform coefficients. Then, by performing quantization processing on the transform coefficients, quantized transform coefficients are derived to signal residual-related information (via bitstream) to the decoding device. Here, residual information may include position information, transform technique, transform kernel and quantization parameters, and the value information of the quantized transform coefficients. The decoding device can perform dequantization / inverse transform processing based on the residual information and derive residual samples (or residual blocks). The decoding device can generate a reconstructed frame based on the prediction block and the residual block. The encoding device can also perform dequantization / inverse transform on the quantized transform coefficients used for inter-frame prediction reference in later frames to derive residual blocks and generate a reconstructed frame based on them.

[0113] In this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the transform coefficients of the quantization may be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or, for consistency of expression, may still be referred to as transform coefficients.

[0114] In this document, quantized transform coefficients and transform coefficients can be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficients, and this information about the transform coefficients can be signaled via residual coding syntax. Transform coefficients can be derived based on residual information (or information about the transform coefficients), and scaled transform coefficients can be derived by the inverse transform (scaling) of the transform coefficients. Residual samples can be derived based on the inverse transform (scaling) of the scaled transform coefficients. This can also be applied / expressed in other parts of this document.

[0115] Intra-frame prediction refers to the prediction of a block based on reference samples in the frame to which the current block belongs (hereinafter referred to as the current frame). When intra-frame prediction is applied to the current block, neighboring reference samples to be used for intra-frame prediction of the current block can be derived. The neighboring reference samples of the current block may include samples adjacent to the left boundary of the current block of size nW×nH and a total of 2×nH samples adjacent to the lower left, samples adjacent to the top boundary of the current block and a total of 2×nW samples adjacent to the upper right, and a sample adjacent to the upper left of the current block. Alternatively, the neighboring reference samples of the current block may include multiple upper neighbor samples and multiple left neighbor samples. In addition, the neighboring reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block of size nW×nH, a total of nW samples adjacent to the bottom boundary of the current block, and a sample adjacent to the lower right of the current block.

[0116] However, some neighboring reference samples of the current block may not yet be decoded or available. In this case, the decoder can configure the neighboring reference samples to be used for prediction by replacing the unavailable samples with the available ones. Alternatively, the neighboring reference samples to be used for prediction can be configured by interpolation of the available samples.

[0117] When deriving neighboring reference samples, (i) the predicted sample can be derived based on the average or interpolation of the neighboring reference samples of the current block, and (ii) the predicted sample can be derived based on reference samples among the peripheral reference samples of the current block that exist in a specific (predictive) direction of the predicted sample. Case (i) can be referred to as non-directional mode or non-angular mode, and case (ii) can be referred to as directional mode or angular mode.

[0118] Furthermore, predicted samples can also be generated by interpolating between the first and second nearest neighbor samples in the neighboring reference samples, where the predicted sample of the current block is located in the direction opposite to the prediction direction of the intra-prediction mode of the current block. This is known as Linear Interpolation Intra-Prediction (LIP). Alternatively, a linear model can be used to generate chroma predicted samples based on luminance samples. This is known as LM mode.

[0119] Additionally, a provisional prediction sample for the current block can be derived based on filtered neighboring reference samples, and at least one reference sample derived from the existing neighboring reference samples according to the intra-prediction mode (i.e., an unfiltered neighboring reference sample) and the provisional prediction sample can be weighted and summed to derive the prediction sample for the current block. This situation can be referred to as position-dependent intra-prediction (PDPC).

[0120] Alternatively, the reference sample line with the highest prediction accuracy among the neighboring multiple reference sample lines of the current block can be selected to derive the prediction sample using reference samples located in the prediction direction on the corresponding line. The reference sample line used in this paper can then be indicated (signaled) to the decoding device to perform intra-frame prediction coding. The above situation can be referred to as multiple reference line (MRL) intra-frame prediction or MRL-based intra-frame prediction.

[0121] Alternatively, intra-prediction can be performed by dividing the current block into vertical or horizontal sub-partitions based on the same intra-prediction mode, and neighboring reference samples can be derived and used on a sub-partition basis. That is, in this case, the intra-prediction mode of the current block also applies to sub-partitions, and in some cases, intra-prediction performance can be improved by deriving and using neighboring reference samples on a sub-partition basis. This prediction method can be called intra-partition (ISP) or ISP-based intra-prediction.

[0122] The intra-prediction methods described above can be separated from intra-prediction modes and referred to as intra-prediction types. Intra-prediction types can be named using various terms such as intra-prediction techniques or additional intra-prediction modes. For example, an intra-prediction type (or additional intra-prediction mode) may include at least one of LIP, PDPC, MRL, and ISP described above. General intra-prediction methods other than specific intra-prediction types such as LIP, PDPC, MRL, or ISP may be referred to as normal intra-prediction types. Normal intra-prediction types can typically be applied without applying specific intra-prediction types and can be used to perform predictions based on the intra-prediction modes described above. Furthermore, post-filtering can be performed on the derived prediction samples as needed.

[0123] Specifically, the intra-frame prediction process may include an intra-frame prediction mode / type determination step, a neighboring reference sample derivation step, and a prediction sample derivation step based on the intra-frame prediction mode / type. Additionally, a post-filtering step may be performed on the derived prediction samples as needed.

[0124] When applying intra-prediction, the intra-prediction mode applied to the current block can be determined using the intra-prediction modes of neighboring blocks. For example, the decoding device can select one of the MPM candidates from a list of MPM candidates derived from the intra-prediction modes of neighboring blocks (e.g., left and / or top neighboring blocks) based on the received Most Probable Mode (MPM) index, and select one of the remaining intra-prediction modes not included in the MPM candidates (and planar modes) based on the remaining intra-prediction mode information. The MPM list can be configured to include or exclude planar modes as candidates. For example, if the MPM list includes planar modes as candidates, it can have six candidates. If the MPM list excludes planar modes as candidates, it can have three candidates. When the MPM list excludes planar modes as candidates, a non-planar flag indicating whether the intra-prediction mode of the current block is not a planar mode can be signaled (e.g., intra_luma_not_planar_flag). For example, the MPM flag can be signaled first, and when the MPM flag is 1, the MPM index and non-planar flag can be signaled. Additionally, when the non-planar flag has a value of 1, the MPM index can be signaled. Here, the MPM list is configured not to include planar patterns as candidates, and the non-planar flag is not signaled first to check if it is a planar pattern, because planar patterns are always considered as MPMs.

[0125] For example, whether the intra-prediction mode applied to the current block is in the MPM candidate (and planar mode) or in the remaining mode can be indicated based on the MPM flag (e.g., Intra_luma_mpm_flag). A value of 1 for the MPM flag indicates that the intra-prediction mode of the current block is in the MPM candidate (and planar mode), and a value of 0 for the MPM flag indicates that the intra-prediction mode of the current block is not in the MPM candidate (and planar mode). A value of 0 for the non-planar flag (e.g., Intra_luma_not_planar_flag) indicates that the intra-prediction mode of the current block is planar, and a value of 1 for the non-planar flag indicates that the intra-prediction mode of the current block is not planar. The MPM index can be signaled in the form of the MPM_idx or Intra_luma_mpm_idx syntax element, and the remaining intra-prediction mode information can be signaled in the form of the Rem_intra_luma_pred_mode or Intra_luma_mpm_remainder syntax element. For example, the remaining intra-prediction mode information can be indexed according to the prediction mode number to indicate one of the remaining intra-prediction modes not included in the MPM candidates (and planar modes). The intra-prediction mode can be an intra-prediction mode for the luma component (sample). Hereinafter, the intra-prediction mode information may include at least one of the following: an MPM flag (e.g., Intra_luma_mpm_flag), a non-planar flag (e.g., Intra_luma_not_planar_flag), an MPM index (e.g., mpm_idx or intra_luma_mpm_idx), and remaining intra-prediction mode information (rem_intra_luma_pred_mode or intra_luma_mpm_remainder). In this document, the MPM list may be referred to by various terms such as the MPM candidate list and candModeList. When MIP is applied to the current block, the individual MPM flags used for MIP (e.g., intra_mip_mpm_flag), MPM index (e.g., intra_mip_mpm_idx), and remaining intra-prediction mode information (e.g., intra_mip_mpm_remainder) can be signaled, and non-plane flags are not signaled.

[0126] In other words, when performing block splitting on an image, the current block to be encoded and its neighboring blocks typically have similar image characteristics. Therefore, there is a high probability that the current block and its neighboring blocks have the same or similar intra-prediction modes. Thus, the encoder can use the intra-prediction modes of neighboring blocks to encode the intra-prediction mode of the current block.

[0127] For example, the encoder / decoder can configure a list of most probable modes (MPMs) for the current block. The MPM list can also be referred to as the MPM candidate list. In this paper, MPM can refer to a mode that considers the similarity between the current block and neighboring blocks to improve coding efficiency in intra-frame predictive mode coding. As mentioned above, the MPM list can be configured to include planar modes or can be configured not to include planar modes. For example, when the MPM list includes planar modes, the number of candidates in the MPM list can be 6. And if the MPM list does not include planar modes, the number of candidates in the MPM list can be 5.

[0128] The encoder / decoder can be configured to include a list of 5 or 6 MPMs.

[0129] To configure the MPM list, three types of modes can be considered: default intra-frame mode, neighbor intra-frame mode, and deduced intra-frame mode.

[0130] For the near-neighbor intra-frame mode, two neighboring blocks can be considered: the left neighboring block and the top neighboring block.

[0131] As described above, if the MPM list is configured not to include flat patterns, then flat patterns are excluded from the list, and the number of candidates in the MPM list can be set to 5.

[0132] In addition, non-directional (or non-corner) modes in intra-frame prediction modes may include DC modes based on the average of neighboring reference samples of the current block or planar modes based on interpolation.

[0133] When inter-frame prediction is applied, the predictor of the encoding / decoding device can derive prediction samples by performing inter-frame prediction on a block-by-block basis. Inter-frame prediction can be derived in a manner that depends on data elements (e.g., sample values ​​or motion information) of frames other than the current frame. When applying inter-frame prediction to the current block, the prediction block (prediction sample array) of the current block can be derived based on the reference block (reference sample array) specified by the motion vector on the reference frame indicated by the reference frame index. Here, to reduce the amount of motion information transmitted in inter-frame prediction mode, the motion information of the current block can be predicted on a block, sub-block, or sample-by-sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. The reference frame including the reference block and the reference frame including the temporally neighboring block may be the same or different. Temporally neighboring blocks can be referred to as collated reference blocks, collated CUs (colCU), etc., and reference frames including temporally neighboring blocks can be referred to as collated frames (colPic). For example, a motion information candidate list can be configured based on the neighboring blocks of the current block, and a signal can be used to indicate which candidate's flag or index information is selected (used) to derive the motion vector of the current block and / or the reference frame index. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the motion information of the current block can be the same as the motion information of neighboring blocks. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, the motion vectors of the selected neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.

[0134] Depending on the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.), motion information may include L0 motion information and / or L1 motion information. Motion vectors in the L0 direction may be referred to as L0 motion vectors or MVL0, and motion vectors in the L1 direction may be referred to as L1 motion vectors or MVL1. Prediction based on L0 motion vectors may be called L0 prediction, prediction based on L1 motion vectors may be called L1 prediction, and prediction based on both L0 and L1 motion vectors may be called bi-prediction. Here, L0 motion vectors may indicate motion vectors associated with a reference frame list L0 (L0), and L1 motion vectors may indicate motion vectors associated with a reference frame list L1 (L1). The reference frame list L0 may include frames that are earlier than the current frame in output order as reference frames, and the reference frame list L1 may include frames that are later than the current frame in output order. Previous frames may be referred to as forward (reference) frames, and subsequent frames may be referred to as backward (reference) frames. The reference frame list L0 may also include frames that are later than the current frame in output order as reference frames. In this scenario, the previous frame can be indexed first in the reference frame list L0, and subsequent frames can be indexed later. The reference frame list L1 may also include previous frames that are earlier than the current frame in the output order. In this case, subsequent frames can be indexed first in the reference frame list L1, and previous frames can be indexed later. The output order can correspond to the Frame Order Count (POC) order.

[0135] Figure 4 An example is shown of the layered structure of an encoded image / video.

[0136] Reference Figure 4 The encoded images / videos are divided into a Video Coding Layer (VCL) that processes the images / videos and performs its own decoding, a subsystem that sends and stores the encoded information, and a Network Abstraction Layer (NAL) that is responsible for the functions and exists between the VCL and the subsystems.

[0137] In VCL, VCL data including compressed image data (slice data) is generated, or parameter sets including Picture Parameter Set (PSP), Sequence Parameter Set (SPS), and Video Parameter Set (VPS) or additional supplemental enhancement information (SEI) messages required for image decoding processing can be generated.

[0138] In NAL, NAL cells are generated by adding header information (NAL cell header) to the raw byte sequence payload (RBSP) generated in VCL. In this case, RBSP refers to slice data, parameter sets, SEI messages, etc., generated in VCL. The NAL cell header may include NAL cell type information specified based on the RBSP data included in the corresponding NAL cell.

[0139] As shown in the figure, NAL units can be classified into VCL NAL units and non-VCL NAL units based on the RBSP generated in the VCL. A VCL NAL unit can refer to a NAL unit that includes information about the image (slice data), while a non-VCL NAL unit can refer to a NAL unit that includes information (parameter set or SEI message) required to decode the image.

[0140] The aforementioned VCL NAL units and non-VCL NAL units can be transmitted over a network by adding header information according to the data standard of the subsystem. For example, NAL units can be converted into predetermined standard data formats such as H.266 / VVC file format, Real-time Transport Protocol (RTP), Transport Stream (TS), etc., and transmitted over various networks.

[0141] As described above, NAL cells can be specified as NAL cell types based on the RBSP data structures included in the corresponding NAL cells, and information about the NAL cell type can be stored in the NAL cell header and notified by signals.

[0142] For example, NAL units can be classified into VCL NAL unit types and non-VCL NAL unit types based on whether the NAL unit includes information about the image (slice data). VCL NAL unit types can be classified based on the nature and type of the image included in the VCL NAL unit, and non-VCL NAL unit types can be classified based on the type of parameter set.

[0143] The following is an example of a NAL cell type specified based on the type of the parameter set included in a non-VCL NAL cell type.

[0144] -APS (Adaptive Parameter Set) NAL Unit: The type of NAL unit including APS.

[0145] -DPS (Decoding Parameter Set) NAL Unit: The type of NAL unit including DPS.

[0146] -VPS (Video Parameter Set) NAL Unit: Includes the type of NAL unit for the VPS.

[0147] -SPS (Sequence Parameter Set) NAL Unit: The type of NAL unit that includes SPS.

[0148] -PPS (Picture Parameter Set) NAL Unit: Includes the types of NAL units for PPS.

[0149] -PH (Header) NAL Unit: Includes the type of NAL unit for PH.

[0150] The aforementioned NAL unit type may contain syntactic information, which can be stored in the NAL unit header and signaled. For example, the syntactic information may be nal_unit_type, and the NAL unit type may be specified by the nal_unit_type value.

[0151] Furthermore, as mentioned above, a frame can include multiple slices, and a slice can include a slice header and slice data. In this case, a frame header can be further added to multiple slices (slice header and slice dataset) within a frame. The frame header (frame header syntax) can include information / parameters typically applicable to the frame. In this document, slices can be mixed with or replaced by tile groups. Additionally, in this document, slice headers can be mixed with or replaced by tile group headers.

[0152] A slice header (slice header syntax) may include information / parameters typically applicable to a slice. APS (APS syntax) or PPS (PPS syntax) may include information / parameters typically applicable to one or more slices or frames. SPS (SPS syntax) may include information / parameters typically applicable to one or more sequences. VPS (VPS syntax) may include information / parameters typically applicable to multiple layers. DPS (DPS syntax) may include information / parameters typically applicable to the total video. DPS may include information / parameters related to the concatenation of encoded video sequences (CVS). The High-Level Syntax (HLS) in this document may include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, and slice header syntax.

[0153] In this document, the image / image information encoded by the encoding device and signaled to the decoding device in the form of a bitstream includes not only segmentation-related information, intra / inter-frame prediction information, residual information, loop filtering information, etc. in the frame, but also information included in the slice header, information included in the APS, information included in the PPS, information included in the SPS and / or information included in the VPS.

[0154] Furthermore, to compensate for differences between the original and reconstructed images caused by errors in compression coding processes such as quantization, loop filtering can be performed on the reconstructed samples or reconstructed images as described above. As mentioned above, loop filtering can be performed by filters from the encoding and decoding devices, and deblocking filters, SAO, and / or adaptive loop filters (ALF) can be applied. For example, ALF processing can be performed after deblocking filtering and / or SAO processing is completed. However, even in this case, deblocking filtering and / or SAO processing can be omitted.

[0155] Furthermore, to increase coding efficiency, Luminance Mapping and Chroma Scaling (LMCS) can be applied as described above. LMCS can be referred to as a loop shaper (shaper). To increase coding efficiency, LMCS control and / or LMCS-related signaling can be performed in layers.

[0156] Figure 5 A hierarchical structure of CVS according to an embodiment of this document is illustrated by way of example.

[0157] Reference Figure 5 A encoded video sequence (CVS) may include a sequence parameter set (SPS), one or more frame parameter sets (PPS), and one or more subsequent encoded frames. Each encoded frame can be divided into rectangular regions. These rectangular regions may be called tiles. One or more tiles can be collected to form a tile group or slice. In this case, the tile group header may be linked to the frame parameter set (PPS), and the PPS may be linked to the SPS.

[0158] Figure 6 A hierarchical structure of a CVS according to an embodiment of this document is illustrated exemplarily. A coded video sequence (CVS) may include a SPS, PPS, a tile group header, tile data, and / or a CTU. Here, the tile group header and tile data may be referred to as a slice header and slice data, respectively.

[0159] The SPS may include local flags to enable the tool to be used in CVS. Additionally, the SPS may be referenced by the PPS, which includes information about parameters that change for each frame. Each encoded frame may include one or more encoded rectangular field tiles. Tiles may be grouped into raster scans forming tile groups. Each tile group is encapsulated with header information called a tile group header. Each tile consists of a CTU including encoded data. Here, the data may include raw sample values, predicted sample values, and their luminance and chrominance components (luminance predicted sample values ​​and chrominance predicted sample values).

[0160] According to existing methods, ALF data (ALF parameters) or LMCS data (LMCS parameters) are included in the tile group header. Considering that a video consists of multiple frames and a frame includes multiple tiles, frequently signaling ALF data (ALF parameters) or LMCS data (LMCS parameters) at the tile group level leads to a degradation in coding efficiency.

[0161] According to the implementation method proposed in this document, ALF parameters or LMCS data (LMCS parameters) can be included in the APS to be signaled as follows.

[0162] Figure 7 A hierarchical structure of CVS according to another embodiment of this document is illustrated by way of example.

[0163] Reference Figure 7 An APS can be defined, and the APS can carry the necessary ALF data (ALF parameters). Additionally, the APS can have self-identification parameters, ALF data, and / or LMCS data. The self-identification parameters of the APS can include an APS ID. That is, the APS can include information representing the APS ID. The tile group header or slice header can use the APS index information to reference the APS. In other words, the tile group header or slice header can include APS index information, and the ALF procedure for the target block can be performed based on the LMCS data (LMCS parameters) included in the APS with the APS ID indicated by the APS index information, or the ALF procedure for the target block can be performed based on the ALF data (ALF parameters) included in the APS with the APS ID indicated by the APS index information. Here, the APS index information can be referred to as APS ID information.

[0164] In the example, the SPS may include flags that allow the use of ALF. For example, the SPS can be checked when CVS starts, and the flags in the SPS can be checked. For example, the SPS may include the syntax of Table 1 below. The syntax of Table 1 may be part of the SPS.

[0165] [Table 1]

[0166]

[0167] For example, the semantics of syntactic elements included in the syntax of Table 1 above can be represented in the following table.

[0168] [Table 2]

[0169]

[0170] That is, the `sps_alf_enabled_flag` syntax element can indicate whether ALF is available based on whether its value is 0 or 1. The `sps_alf_enabled_flag` syntax element can be referred to as the ALF availability flag (first ALF availability flag) and can be included in the SPS. That is, the ALF availability flag can be signaled in the SPS (or at the SPS level). If the value of the ALF availability flag signaled in the SPS is 1, then ALF can be essentially determined to be available relative to the reference SPS in the CVS. Furthermore, as mentioned above, ALF can be individually handled as on / off by signaling an additional availability flag at a level lower than the SPS.

[0171] For example, if the ALF tool is available for CVS, an additional availability flag (which may be referred to as the second ALF availability flag) can be signaled in the tile group header or slice header. For example, if ALF is available at the SPS level, the second ALF availability flag can be parsed / signed. If the second ALF availability flag is 1, the ALF data can be parsed through the tile group header or slice header. For example, the second ALF availability flag can specify the ALF availability conditions for the luma and chroma components. ALF data can be accessed via APSID information.

[0172] [Table 3]

[0173]

[0174] [Table 4]

[0175]

[0176] For example, the semantics of the syntactic elements of the syntax in Table 3 or Table 4 above can be represented in the following table.

[0177] [Table 5]

[0178]

[0179] [Table 6]

[0180]

[0181]

[0182] The second ALF enabled flag may include either the tile_group_alf_enabled_flag syntax element or the slice_alf_enabled_flag syntax element.

[0183] Based on APS ID information (e.g., the tile_group_aps_id syntax element or the slice_aps_id syntax element), the APS referenced by the corresponding tile group or slice can be identified. APS may include ALF data.

[0184] Furthermore, for example, the structure of an APS that includes ALF data can be described based on the following syntax and semantics. The syntax in Table 7 can be part of the APS.

[0185] [Table 7]

[0186]

[0187] [Table 8]

[0188]

[0189] As described above, the `adaptation_parameter_set_id` syntax element can represent the identifier of the corresponding APS. That is, the APS can be identified based on the `adaptation_parameter_set_id` syntax element. The `adaptation_parameter_set_id` syntax element can be referred to as the APS ID information. Additionally, the APS may include an ALF data field. The ALF data field can be parsed / signaled after the `adaptation_parameter_set_id` syntax element.

[0190] Additionally, for example, APS extension flags (e.g., the aps_extension_flag syntax element) can be resolved / signed in the APS. APS extension flags can indicate the presence of the APS extension data flag (aps_extension_data_flag) syntax element. For example, APS extension flags can be used to provide extension points for later versions of the VVC standard.

[0191] Figure 8 An exemplary LMCS structure according to an embodiment of this document is shown. Figure 8 The LMCS structure 800 includes a loop mapping unit 810 for the luminance component based on an adaptive piecewise linear (adaptive PWL) model and a luminance-dependent chrominance residual scaling unit 820 for the chrominance component. The dequantization and inverse transform 811, reconstruction 812, and intra-frame prediction blocks of the loop mapping unit 810 represent the processing applied in the mapping (shaping) domain. The loop filter 815, the motion compensation or inter-frame prediction 817 block of the loop mapping unit 810, and the reconstruction 822, intra-frame prediction 823, motion compensation or inter-frame prediction 824, and loop filter 825 blocks of the chrominance residual scaling unit 820 represent the processing applied in the original (non-mapped, non-shaping) domain.

[0192] like Figure 8 As shown, when LMCS is enabled, at least one of the following can be applied: reverse mapping (shaping) processing 814, forward mapping (shaping) processing 818, and chroma scaling processing 821. For example, reverse mapping processing can be applied to (reconstructed) luminance samples (or luminance samples or arrays of luminance samples) in a reconstructed image. Reverse mapping processing can be performed based on the piecewise function (reverse) index of the luminance sample. The piecewise function (reverse) index identifies the segment to which the luminance sample belongs. The output of the reverse mapping processing is a modified (reconstructed) luminance sample (or modified luminance sample or modified array of luminance samples). LMCS can be enabled or disabled at the tile group (or slice), image, or higher levels.

[0193] Forward mapping and / or chroma scaling can be applied to generate a reconstructed frame. The frame may include luma samples and chroma samples. A reconstructed frame based on luma samples is called a reconstructed luma frame, and a reconstructed frame based on chroma samples is called a reconstructed chroma frame. A combination of a reconstructed luma frame and a reconstructed chroma frame is called a reconstructed frame. The reconstructed luma frame can be generated based on forward mapping. For example, if inter-frame prediction is applied to the current block, forward mapping is applied to luma prediction samples derived from (reconstructed) luma samples in a reference frame. Since the (reconstructed) luma samples in the reference frame are generated based on backward mapping, forward mapping can be applied to the luma prediction samples, thus deriving mapped (shaped) luma prediction samples. Forward mapping can be performed based on the piecewise function index of the luma prediction samples. The piecewise function index can be derived based on the value of the luma prediction samples or the value of the luma samples in the reference frame used for inter-frame prediction. If intra-frame prediction (or intra-block copy (IBC)) is applied to the current block, forward mapping is not required because backward mapping has not yet been applied to the reconstructed samples in the current frame. The (reconstructed) brightness samples in the reconstructed brightness image are generated based on the mapped brightness prediction samples and the corresponding brightness residual samples.

[0194] The reconstructed chroma image can be generated based on chroma scaling. For example, the (reconstructed) chroma samples in the reconstructed chroma image can be based on the chroma residual samples (c) in the current block. res The derivation is based on the (scaled) chromaticity residual samples of the current block (c). resScale The chromaticity residual sample (c) is derived using the chromaticity residual scaling factor (cScaleInv, which can be called varScale) and the chromaticity residual scaling factor (cScaleInv, which can be called varScale). res The chromaticity residual scaling factor can be calculated based on the luminance prediction sample values ​​of the current block. For example, the chromaticity residual scaling factor can be calculated based on the luminance prediction sample value Y'. pred The average brightness value ave(Y' pred The scaling factor is calculated using this method. For reference, the (scaled) chromaticity residual sample derived based on the inverse transform / dequantization derivation can be referred to as c. resScale The chromaticity residual sample derived by performing (inverse) scaling on the (scaled) chromaticity residual sample can be called c. res .

[0195] Figure 9 An LMCS structure according to another embodiment of this document is shown. Figure 9 Reference Figure 8 To describe. Here, the main description is... Figure 9 LMCS structure and Figure 8 The differences between the LMCS structure 800. Figure 9 The loop mapping section and the luminance-dependent chrominance residual scaling section can be combined with Figure 8The loop mapping unit 810 and the luminance-related chromaticity residual scaling unit 820 operate in the same (similar) manner.

[0196] Reference Figure 9 The chroma residual scaling factor can be derived based on the luminance reconstruction samples. In this case, the average luminance value (avgYr) can be obtained (derived) based on neighboring luminance reconstruction samples outside the reconstruction block, rather than the luminance reconstruction samples inside the reconstruction block, and the chroma residual scaling factor is derived based on the average luminance value (avgYr). Here, the neighboring luminance reconstruction samples can be neighboring luminance reconstruction samples of the current block, or neighboring luminance reconstruction samples that include the Virtual Pipeline Data Unit (VPDU) of the current block. For example, when applying intra-frame prediction to the target block, the reconstruction samples can be derived from the prediction samples derived based on intra-frame prediction. In another example, when applying inter-frame prediction to the target block, a forward mapping is applied to the prediction samples derived based on inter-frame prediction, and the reconstruction samples are generated (derived) based on the shaped (or forward-mapped) luminance prediction samples.

[0197] Video / image information signaled via bitstream may include LMCS parameters (information about the LMCS). LMCS parameters can be configured with high-level syntax (HLS, including slice header syntax), etc. A detailed description and configuration of LMCS parameters will be described later. As mentioned above, the syntax table described in this document (and the following embodiments) can be configured / encoded at the encoder end and signaled to the decoder end via bitstream. The decoder can parse / decode the LMCS information (in the form of syntax components) in the syntax table. One or more embodiments described below can be combined. The encoder can encode the current frame based on the information about the LMCS, and the decoder can decode the current frame based on the information about the LMCS.

[0198] Loop mapping of the luminance component can improve compression efficiency by redistributing codewords across the dynamic range to adjust the dynamic range of the input signal. For luminance mapping, a forward mapping (shaping) function (FwdMap) and a corresponding backward mapping (shaping) function (InvMap) can be used. A piecewise linear model can be used to signal the FwdMap function; for example, the piecewise linear model may have 16 segments or bins, with segments of equal length. In one example, the InvMap function does not need to be signaled but is derived from the FwdMap function. That is, the backward mapping can be a forward mapping function. For example, the backward mapping function can be mathematically constructed as a symmetric function of the forward mapping, as reflected by the line y = x.

[0199] Loop (luminance) shaping can be used to map input luminance values ​​(samples) to modified values ​​in the shaping domain. The shaped values ​​can be encoded and then mapped back to the original (unmapped, unshaped) domain after reconstruction. To compensate for the interaction between the luminance and chrominance signals, chrominance residual scaling can be applied. Loop shaping is accomplished by specifying a high-level syntax for the shaper model. The shaper model syntax can signal the piecewise linear model (PWL model). For example, the shaper model syntax can signal a PWL model with 16 bins or segments of equal length. Forward lookup tables (FwdLUTs) and / or backward lookup tables (InvLUTs) can be derived based on the piecewise linear model. For example, the PWL model pre-computes 1024 forward (FwdLUT) and backward (InvLUT) lookup tables (LUTs). As an example, when deriving the forward lookup table FwdLUT, the backward lookup table InvLUT can be derived from the forward lookup table FwdLUT. A forward lookup table (FwdLUT) maps the input brightness value Yi to the changed value Yr, and a backward lookup table (InvLUT) maps the changed value Yr to the reconstructed value Y'i. The reconstructed value Y'i can be derived based on the input brightness value Yi.

[0200] In one example, SPS may include the syntax of Table 9 below. The syntax of Table 9 may include `sps_reshaper_enabled_flag` as a tool enabling flag. Here, `sps_reshaper_enabled_flag` can be used to specify whether a shaper is used in the encoded video sequence (CVS). That is, `sps_reshaper_enabled_flag` can be a flag that enables shaping in SPS. In one example, the syntax of Table 9 may be part of SPS.

[0201] [Table 9]

[0202]

[0203] In one example, the semantics of the syntactic elements sps_seq_parameter_set_id and sps_reshaper_enabled_flag can be shown in Table 10 below.

[0204] [Table 10]

[0205]

[0206]

[0207] In one example, the piece group header or slice header may include the syntax of Table 11 or Table 12 below.

[0208] [Table 11]

[0209]

[0210] [Table 12]

[0211]

[0212] The semantics of syntactic elements included in the syntax of Table 11 or Table 12 may include, for example, the matters disclosed in the following tables.

[0213] [Table 13]

[0214]

[0215] [Table 14]

[0216]

[0217]

[0218] As an example, once the integer reshaper flag (i.e., `sps_reshaper_enabled_flag`) is resolved in SPS, the tile group header can resolve additional data (i.e., information included in Table 13 or Table 14 above) used to construct the lookup table (FwdLUT and / or InvLUT). For this, the state of the SPS reshaper flag (`sps_reshaper_enabled_flag`) can be checked first in the slice header or tile group header. When `sps_reshaper_enabled_flag` is true (or 1), the additional flag, `tile_group_reshaper_model_present_flag` (or `slice_reshaper_model_present_flag`), can be resolved. The purpose of `tile_group_reshaper_model_present_flag` (or `slice_reshaper_model_present_flag`) can be to indicate the presence of an integer model. For example, when `tile_group_reshaper_model_present_flag` (or `slice_reshaper_model_present_flag`) is true (or 1), it indicates that an shaper exists for the current tile group (or the current slice). When `tile_group_reshaper_model_present_flag` (or `slice_reshaper_model_present_flag`) is false (or 0), it indicates that an shaper does not exist for the current tile group (or the current slice).

[0219] If an integer reshaper exists and is enabled in the current tile group (or slice), the integer reshaper model can be processed (i.e., `tile_group_reshaper_model()` or `slice_reshaper_model()`). Furthermore, the additional flag `tile_group_reshaper_enable_flag` (or `slice_reshaper_enable_flag`) can be resolved. `tile_group_reshaper_enable_flag` (or `slice_reshaper_enable_flag`) indicates whether the integer reshaping model is used in the current tile group (or slice). For example, if `tile_group_reshaper_enable_flag` (or `slice_reshaper_enable_flag`) is 0 (or false), it indicates that the integer reshaping model is not used in the current tile group (or slice). If `tile_group_reshaper_enable_flag` (or `slice_reshaper_enable_flag`) is 1 (or true), it indicates that the integer reshaping model is used in the current tile group (or slice).

[0220] As an example, `tile_group_reshaper_model_present_flag` (or `slice_reshaper_model_present_flag`) can be true (or 1), and `tile_group_reshaper_enable_flag` (or `slice_reshaper_enable_flag`) can be false (or 0). This means the integer model exists but is not used in the current tile group (or slice). In this case, the integer model can be used in future tile groups (or slices). As another example, `tile_group_reshaper_enable_flag` can be true (or 1), and `tile_group_reshaper_model_present_flag` can be false (or 0). In this case, the decoder uses the integer model from the previously initialized one.

[0221] When resolving the integer model (i.e., `tile_group_reshaper_model()` or `slice_reshaper_model()`) and `tile_group_reshaper_enable_flag` (or `slice_reshaper_enable_flag`), the conditions required for chroma scaling can be determined (evaluated). These conditions include condition 1 (the current tile group / slice has not yet been intra-coded) and / or condition 2 (the current tile group / slice has not yet been split into two separate encoded quadtree structures for luma and chroma; i.e., the block structure of the current tile group / slice is not a bitree structure). If condition 1 and / or condition 2 are true and / or `tile_group_reshaper_enable_flag` (or `slice_reshaper_enable_flag`) is true (or 1), then `tile_group_reshaper_chroma_residual_scale_flag` (or `slice_reshaper_chroma_residual_scale_flag`) can be resolved. When `tile_group_reshaper_chroma_residual_scale_flag` (or `slice_reshaper_chroma_residual_scale_flag`) is enabled (if 1 or true), chroma residual scaling is enabled for the current tile group (or slice). When `tile_group_reshaper_chroma_residual_scale_flag` (or `slice_reshaper_chroma_residual_scale_flag`) is disabled (if 0 or false), chroma residual scaling is disabled for the current tile group (or slice).

[0222] The purpose of tile group shaping models is to parse the data needed to construct lookup tables (LUTs). The idea behind these LUTs is that the distribution of the allowed range of luminance values ​​can be divided into multiple bins (e.g., 16 bins), which can be represented using a set of 16 PWL equations. Therefore, any luminance value falling within a given bin can be mapped to a changed luminance value.

[0223] Figure 10 A diagram illustrating an exemplary forward mapping is shown. Figure 10 Five warehouses are shown as an example.

[0224] Reference Figure 10The x-axis represents the input brightness value, and the y-axis represents the changed output brightness value. The x-axis is divided into 5 bins or slices, each of length L. That is, the five bins mapped to the changed brightness value have the same length. The forward lookup table (FwdLUT) can be constructed using data available from the tile group header (i.e., shaper data), thus facilitating mapping.

[0225] In one implementation, an output pivot point associated with the bin index can be calculated. The output pivot point can set (mark) the minimum and maximum boundaries of the output range for luminance codeword shaping. The calculation of the output pivot point can be performed by calculating the segmented cumulative distribution function (CDF) of the number of codewords. The output pivot range can be sliced ​​based on the maximum number of bins to be used and the size of the lookup table (FwdLUT or InvLUT). As an example, the output pivot range can be sliced ​​based on the product of the maximum number of bins and the size of the lookup table (LUT size * maximum number of bin indices). For example, if the product of the maximum number of bins and the size of the lookup table is 1024, the output pivot range can be sliced ​​into 1024 entries. This zigzag pattern of the output pivot range can be performed (applied or implemented) based on a scaling factor. In one example, the scaling factor can be derived based on Equation 1 below.

[0226] [Formula 1]

[0227] SF=(y2-y1)*(1<<FP_PREC)+c

[0228] In Equation 1, SF represents the scaling factor, and y1 and y2 represent the output pivot points corresponding to each compartment. Additionally, FP_PREC and c can be predetermined constants. The scaling factor determined based on Equation 1 can be referred to as the scaling factor used for forward shaping.

[0229] In another implementation, regarding reverse shaping (reverse mapping), for the defined range of bins to be used (i.e., from reshaper_model_min_bin_idx to reshape_model_max_bin_idx), the input shaping pivot point and the mapped reverse output pivot point corresponding to the mapping pivot point of the forward LUT are obtained (given by the bin index considered * initial codeword count). In another example, the scaling factor SF can be derived based on Equation 2 below.

[0230] [Equation 2]

[0231] SF=(y2-y1)*(1<<FP_PREC) / (x2-x1)

[0232] In Equation 2, SF represents the scaling factor, x1 and x2 represent the input pivot points, and y1 and y2 represent the output pivot points (reverse-mapped output pivot points) corresponding to each segment (bin). Here, the input pivot point can be a pivot point mapped based on a forward lookup table (FwdLUT), and the output pivot point can be a pivot point reverse-mapped based on a reverse lookup table (InvLUT). Additionally, FP_PREC can be a predetermined constant value. The FP_PREC in Equation 2 can be the same as or different from the FP_PREC in Equation 1. The scaling factor determined based on Equation 2 can be referred to as the scaling factor used for reverse shaping. During reverse shaping, the segmentation of the input pivot points can be performed based on the scaling factor of Equation 2. The scaling factor SF is used to slice the range of the input pivot points. Based on the segmented input pivot point, bin indices in the range from 0 to the minimum bin index (reshaper_model_min_bin_idx) and / or from the minimum bin index (reshaper_model_min_bin_idx) to the maximum bin index (reshape_model_max_bin_idx) are assigned pivot values ​​corresponding to the minimum and maximum bin values.

[0233] In one example, LMCS data (lmcs_data) can be included in the APS. For instance, the semantics of the APS could be 32 APSs for signaling used in encoding.

[0234] The following illustrates the syntax and semantics of an exemplary APS according to the implementation of this document.

[0235] [Table 15]

[0236]

[0237] [Table 16]

[0238]

[0239] Referring to Table 15 above, the type information of APS parameters (e.g., aps_params_type) can be parsed / signed in the APS. The type information of APS parameters can be parsed / signed after `adaptation_parameter_set_id`.

[0240] The APS_params_type, ALF_APS, and LMCS_APS included in Table 15 above can be described according to Table 3.2 included in Table 16. That is, the type of APS parameter applied to APS can be configured as shown in Table 3.2 included in Table 16, based on the APS_params_type included in Table 15 above. The syntactic elements included in Table 15 can be described with reference to Table 8. The descriptions related to APS are supported by the descriptions above along with Tables 1 to 8.

[0241] Referring to Table 16, for example, `aps_params_type` can be a syntactic element used to classify the type of the corresponding APS parameter. The type of an APS parameter can include ALF parameters and LMCS parameters. Referring to Table 16, if the value of the type information `aps_params_type` is 0, then the name of `aps_params_type` can be determined as `ALF_APS` (or `ALF APS`), and the type of the APS parameter can be determined as an ALF parameter (an APS parameter can represent an ALF parameter). In this case, the ALF data field (i.e., `alf_data()`) can be parsed / signaled to the APS. If the value of the type information `aps_params_type` is 1, then the name of `aps_params_type` can be determined as `LMCS_APS` (or `LMCS APS`), and the type of the APS parameter can be determined as an LMCS parameter (an APS parameter can represent an LMCS parameter). In this case, the LMCS data field (i.e., `lmcs_data()`) can be parsed / signaled to the APS.

[0242] Tables 17 and / or 18 below illustrate the syntax of the shaper model according to the embodiments. The shaper model may be referred to as an LMCS model. Here, although the shaper model is illustrated here as a tile group shaper by way of example, this specification is not necessarily limited to this embodiment. For example, the shaper model may be included in the APS, or the tile group shaper model may be referred to as a slice shaper model or LMCS data (LMCS data field). In addition, the prefix "reshaper_model" or "Rsp" may be used interchangeably with "lmcs". For example, in the following tables and the following description, reshaper_model_min_bin_idx, reshaper_model_delta_max_bin_idx, reshaper_model_max_bin_idx, RspCW, and RsepDeltaCW may be used interchangeably with lmcs_min_bin_idx, lmcs_delta_max_bin_idx, lmcs_max_bin_idx, lmcsCW, and lmcsDeltaCW, respectively.

[0243] The LMCS data (lmcs_data()) or shaper model (piece group shaper or slice shaper) included in Table 15 above can be represented as the syntax included in the table below.

[0244] [Table 17]

[0245]

[0246] [Table 18]

[0247]

[0248] The semantics of syntactic elements included in the syntax of Tables 17 and / or 18 may include, for example, the matters disclosed in the following tables.

[0249] [Table 19]

[0250]

[0251]

[0252] [Table 20]

[0253]

[0254]

[0255] The inverse mapping processing of the luminance samples in this paper can be described in the form of a standard document as shown in the table below.

[0256] [Table 21]

[0257]

[0258]

[0259] The identifiers for the piecewise function indexing of the luminance samples in this paper can be described in the form of standard literature as shown in the table below. In Table 22, idxYInv can be referred to as the inverse mapping index, and the inverse mapping index can be derived based on the reconstructed luminance sample (lumaSample).

[0260] [Table 22]

[0261]

[0262] Luminance mapping can be performed based on the above embodiments and examples, and the above syntax and components included therein are merely exemplary representations. The embodiments in this document are not limited to the above tables or formulas. Hereinafter, a method for performing chroma residual scaling (scaling of the chroma components of the residual sample) based on luminance mapping is described.

[0263] (Luminance-dependent) chroma residual scaling is designed to compensate for the interaction between the luminance signal and its corresponding chroma signal. For example, at the tile group level, a signal is also used to indicate whether chroma residual scaling is enabled. In one example, if luminance mapping is enabled and no dual-tree partitioning (also known as a separate chroma tree) is applied to the current tile group, an additional flag is signaled to indicate whether luminance-dependent chroma residual scaling is enabled. In other examples, luminance-dependent chroma residual scaling is disabled when luminance mapping is not used, or when dual-tree partitioning is used in the current tile group. In yet another example, luminance-dependent chroma residual scaling is always disabled for chroma tiles with an area less than or equal to 4.

[0264] Chromaticity residual scaling can be based on the average value of the corresponding luma prediction block (the luma component of the prediction block for which intra-frame prediction mode and / or inter-frame prediction mode are applied). The scaling operation at the encoder and / or decoder side can be implemented using fixed-point integer arithmetic based on Equation 3 below.

[0265] [Formula 3]

[0266] c'=sign(c)*((abs(c)*s+2CSCALE_FP_PREC-1)>>CSCALE_FP_PREC)

[0267] In Equation 3, c' represents the scaled chroma residual sample (scaled chroma components of the residual sample), c represents the chroma residual sample (chroma residual sample, chroma components of the residual sample), s represents the chroma residual scaling factor, and CSCALE_FP_PREC represents a (predefined) constant value to specify precision. For example, CSCALE_FP_PREC can be 11.

[0268] Figure 11 This is a flowchart illustrating a method for deriving a chromaticity residual scaling index according to an embodiment of this document. Figure 11 The method in can be based on Figure 8 And included in with Figure 8 The relevant descriptions use tables, formulas, variables, arrays, and functions to perform the operation.

[0269] In step S1110, the prediction mode of the current block can be determined based on the prediction mode information as either intra-frame prediction mode or inter-frame prediction mode. If the prediction mode is intra-frame prediction mode, the current block or its prediction samples are considered to be in the reshaped (mapped) region. If the prediction mode is inter-frame prediction mode, the current block or its prediction samples are considered to be in the original (unmapped, unreshaped) region.

[0270] In step S1120, when the prediction mode is intra-frame prediction mode, the average (or the brightness prediction sample of the current block) of the current block can be calculated (derived). That is, the average of the current block in the already shaped region is directly calculated. The average can also be called the mean or average value.

[0271] In step S1121, when the prediction mode is inter-frame prediction mode, forward shaping (forward mapping) can be performed on the luminance prediction samples of the current block. Forward shaping maps the luminance prediction samples based on the inter-frame prediction mode from the original region to the shaped region. In one example, the forward shaping of the luminance prediction samples can be performed based on the shaping models described in Tables 17 and / or 18 above.

[0272] In step S1122, the average of the forward-shaped (forward-mapped) brightness prediction samples can be calculated (derived). That is, the averaging process of the forward-shaped results can be performed.

[0273] In step S1130, the chroma residual scaling index can be calculated. When the prediction mode is intra-frame prediction mode, the chroma residual scaling index can be calculated based on the average of the luminance prediction samples. When the prediction mode is inter-frame prediction mode, the chroma residual scaling index can be calculated based on the average of the forward-shaped luminance prediction samples.

[0274] In an implementation, the chroma residual scaling index can be calculated using a for loop syntax. An exemplary for loop syntax for deriving (calculating) the chroma residual scaling index is shown below.

[0275] [Table 23]

[0276]

[0277] In Table 23, idxS represents the chroma residual scaling index, idxFound represents the index that indicates whether the chroma residual scaling index that satisfies the conditions of the if statement has been obtained, S represents a predefined constant value, and MaxBinIdx represents the maximum allowed bin index. ReshapPivot[idxS+1] (in other words, LmcsPivot[idxS+1]) can be derived based on Table 19 and / or Table 20 above.

[0278] In an implementation, the chroma residual scaling factor can be derived based on the chroma residual scaling index. Equation 4 is an example for deriving the chroma residual scaling factor.

[0279] [Formula 4]

[0280] s = ChromaScaleCoef[idxS]

[0281] In Equation 4, s represents the chroma residual scaling factor, and ChromaScaleCoef can be a variable (or array) derived based on Table 19 and / or Table 20 above.

[0282] As described above, the average luminance value of the reference sample can be obtained, and the chromaticity residual scaling factor can be derived based on the average luminance value. As described above, the chromaticity component residual sample can be scaled based on the chromaticity residual scaling factor, and a chromaticity component reconstruction sample can be generated based on the scaled chromaticity component residual sample.

[0283] In one embodiment of this document, a signaling structure for efficiently applying the aforementioned LMCS is proposed. According to this embodiment, for example, LMCS data can be included in the HLS (i.e., APS), and the LMCS model (shaper model) can be adaptively derived by signaling the APS ID (called the header information) through the lower-level header information of the APS (i.e., frame header, slice header). The LMCS model can be derived based on LMCS parameters. Furthermore, for example, multiple APS IDs can be signaled through the header information, and thus, different LMCS models can be applied on a block-by-block basis within the same frame / slice.

[0284] In one embodiment of this document, a method for efficiently performing the operations required for LMCS is proposed. Based on the semantics described above in Tables 19 and / or 20, InvScaleCoeff[i] needs to be derived through a division operation of the fragment length lmcsCW[i] (also referred to as RspCW[i] in this document). The fragment length of the reverse-mapped sequence may not be a power of 2, meaning that division cannot be performed via bit shifting.

[0285] For example, calculating InvScaleCoeff might require up to 16 divisions per slice. According to Tables 19 and / or 20 above, for 10-bit encoding, lmcsCW[i] ranges from 8 to 511; therefore, to implement division using the LUT according to lmcsCW[i], the LUT size must be 504. Furthermore, for 12-bit encoding, lmcsCW[i] ranges from 32 to 2047; therefore, the LUT size needs to be 2016 to implement division using the LUT according to lmcsCW[i]. That is, division is expensive in terms of hardware implementation; therefore, it is preferable to avoid division whenever possible.

[0286] In one aspect of this implementation, lmcsCW[i] can be constrained to a fixed number (or a predetermined number) multiple. Therefore, the lookup table (LUT) for division (the capacity or size of the LUT) can be reduced. For example, if lmcsCW[i] becomes a multiple of 2, the size of the LUT replacing the division process can be halved.

[0287] In another aspect of this implementation, for encoding with a higher internal bit depth, it is proposed that, beyond the existing constraint that "the value of lmcsCW[i] should be in the range of (OrgCW>>3) to (OrgCW<<3-1)," if the encoding bit depth is higher than 10, lmcsCW[i] is further constrained to be a multiple of 1<<(BitDepthY-10). Here, BitDepthY can be the luminance bit depth. Therefore, the possible number of lmcsCW[i] does not change with the encoding bit depth, and the size of the LUT required to calculate InvScaleCoeff does not increase due to the higher encoding bit depth. For example, for a 12-bit internal encoding bit depth, the limit of lmcsCW[i] is a multiple of 4, so the LUT that replaces the division process will be the same as that used for 10-bit encoding. This aspect can be implemented alone, but it can also be implemented in combination with the aspects described above.

[0288] In another aspect of this implementation, lmcsCW[i] can be constrained to a narrower range. For example, lmcsCW[i] can be constrained to the range from (OrgCW>>1) to (OrgCW<<1)-1. Then, for 10-bit encoding, the range of lmcsCW[i] can be [32,127], thus requiring only a LUT of size 96 to compute InvScaleCoeff.

[0289] In another aspect of this implementation, lmcsCW[i] can be approximated as the number closest to a power of 2 and used in shaper design. Therefore, division in the reverse mapping can be performed (and replaced) via bit shifting.

[0290] In one embodiment according to this document, a constraint on the range of LMCS codewords is proposed. According to Table 8 above, the values ​​of the LMCS codewords are in the range from (OrgCW>>3) to (OrgCW<<3)-1. This codeword range is too wide. When there is a large difference between RspCW[i] and OrgCW, it may lead to visual artifacts.

[0291] According to one embodiment of this document, it is proposed to constrain the codewords of the LMCS PWL mapping to a narrow range. For example, the range of lmcsCW[i] can be from (OrgCW>>1) to (OrgCW<<1)-1.

[0292] In one embodiment according to this document, a single chroma residual scaling factor is proposed for chroma residual scaling in LMCS. Existing methods for deriving the chroma residual scaling factor use the average value of the corresponding luma block and derive the slope of each segment of the inverse luma map as the corresponding scaling factor. Furthermore, the process of identifying the segment index requires the availability of the corresponding luma block, which introduces latency issues. This is undesirable for hardware implementation. According to this embodiment, scaling in the chroma block may not depend on the luma block values ​​and may not require identifying the segment index. Therefore, chroma residual scaling processing in LMCS can be performed without latency issues.

[0293] In one embodiment according to this document, a single chroma scaling factor can be derived from both the encoder and decoder based on luminance LMCS information. The chroma residual scaling factor can be updated when the LMCS luminance model is received. For example, the single chroma residual scaling factor can be updated when the LMCS model is updated.

[0294] The following shows an example of obtaining a single chroma scaling factor according to this embodiment.

[0295] [Table 24]

[0296]

[0297] Referring to Table 24, a single chroma scaling factor (e.g., ChromaScaleCoeff or ChromaScaleCoeffSingle) can be obtained by averaging the inverse luminance mapping slopes of all segments within LMCS_min_bin_idx and lmcs_max_bin_idx.

[0298] Figure 12 Linear fitting of the pivot point according to the embodiment of this document is shown. Figure 12 The pivot points P1, Ps, and P2 are shown in the diagram. The following implementations or examples thereof will be described using... Figure 12 To describe.

[0299] In this example implementation, a single chroma scaling factor can be obtained based on a linear approximation of the luminance PWL mapping between pivot points lmcs_min_bin_idx and lmcs_max_bin_idx+1 (LmcsMaxBinIdx+1). That is, the inverse slope of the linear mapping can be used as the chroma residual scaling factor. For example, Figure 12 Line 1 can be a straight line connecting pivot points P1 and P2. (Refer to...) Figure 12In P1, the input value is x1 and the mapping value is 0, and in P2, the input value is x2 and the mapping value is y2. The inverse slope (inverse scale) of linear line 1 is (x2-x1) / y2, and the single chroma scaling factor ChromaScaleCoeffSingle can be calculated based on the input and mapping values ​​of pivot points P1 and P2 and the following formula.

[0300] [Formula 5]

[0301] ChromaScaleCoeffSingle = (x2 - x1) * (1 << CSCALE_FP_PREC) / y2 In Equation 5, CSCALE_FP_PREC represents the shift factor; for example, CSCALE_FP_PREC can be a predetermined constant value. In one example, CSCALE_FP_PREC could be 11.

[0302] In another example according to this implementation, refer to Figure 12 The input value at pivot point Ps is min_bin_idx+1, and the mapping value at pivot point Ps is ys. Therefore, the inverse slope (inverse scale) of linear line 1 can be calculated as (xs-x1) / ys, and the single chroma scaling factor ChromaScaleCoeffSingle can be calculated based on the input and mapping values ​​of pivot points P1 and Ps, as well as the following formula.

[0303] [Formula 6]

[0304] ChromaScaleCoeffSingle=(xs-x1)*(1<<CSCALE_FP_PREC) / ys

[0305] In Equation 6, CSCALE_FP_PREC represents the shift factor (bit shift factor). For example, CSCALE_FP_PREC can be a predetermined constant value. In one example, CSCALE_FP_PREC can be 11, and a reverse-scaled bit shift can be performed based on CSCALE_FP_PREC.

[0306] In another example according to this embodiment, a single chroma residual scaling factor can be derived based on a linear approximation line. Examples of deriving a linear approximation line may include a linear connection of pivot points (i.e., lmcs_min_bin_idx, lmcs_max_bin_idx+1). For example, the linear approximation result can be represented by codewords of a PWL mapping. The mapping value y2 at P2 can be the sum of codewords for all bins (fragments), and the difference (x2-x1) between the input value at P2 and the input value at P1 is OrgCW*(lmcs_max_bin_idx-lmcs_min_bin_idx+1) (for OrgCW, see Table 19 and / or Table 20 above). An example of obtaining a single chroma scaling factor according to the above embodiment is shown below.

[0307] [Table 25]

[0308]

[0309] Referring to Table 25, a single chroma scaling factor (e.g., ChromaScaleCoeffSingle) can be obtained from two pivot points (i.e., lmcs_min_bin_idx, lmcs_max_bin_idx). For example, the inverse slope of a linear map can be used as a chroma scaling factor.

[0310] In another example of this implementation, a single chroma scaling factor can be obtained by linearly fitting the pivot point to minimize the error (or mean square error) between the linear fit and the existing PWL mapping. This example is more accurate than simply connecting the two pivot points at lmcs_min_bin_idx and lmcs_max_bin_idx. There are many ways to find the optimal linear mapping; an example of this is described below.

[0311] In one example, the parameters b1 and b0 used to minimize the sum of least squares errors in the linear fit y = b1*x + b0 can be calculated based on Equations 7 and / or 8 below.

[0312] [Formula 7]

[0313]

[0314] [Formula 8]

[0315]

[0316] In Equations 7 and 8, x is the original luminance value, and y is the shaped luminance value. and It is the mean of x and y, x i and y i This represents the value of the i-th pivot point.

[0317] Reference Figure 12 Another simple approximation for identifying linear mappings is given as:

[0318] - Linear line 1 is obtained by connecting the pivot points of the PWL mapping at lmcs_min_bin_idx and lmcs_max_bin_idx+1, and lmcs_pivots_linear[i] is calculated using input values ​​that are multiples of OrgW.

[0319] - Use linear line 1 and sum the differences in pivot point mapping values ​​using PWL mapping.

[0320] - Obtain the average difference avgDiff.

[0321] - Adjust the final pivot point of the linear line based on the average difference, for example, 2*avgDiff

[0322] - Use the backslope of the adjusted linear line as the chromaticity residual scale.

[0323] Based on the linear fit described above, the chroma scaling factor (i.e., the inverse slope of the forward mapping) can be derived (obtained) based on Equation 9 or Equation 10 below.

[0324] [Formula 9]

[0325] ChromaScaleCoeffSingle=OrgCW*(1<<CSCALE_FP_PREC)

[0326] / lmcs_pivots_linear[lmcs_min_bin_idx+1]

[0327] [Formula 10]

[0328] ChromaScaleCoeffSingle=OrgCW*(lmcs_max_bin_idx-lmcs_max_bin_idx+1)

[0329] *(1< <CSCALE_FP_PREC) / lmcs_pivots_linear[lmcs_max_bin_idx+1]

[0330] In the above equation, lmcs_pivots_lienar[i] can be the mapping value of a linear mapping. For a linear mapping, all segments of the PWL mapping between the minimum bin index and the maximum bin index can have the same LMCS codeword (lmcsCW). That is, lmcs_pivots_linear[lmcs_min_bin_idx+1] can be the same as lmcsCW[lmcs_min_bin_idx].

[0331] Additionally, in Equations 9 and 10, CSCALE_FP_PREC represents the shift factor (bit shift factor), which can be a predetermined constant value. In one example, CSCALE_FP_PREC could be 11.

[0332] By using a single chroma residual scaling factor (ChromaScaleCoeffSingle), it is no longer necessary to calculate the average of corresponding luma blocks and find the index in the PWL linear mapping to obtain the chroma residual scaling factor. Therefore, the coding efficiency of using chroma residual scaling can be increased. This not only eliminates the dependency on corresponding luma blocks and solves the latency problem, but also reduces complexity.

[0333] The LMCS data-related semantics and / or chromaticity sample-related luminance-dependent chromaticity residual scaling process according to the above-described embodiment can be described in the following table in the form of standard literature.

[0334] [Table 26]

[0335]

[0336] [Table 27]

[0337]

[0338]

[0339] In another embodiment of this document, the encoder can determine parameters related to a single chroma scaling factor and signal these parameters to the decoder. Using the signaling, the encoder can derive the chroma scaling factor using other information available at the encoder. This embodiment aims to eliminate the chroma residual scaling delay problem.

[0340] For example, another example of identifying a linear mapping to be used to determine the chromaticity residual scaling factor is given below:

[0341] -Calculate lmcs_pivots_linear[i] using input values ​​that are multiples of OrgW, by connecting the pivot points of the PWL mapping at lmcs_min_bin_idx and lmcs_max_bin_idx+1.

[0342] - Use the linear line 1 and the brightness PWL mappings to obtain a weighted sum of the differences in the pivot point mapping values. The weights may be based on encoder statistics (e.g., a histogram of the bin).

[0343] - Obtain the weighted average difference avgDiff.

[0344] - Adjust the final pivot point of linear line 1 based on the weighted average difference, for example, 2*avgDiff

[0345] - The chromaticity residual scale is calculated using the reverse slope of the adjusted linear line.

[0346] The following is an example of the syntax for signaling the y-value used in chroma scaling factor derivation.

[0347] [Table 28]

[0348]

[0349] In Table 28, the syntax element `lmcs_chroma_scale` specifies a single chroma (residual) scaling factor (ChromaScaleCoeffSingle = lmcs_chroma_scale) used for LMCS chroma residual scaling. That is, information about the chroma residual scaling factor can be directly signaled, and the signaled information can be derived from the chroma residual scaling factor. In other words, the value of the signaled information about the chroma residual scaling factor can be (directly) derived from the value of a single chroma residual scaling factor. Here, the syntax element `lmcs_chroma_scale` can be signaled along with other LMCS data (i.e., syntax elements related to the absolute value and symbol of codewords, etc.).

[0350] Alternatively, the encoder may signal only the necessary parameters to derive the chroma residual scaling factor at the decoder. To derive the chroma residual scaling factor at the decoder, an input value x and a mapping value y are required. Since the x value is the bin length and is a known number, it does not need to be signaled. After all, only the y value needs to be signaled to derive the chroma residual scaling factor. Here, the y value can be the mapping value of any pivot point in a linear mapping (i.e., Figure 12 (The mapping value of P2 or Ps in the middle).

[0351] The following is an example of using a signal to notify the mapping value used to derive the chroma residual scaling factor.

[0352] [Table 29]

[0353]

[0354] [Table 30]

[0355]

[0356] One of the syntaxes in Tables 29 and 30 above can be used to signal the y-value at any linear pivot point specified by the encoder and decoder. That is, the encoder and decoder can use the same syntax to derive the y-value.

[0357] First, the implementation according to Table 29 is described. In Table 29, lmcs_cw_linear can represent the mapping value at Ps or P2. That is, in the implementation according to Table 29, a fixed number can be notified by signaling lmcs_cw_linear.

[0358] In the example according to this implementation, if lmcs_cw_linear represents a warehouse mapping value (i.e., Figure 12 The chroma scaling factor can be derived from the following formula: lmcs_pivots_linear[lmcs_min_bin_idx+1] in Ps.

[0359] [Equation 11]

[0360] ChromaScaleCoeffSingle=OrgCW*(1<<CSCALE_FP_PREC) / lmcs_cw_linear

[0361] In another example according to this implementation, if lmcs_cw_linear represents lmcs_max_bin_idx+1 (i.e., Figure 12 The chroma scaling factor can be derived from the following formula based on lmcs_pivots_linear[lmcs_max_bin_idx+1] in P2.

[0362] [Equation 12]

[0363] ChromaScaleCoeffSimgle=OrgCW*(lmcs_max_bin_idx-lmcs_max_bin_idx+1)

[0364] *(1< <CSCALE_FP_PREC) / lmcs_cw_linear

[0365] In the above formula, CSCALE_FP_PREC represents the shift factor (bit shift factor). For example, CSCALE_FP_PREC can be a predetermined constant value. In one example, CSCALE_FP_PREC could be 11.

[0366] Next, the implementation according to Table 30 is described. In this implementation, lmcs_cw_linear can be signaled as a Δ value relative to a fixed number (i.e., lmcs_delta_abs_cw_linear, lmcs_delta_sign_cw_linear_flag). In the example of this implementation, lmcs_cw_linear represents lmcs_pivots_linear[lmcs_min_bin_idx+1] (i.e., Figure 12 When mapping values ​​in Ps), lmcs_cw_linear_delta and lmcs_cw_linear can be derived based on the following formula.

[0367] [Equation 13]

[0368] lmcs_cw_linear_delta=(1-2*lmcs_delta_sign_cw_linear_flag)*lmcs_delta_abs_linear_cw

[0369] [Formula 14]

[0370] lmcs_cw_linear=lmcs_cw_linear_delta+OrgCW

[0371] In another example of this implementation, when lmcs_cw_linear represents lmcs_pivots_linear[lmcs_max_bin_idx+1] (i.e., Figure 12 When mapping values ​​in P2), lmcs_cw_linear_delta and lmcs_cw_linear can be derived based on the following formula.

[0372] [Formula 15]

[0373] lmcs_cw_linear_delta=(1-2*lmcs_delta_sign_cw_linear_flag)*lmcs_delta_abs_linear_cw

[0374] [Formula 16]

[0375] lmcs_cw_linear=lmcs_cw_linear_delta

[0376] +OrgCW*(lmcs_max_bin_idx-lmcs_max_bin_idx+1)

[0377] In the above formula, OrgCW can be a value derived based on Table 19 and / or Table 20 above.

[0378] The LMCS data-related semantics and / or chromaticity sample-related luminance-dependent chromaticity residual scaling process according to the above-described embodiment can be described in the following table in the form of standard literature.

[0379] [Table 31]

[0380]

[0381]

[0382] [Table 32]

[0383]

[0384] Figure 13 An example of linear shaping (or linear mapping, linear transformation) according to an embodiment of this document is shown. That is, in this embodiment, the use of a linear shaper in LMCS is proposed. For example, Figure 13 This example could involve forward linear shaping (mapping).

[0385] Reference Figure 13 The linear shaper can include two pivot points, P1 and P2. P1 and P2 can represent input and mapped values, respectively; for example, P1 can be (min_input, 0) and P2 can be (max_input, max_mapped). Here, min_input represents the minimum input value, and max_input represents the maximum input value. Any input value less than or equal to min_input is mapped to 0, and any input value greater than max_input is mapped to max_mapped. Any input brightness value within min_input and max_input is linearly mapped to the other values. Figure 13 An example of mapping is shown. Pivot points P1 and P2 can be determined at the encoder, and a piecewise linear mapping can be approximated using linear fitting.

[0386] In another embodiment according to this document, another example of a method for signaling a linear shaper can be proposed. The pivot points P1 and P2 of the linear shaper model can be explicitly signaled. The following shows an example of the syntax and semantics of explicitly signaling the linear shaper model according to this example.

[0387] [Table 33]

[0388] lmcs_data(){ descriptor lmcs_min_input ue(v) lmcs_max_input ue(v) lmcs_max_mapped ue(v)

[0389] [Table 34]

[0390]

[0391] Referring to Table 33 and Table 34, the input value of the first pivot point can be derived based on the syntactic element lmcs_min_input, and the input value of the second pivot point can be derived based on the syntactic element lmcs_max_input. The mapped value of the first pivot point can be a predetermined value (a value known to both the encoder and the decoder). For example, the mapped value of the first pivot point is 0. The mapped value of the second pivot point can be derived based on the syntactic element lmcs_max_mapped. That is, the linear shaping model can be explicitly (directly) signaled based on the information signaled in the syntax of Table 33.

[0392] Alternatively, lmcs_max_input and lmcs_max_mapped can be signaled as Δ values. The following table shows an example of the syntax and semantics of signaling the linear shaping model as Δ values.

[0393] [Table 35]

[0394] lmcs_data(){ descriptor lmcs_min_input ue(v) lmcs_max_input_delta ue(v) lmcs_max_mapped_delta ue(v)

[0395] [Table 36]

[0396]

[0397] Referring to Table 36, the input value of the first pivot point can be derived based on the syntactic element lmcs_min_input. For example, lmcs_min_input can have a mapped value of 0. lmcs_max_input_delta can specify the difference between the input value of the second pivot point and the maximum luminance value (i.e., (1<<bitdepthY)-1). lmcs_max_mapped_delta can specify the difference between the mapped value of the second pivot point and the maximum luminance value (i.e., (1<<bitdepthY)-1).

[0398] According to the implementation of this document, forward mapping of luminance prediction samples, backward mapping of luminance reconstruction samples, and chroma residual scaling can be performed based on the above example of a linear shaper. In one example, backward scaling of luminance (reconstructed) samples (pixels) in backward mapping based on a linear shaper may require only one backward scaling factor. The same applies to forward mapping and chroma residual scaling. That is, the steps of determining ScaleCoeff[i], InvScaleCoeff[i], and ChromaScaleCoeff[i] (where i is the bin index) can be replaced by a single factor. Here, a single factor is a fixed-point representation of the (positive) slope or negative slope of the linear mapping. In one example, the backward luminance mapping scaling factor (the backward scaling factor in the backward mapping of luminance reconstruction samples) can be derived based on at least one of the following formulas.

[0399] [Equation 17]

[0400] InvScaleCoeffSingle=OrgCW / lmcsCWLinear

[0401] [Formula 18]

[0402] InvScaleCoeffSingle=OrgCW*(lmcs_max_bin_idx-lmcs_max_bin_idx+1)

[0403] / lmcsCWLinearAll

[0404] [Formula 19]

[0405] InvScaleCoeffSingle=(lmcs_max_input-lmcs_min_input) / lmcsCWLinearAll

[0406] The lmcsCWLinear of Equation 17 can be derived from Table 31 above. The lmcsCWLinearALL of Equations 18 and 19 can be derived from at least one of Tables 33 to 36 above. In Equation 17 or Equation 18, OrgCW can be derived from Table 19 and / or Table 20.

[0407] The table below describes the formulas and syntax (conditional statements) that instruct the forward mapping processing of luminance samples (i.e., luminance prediction samples) in image reconstruction. In the table and formulas below, FP_PREC is a constant value for bit shifting and can be a predetermined value. For example, FP_PREC can be 11 or 15.

[0408] [Table 37]

[0409]

[0410] [Table 38]

[0411]

[0412] Table 37 can be used to derive the luminance samples of the forward mapping in the luminance mapping process based on Tables 17 to 20 above. That is, Table 37 can be described together with Tables 19 and 20. In Table 37, the luminance (prediction) samples PredMapPSamples[i][j] of the forward mapping as the output can be derived from the luminance (prediction) samples predSamples[i][j] as the input. The idxY of Table 37 can be called the (forward) mapping index, and the mapping index can be derived based on the predicted luminance samples.

[0413] Table 38 can be used to derive the luminance samples of the forward mapping in the luminance mapping based on the linear shaper. For example, lmcs_min_input, lmcs_max_input, lmcs_max_mapped, and ScaleCoeffSingle in Table 38 can be derived from at least one of Tables 33 to 36. In Table 38, when "lmcs_min_input < predSamples[i][j] < lmcs_max_input", the luminance (prediction) samples PredMapSamples[i][j] of the forward mapping can be derived from the input luminance (prediction) samples predSamples[i][j] as the output. Comparing between Table 37 and Table 38, the change relative to the existing LMCS according to the application of the linear shaper can be seen from the perspective of the forward mapping.

[0414] The following formula and table describe the inverse mapping process of the luminance samples (i.e., the luminance reconstruction samples). In the following formula and table, the "lumaSample" as the input can be the luminance reconstruction sample before the inverse mapping (before modification). The "invSample" as the output can be the (modified) luminance reconstruction sample of the inverse mapping. In other cases, the clipped invSample can be called the modified luminance reconstruction sample.

[0415] [Equation 20]

[0416] invSample = InputPivot[idxYInv] + (InvScaleCoeff[idxYInv] *

[0417] (lumaSample - LmcsPiVot[idxYInv]) + (1 << (FP_PREC - 1))) >> FP_PREC

[0418] [Equation 21]

[0419] invSample = lmcs_min_input

[0420] +(InvScaleCoeffSingle*(lumaSample-lmcs min input)+(1<<(FP_PREC-1)))

[0421] >>FP_PREC

[0422] [Table 39]

[0423]

[0424] [Table 40]

[0425]

[0426] Equation 21 can be used to derive the luminance sample of the inverse mapping in the luminance mapping according to this document. In Equation 20, the index idxInv can be derived based on Tables 50, 51 or 52, which are described later.

[0427] Equation 21 can be used to derive the luminance sample of the reverse mapping from the luminance mapping based on the application of the linear shaper. For example, lmcs_min_input of Equation 21 can be derived from at least one of Tables 33 to 36. By comparing Equations 20 and 21, the changes relative to the existing LMCS based on the application of the linear shaper can be seen from the perspective of the forward mapping.

[0428] Table 39 may include examples of formulas for deriving the luminance samples of the inverse mapping in the luminance mapping. For example, the index idxInv may be derived based on Tables 50, 51, or 52, which are described later.

[0429] Table 40 may include other examples of formulas for deriving the luminance samples of the inverse mapping in the luminance mapping. For example, lmcs_min_input and / or lmcs_max_mapped in Table 40 may be derived from at least one of Tables 33 to 36, and / or InvScaleCoeffSingle in Table 40 may be at least Tables 33 to 36, and / or Equations 17 to 19 may be derived from one.

[0430] Based on the example above using a linear shaper, the segmented index recognition process can be omitted. That is, in this example, since there is only one segment with a valid luminance pixel for shaping, the segmented index recognition process used for inverse luminance mapping and chroma residual scaling can be eliminated. Therefore, the complexity of inverse luminance mapping can be reduced. Furthermore, the latency issue caused by relying on luminance segmented index recognition during chroma residual scaling can be eliminated.

[0431] According to the embodiments using the above linear shaper, the following advantages can be provided for LMCS: i) The encoder shaper design can be simplified, thus preventing possible artifacts caused by sudden changes between piecewise linear segments; ii) The decoder inverse mapping process that can remove the piecewise index identification process can be simplified by eliminating the piecewise index identification process; iii) By removing the piecewise index identification process, the delay problem caused by depending on the corresponding luminance block in chrominance residual scaling can be removed; iv) The signaling overhead can be reduced, and more frequent updates of the shaper can be made more feasible; v) For many places where a cycle of 16 segments was required in the past, the cycle can be eliminated. For example, in order to derive InvScaleCoeff[i], the number of division operations according to lmcsCW[i] can be reduced to 1.

[0432] In another embodiment according to this document, a flexible bin-based LMCS is proposed. Here, a flexible bin may refer to a bin whose number is not fixed to a predetermined (predefined, specific) number. In the existing embodiments, the number of bins in LMCS is fixed to 16, and for the input sample values, these 16 bins are equally distributed. In this embodiment, a flexible number of bins is proposed, and these segments (bins) may not be equally distributed in terms of the original pixel values.

[0433] The following table exemplarily shows the syntax of the LMCS data (data field) according to this embodiment and the semantics of the syntax elements included therein.

[0434] [Table 41]

[0435]

[0436] [Table 42]

[0437]

[0438]

[0439] Referring to Table 41, the information lmcs_num_bins_minus1 regarding the number of bins can be signaled. Referring to Table 42, lmcs_num_bins_minus1 + 1 can be equal to the number of bins, and the number of bins can be in the range from 1 to (1 << BitDepthY) - 1. For example, lmcs_num_bins_minus1 or lmcs_num_bins_minus1 + 1 can be a multiple of a power of 2.

[0440] In the implementation described with Tables 41 and 42, regardless of whether the shaper is linear (signaling of lmcs_num_bins_minus1), the number of pivot points can be derived based on lmcs_num_bins_minus1 (information about the number of bins), and the input and mapped values ​​of the pivot points (LmcsPivot_input[i], LmcsPivot_mapped[i]) can be derived based on the sum of the codeword values ​​(lmcs_delta_input_cw[i], lmcs_delta_mapped_cw[i]) notified by the signal (here, the initial input value LmcsPivot_input[0] and the initial output value LmcsPivot_mapped[0] are 0).

[0441] Figure 14 An example of a linear forward mapping in the implementation of this document is shown. Figure 15 An example of inverse forward mapping in the implementation of this document is shown.

[0442] According to Figure 14 and Figure 15 In this implementation, a method supporting both regular LMCS and linear LMCS is proposed. In an example according to this implementation, regular LMCS and / or linear LMCS can be indicated based on the syntactic element lmcs_is_linear. In the encoder, after determining the linear LMCS line, the mapping value (i.e., Figure 14 and Figure 15 The mapped values ​​in pL can be divided into equal segments (i.e., LmcsMaxBinIdx - lmcs_min_bin_idx + 1). The codewords in the LmcsMaxBinIdx bin can be signaled using the syntax of the LMCS data or shaper mode described above.

[0443] The following table exemplarily illustrates the syntax of an example LMCS data (data field) according to this implementation and the semantics of the syntactic elements included therein.

[0444] [Table 43]

[0445]

[0446] [Table 44]

[0447]

[0448]

[0449] The following table exemplarily illustrates the syntax of LMCS data (data fields) and the semantics of the syntactic elements included therein, in another example according to this implementation.

[0450] [Table 45]

[0451]

[0452] [Table 46]

[0453]

[0454]

[0455] Referring to Tables 43 to 46, when lmcs_is_linear_flag is true, all LMCSDeltaCW[i] between lmcs_min_bin_idx and LmcsMaxBinIdx can have the same value. That is, the lmcsCW[i] of all segments between lmcs_min_bin_idx and LmcsMaxBinIdx can have the same value. The scale, inverse scale, and chromaticity scale of all segments between lmcs_min_bin_idx and lmcsMaxBinIdx can be the same. Then, if the linear shaper is true, there is no need to derive the segment index; it can use the scale and inverse scale from only one segment.

[0456] The following table exemplarily illustrates the segmentation index identification process according to this embodiment.

[0457] [Table 47]

[0458]

[0459] According to another embodiment of this document, the application of conventional 16-segment PWL LMCS and linear LMCS may depend on a higher level of syntax (i.e., sequence level).

[0460] The following table exemplarily illustrates the syntax of the SPS according to this embodiment and the semantics of the syntactic elements included therein.

[0461] [Table 48]

[0462]

[0463] [Table 49]

[0464]

[0465] Referring to Tables 48 and 49, the enabling of regular LMCS and / or linear LMCS can be determined (notified by signals) by the syntactic elements included in the SPS. Referring to Table 48, based on the syntactic element sps_linear_lmcs_enabled_flag, one of the regular LMCS or linear LMCS can be used sequentially.

[0466] Additionally, whether only linear LMCS, regular LMCS, or both are enabled can also depend on the profile level. In one example, for a particular profile (i.e., the SDR profile), only linear LMCS may be allowed; for another profile (i.e., the HDR profile), only regular LMCS may be allowed; and for yet another profile, both regular LMCS and / or linear LMCS may be allowed.

[0467] According to another embodiment of this document, LMCS segmented index identification processing can be used in inverse luminance mapping and chroma residual scaling. In this embodiment, segmented index identification processing can be used for blocks where chroma residual scaling is enabled, and is also invoked for all luminance samples in the shaping (mapping) domain. This embodiment aims to keep its complexity low.

[0468] The following describes the identification and processing (derivation) of piecewise function indices.

[0469] [Table 50]

[0470]

[0471] In the example, during segmented index recognition processing, input samples can be classified into at least two categories. For instance, input samples can be classified into three categories: a first category, a second category, and a third category. For example, the first category could represent samples (with values) less than LmcsPivot[lmcs_min_bin_idx+1], the second category could represent samples (with values) greater than or equal to LmcsPivot[LmcsMaxBinIdx], and the third category could indicate samples (with values) between LmcsPivot[lmcs_min_bin_idx+1] and LmcsPivot[LmcsMaxBinIdx].

[0472] In this implementation, an optimization of the recognition process is proposed by eliminating category classification. This is because the input to the segmented index recognition process is the brightness value in the integer (mapping) domain, and no value should exceed the mapping value at the pivot point lmcs_min_bin_idx and LmcsMaxBinIdx+1. Therefore, the conditional processing of classifying samples into categories in existing segmented index recognition processes is unnecessary. Specific examples are described in tables below for further details.

[0473] In the example according to this embodiment, the identification process included in Table 50 can be replaced by one of Tables 51 or 52 below. Referring to Tables 51 and 52, the first two categories of Table 50 can be removed, and for the last category, the boundary value (second boundary value or endpoint) in the iterative for loop is changed from LmcsMaxBinIdx to LmcsMaxBinIdx+1. That is, the identification process can be simplified, and the complexity of segmented index derivation can be reduced. Therefore, LMCS-related coding can be performed efficiently according to this embodiment.

[0474] [Table 51]

[0475]

[0476] [Table 52]

[0477]

[0478] Referring to Table 51, comparisons corresponding to the conditions of the if statement (expressions corresponding to the conditions of the if statement) can be iteratively performed on all warehouse indices from the minimum warehouse index to the maximum warehouse index. If the expression corresponding to the conditions of the if statement is true, the warehouse index can be derived as the inverse mapping index of the inverse luminance mapping (or the inverse scaling index of the chroma residual scaling). Based on the inverse mapping index, the modified reconstructed luminance sample (or the scaled chroma residual sample) can be derived.

[0479] According to the implementation in Table 51, problems that may arise during the identification of (reverse) piecewise function indices can be resolved. Due to the implementation in Table 51, arithmetic loopholes can be eliminated and / or duplicate operations due to overlapping boundary conditions can be omitted. If this document does not follow the implementation in Table 51, the mapping value used in the LMCS of the current block in the for syntax (loop syntax) used to identify the piecewise index (reverse mapping index) may exceed (deviate from) LmcsPivot[idxYInv+1]. According to this implementation, mapping values ​​within an appropriate range used in the LMCS of the current block during index identification can be used.

[0480] The following table shows the impact of the improved coding performance achieved through the implementation method described in Table 51.

[0481] [Table 53]

[0482]

[0483]

[0484] [Table 54]

[0485]

[0486] [Table 55]

[0487]

[0488] Referring to Tables 53 to 55, coding performance can be improved according to the embodiments in Table 51. Furthermore, the embodiments in Table 51 can remove arithmetic loopholes while maintaining coding performance, and / or overlapped arithmetic operations can be omitted.

[0489] In the embodiments according to this document, the latency is higher than that of slices encoded by a single block tree (e.g., intra-frame slices) in existing embodiments due to the chroma residual scaling dependency in the corresponding luma block. Therefore, in existing embodiments, LMCS chroma residual scaling is not applied to slices encoded by a single block tree.

[0490] In this embodiment, chroma residual scaling can be applied even to slices encoded by a single tree. When using a single chroma residual scaling factor as described above, there is no dependency between chroma residual scaling in the corresponding luma blocks, and therefore there may be no delay in the application of chroma residual scaling.

[0491] The table below shows the syntax and semantics of the slice header according to this embodiment.

[0492] [Table 56]

[0493]

[0494] [Table 57]

[0495]

[0496] Referring to Table 57, regardless of whether the conditional clause (or its flag) indicates that the current block has a two-tree structure or a single-tree structure, the chroma residual scaling flag can be signaled.

[0497] In the embodiments according to this document, ALF data and / or LMCS data can be signaled in the APS. For example, 32 APS can be used. In the example, with all APS used for ALF and / or LMCS, the buffer of the APS may require approximately 10KB of on-chip memory. To limit (reduce) the memory required to store ALF / LMCS parameters and the computational complexity required for LMCS, a scheme to limit the number of ALF and / or LMCS APS is proposed in this embodiment.

[0498] In the example according to this embodiment, regardless of the number of slices or tiles included in the image, one LMCS model can be used (allowed) for one image. The number of APS used for the LMCS can be less than 32. For example, the number of APS used for the LMCS can be 4.

[0499] The following table shows the semantics related to APS according to this embodiment.

[0500] [Table 58]

[0501]

[0502] Referring to Table 58, the maximum number of LMCS APSs can be predetermined. For example, the maximum number of LMCS APSs can be 4. Referring to Tables 15 and 58, multiple APSs can include LMCS APSs. The (maximum) number of LMCS APSs can be 4. In the example, the LMCS data field included in one of the LMCSAPSs is available in the LMCS process of the current block in the current screen.

[0503] The table below shows the semantics of the syntactic elements included in the slice header (or screen header).

[0504] [Table 59]

[0505]

[0506] Refer to Table 56 for an explanation of the syntax elements described in Table 59. In the example, the syntax element slice_lmcs_aps_id can be included in the slice header. In another example, the syntax element slice_lmcs_aps_id from Table 59 can be included in the frame header, and in this case, slice_lmcs_aps_id can be modified to ph_lmcs_aps_id.

[0507] According to the embodiments described in this document, the total number of LMCS APS can be equal to or less than 4. Furthermore, in the example of this embodiment, only one LMCS model can be used for one frame. In this case, regardless of the image / video resolution, only one LMCS model can be used for one frame. According to this embodiment, due to the limitation on LMCS APS, it is easy to implement and the problem of resource (memory) overspending can be solved.

[0508] The table below shows the semantics of examples of the syntactic elements disclosed in this document.

[0509] [Table 60]

[0510]

[0511] The syntax elements described in Table 60 can be further explained with reference to Table 56. In the example, the syntax element slice_lmcs_aps_id can be included in the slice header. In another example, the syntax element slice_lmcs_aps_id from Table 60 can be included in the frame header, and in this case, slice_lmcs_aps_id can be modified to ph_lmcs_aps_id.

[0512] Referring to Table 60, the adaptation_parameter_set_id of the LMCS APS referenced by the current slice (or current frame) can be identified by the syntax element slice_lmcs_aps_id. That is, slice_lmcs_aps_id represents the adaptation_parameter_set_id of the LMCS APS referenced by the current slice (or current frame). The value of slice_lmcs_aps_id can be in the range of 0 to 3. Specifically, the value of slice_lmcs_aps_id can be 0, 1, 2, or 3. For example, the 0th LMCS APS (or the first LMCSAPS) indicated by slice_lmcs_aps_id in the 0th to third LMCS APS (or the first to fourth LMCS APS) can be referenced by the current slice (or the current frame), and the LMCS data (LMCS data field or shaper model) included in the 0th LMCS APS (or the first LMCS APS) can be used for the LMCS procedure of the current slice (or the current frame) (the LMCS procedure of the current slice (or the current frame) can be performed based on the LMCS data (LMCS data field or shaper model) included in the 0th LMCS APS (or the first LMCS APS)).

[0513] According to at least one embodiment disclosed in this document, the LMCS-related coding process can be cleaned up or simplified.

[0514] The following figures were created to illustrate specific examples of this specification. Since the names of specific devices or signals / messages / fields described in the figures are presented as examples, the technical features of this specification are not limited to the specific names used in the following figures.

[0515] Figure 16 and Figure 17 Examples of video / image coding methods and related components according to embodiments of this document are illustrated schematically. Figure 16 The method disclosed in the article can be derived from Figure 2 The encoding device disclosed in the document executes the code. Specifically, for example, Figure 16S1600 can be executed by the predictor 220 of the encoding device, S1610 can be executed by the residual processor 230, predictor 220 and / or adder 250 of the encoding device, S1620 can be executed by the residual processor 230 or predictor 220 of the encoding device, S1630, S1640 and / or S1650 can be executed by the residual processor 230 of the encoding device, and S1660 can be executed by the entropy encoder 240 of the encoding device. Figure 16 The methods disclosed herein may include the embodiments described in detail above.

[0516] Reference Figure 16 The encoding device can generate predicted brightness samples (S1600). Regarding the predicted brightness samples, the encoding device can derive the predicted brightness samples for the current block based on the prediction mode. In this case, various prediction methods disclosed in this document, such as inter-frame prediction or intra-frame prediction, can be applied.

[0517] The encoding device can derive predicted chroma samples. The encoding device can derive residual chroma samples based on the original chroma samples and predicted chroma samples of the current block. For example, the encoding device can derive residual chroma samples based on the difference between predicted chroma samples and original chroma samples.

[0518] The encoding device can derive LMCS codewords (S1610). The encoding device can derive individual LMCS codewords for multiple binaries. For example, the above lmcsCW[i] can correspond to the LMCS codewords derived by the encoding device.

[0519] The encoding device can generate a predicted luminance sample of the mapping (S1620). The encoding device can generate the predicted luminance sample of the mapping based on LMCS codewords. For example, the encoding device can derive the input value and mapping value (output value) of the pivot point for luminance mapping, and can generate a predicted luminance sample of the mapping based on the input value and mapping value. Here, the input value and mapping value can be derived based on LMCS codewords. In the example, the encoding device can derive the mapping index idxY based on the first predicted luminance sample, and can generate a predicted luminance sample of the first mapping based on the input value and mapping value of the pivot point corresponding to the mapping index. In another example, a linear mapping (linear integer or linear LMCS) can be used, and the predicted luminance sample of the mapping can be generated based on the forward mapping scaling factor derived from the two pivot points in the linear mapping. Therefore, due to the linear mapping, the index derivation process can be omitted.

[0520] The encoding device can generate scaled residual chroma samples. Specifically, the encoding device can derive a chroma residual scaling factor and generate scaled residual chroma samples based on the chroma residual scaling factor. Here, the encoding-level chroma residual scaling can be referred to as forward chroma residual scaling. Therefore, the chroma residual scaling factor derived by the encoding device can be referred to as the forward chroma residual scaling factor, and forward-scaled residual chroma samples can be generated.

[0521] The encoding device can derive LMCS-related information based on LMCS codewords (S1630). Furthermore, the encoding device can generate LMCS-related information based on the mapped predicted luminance samples and / or scaled residual chrominance samples. The encoding device can generate LMCS-related information for the reconstructed samples. The encoding device can derive LMCS-related parameters for filtering applicable to the reconstructed samples and can generate LMCS-related information based on these parameters. For example, LMCS-related information may include information about the aforementioned luminance mapping (e.g., forward mapping, backward mapping, and linear mapping), information about chrominance residual scaling, and / or indices related to LMCS (or shaping or shaper) (e.g., maximum bin index and minimum bin index).

[0522] The encoding device can generate residual luminance samples based on the mapped predicted luminance samples (S1640). For example, the encoding device can derive the residual luminance samples based on the difference between the mapped predicted luminance samples and the original luminance samples.

[0523] The encoding device can derive residual information (S1650). The encoding device can derive residual information based on scaled residual chroma samples and / or residual luminance samples. The encoding device can derive transform coefficients based on a transform process applied to the scaled residual chroma and luminance residual samples. For example, the transform process can include at least one of DCT, DST, GBT, or CNT. The encoding device can derive quantized transform coefficients based on a quantization process applied to the transform coefficients. The quantized transform coefficients can have a one-dimensional vector form based on the coefficient scan order. The encoding device can generate residual information for the specified quantized transform coefficients. The residual information can be generated using various encoding methods such as exponential Golomb, CAVLC, CABAC, etc.

[0524] The encoding device can encode image / video information (S1660). The image information may include LMCS-related information and / or residual information. For example, LMCS-related information may include information about linear LMCS. In one example, at least one LMCS codeword can be derived based on information about linear LMCS. The encoded video / image information can be output in the form of a bitstream. The bitstream can be transmitted to the decoding device via a network or storage medium.

[0525] According to embodiments of this document, image / video information may include various types of information. For example, image / video information may include information disclosed in at least one of Tables 1 to 60 above.

[0526] In this implementation, the image information may include an APS (Application Profiler). Each APS may include type information (e.g., aps_params_type) indicating whether the APS is an LMCS APS. The LMCS data field included in the APS indicating that the type information is an LMCS APS (e.g., the first LMCSAPS) may include LMCS-related information. The LMCS data field may include information about the LMCS codeword. In this example, the ID value of the APS indicating that the type information is an LMCS APS may be within a predetermined range. For example, the predetermined range may be from 0 to 3. That is, the ID value of the APS indicating that the type information is an LMCS APS may be 0, 1, 2, or 3. In another example, the maximum number of LMCS APSs in the APS may be a predetermined value. For example, the maximum number of LMCS APSs (the predetermined value) may be 4.

[0527] In an implementation, based on a value of 1 for the type information, the APS may include an LMCS data field, which includes LMCS parameters.

[0528] In this implementation, the image information may include header information. The header information may include LMCS-related APS ID information. The LMCS-related APS ID information may represent the ID of the LMCS APS for the current frame or current block. For example, the header information may be a frame header (or slice header). In this example, the value of the LMCS-related APS ID information may be within a predetermined range related to the ID value of the APS representing the type information as an LMCS APS. For example, the value of the LMCS-related APS ID information may be in the range of 0 to 3. That is, the value of the LMCS-related APS ID information may be 0, 1, 2, or 3.

[0529] In this implementation, the image information may include an SPS (Special Status Program). The SPS may include a first LMCS availability flag indicating whether an LMCS data field is available.

[0530] In one implementation, based on the value of the first LMCS availability flag being 1, the header information may include a second LMCS availability flag indicating whether the LMCS data field is available in the frame.

[0531] In this implementation, based on the value of the second LMCS available flag being 1, the LMCS-related APS ID information can be included in the header information.

[0532] In the implementation, when the current block has a single-tree structure or a dual-tree structure (when the current block has a separate tree structure, or when the current block is encoded by a separate tree), the encoding device can generate a chroma residual scaling availability flag indicating whether chroma residual scaling is applied to the current block. When chroma residual scaling is applied to the current frame, the current slice, and / or the current block, the value of the chroma residual scaling availability flag can be 1.

[0533] In implementation, the minimum bin index (e.g., lmcs_min_bin_idx) and / or the maximum bin index (e.g., LmcsMaxBinIdx) can be derived based on LMCS-related information. Based on the minimum bin index, a first mapping value LmcsPivot[lmcs_min_bin_idx] can be derived. Based on the maximum bin index, a second mapping value LmcsPivot[LmcsMaxBinIdx] or LmcsPivot[LmcsMaxBinIdx+1] can be derived. The values ​​of the reconstructed luminance samples (e.g., lumaSample from Table 51 or Table 52) can be within the range of the first mapping value to the second mapping value. In one example, the values ​​of all reconstructed luminance samples are within the range of the first mapping value to the second mapping value. In another example, some sample values ​​of the reconstructed luminance samples are within the range of the first mapping value to the second mapping value.

[0534] In this implementation, the image information may include a sequence parameter set (SPS). The SPS may include a linear LMCS availability flag indicating whether a linear LMCS is available.

[0535] In this implementation, the encoding device can generate a segmented index for chroma residual scaling. The encoding device can derive a chroma residual scaling factor based on the segmented index. The encoding device can generate scaled residual chroma samples based on the residual chroma samples and the chroma residual scaling factor.

[0536] In an implementation, the chromaticity residual scaling factor can be a single chromaticity residual scaling factor.

[0537] In this implementation, LMCS-related information may include an LMCS data field and information about linear LMCS. The information about linear LMCS may be referred to as information about linear mapping. The LMCS data field may include a linear LMCS flag indicating whether linear LMCS is applied. If the value of the linear LMCS flag is 1, then a predicted brightness sample for the mapping can be generated based on the information about linear LMCS.

[0538] In an implementation, information about the linear LMCS may include information about the first pivot point (e.g., Figure 12 Information about P1 and about the second pivot point (e.g., Figure 12The information from P2). For example, the input value and mapped value of the first pivot point can be the minimum input value and the minimum mapped value, respectively. The input value and mapped value of the second pivot point can be the maximum input value and the maximum mapped value, respectively. The input values ​​between the minimum and maximum input values ​​can be linearly mapped.

[0539] In one implementation, the image information includes information about the maximum input value and information about the maximum mapped value. The maximum input value is equal to the value of the information about the maximum input value (i.e., lmcs_max_input in Table 33). The maximum mapped value is equal to the value of the information about the maximum mapped value (i.e., lmcs_max_mapped in Table 33).

[0540] In one implementation, the information regarding the linear mapping includes information about the input Δ value at the second pivot point (i.e., lmcs_max_input_delta in Table 35) and information about the mapping Δ value at the second pivot point (i.e., lmcs_max_mapped_delta in Table 35). The maximum input value can be derived based on the input Δ value at the second pivot point, and the maximum mapping value can be derived based on the mapping Δ value at the second pivot point.

[0541] In one implementation, the maximum input value and the maximum mapping value can be derived based on at least one formula included in Table 36 above.

[0542] In one implementation, generating a predicted brightness sample for the map includes: deriving a forward mapping scaling factor (i.e., ScaleCoeffSingle) for the predicted brightness sample; and generating the predicted brightness sample for the map based on the forward mapping scaling factor. The forward mapping scaling factor can be a single factor of the predicted brightness sample.

[0543] In one implementation, the forward mapping scaling factor can be derived based on at least one formula included in Tables 36 and / or 38 above.

[0544] In one implementation, the predicted brightness sample of the map can be derived based on at least one formula included in Table 38 above.

[0545] In one implementation, the encoding device may derive a reverse mapping scaling factor (i.e., InvScaleCoeffSingle) for the reconstructed luma sample (i.e., lumaSample). Alternatively, the encoding device may generate a modified reconstructed luma sample (i.e., invSample) based on the reconstructed luma sample and the reverse mapping scaling factor. The reverse mapping scaling factor may be a single factor of the reconstructed luma sample.

[0546] In one implementation, a segmented index derived from reconstructed luminance samples can be used to derive the inverse mapping scaling factor.

[0547] In one embodiment, the segment index may be derived based on Table 51 above. That is, the comparison process (lumaSample < LmcsPivot[idxYInv + 1]) included in Table 51 may be iteratively performed from the segment index as the minimum bin index to the segment index as the maximum bin index.

[0548] In one embodiment, the inverse mapping scaling factor may be derived based on at least one of the equations included in Table 33, Table 34, Table 35, and Table 36 or Equation 11 or Equation 12 above.

[0549] In one embodiment, the modified reconstructed luma sample may be derived based on Equation 20, Equation 21, Table 39, and / or Table 40 above.

[0550] In one embodiment, the LMCS-related information may include information about the number of bins for the predicted luma samples used for derivation of the mapping (i.e., lmcs_num_bins_minus1 in Table 41). For example, the number of pivot points of the luma mapping may be set to be equal to the number of bins. In one example, the encoding device may generate the Δ input value and the Δ mapping value of the pivot point according to the number of bins, respectively. In one example, the input value and the mapping value of the pivot point are derived based on the Δ input value (i.e., lmcs_delta_input_cw[i] in Table 41) and the Δ mapping value (i.e., lmcs_delta_mapped_cw[i] in Table 41), and the mapped predicted luma sample may be generated based on the input value (i.e., LmcsPivot_input[i] in Table 42) and the mapping value (i.e., LmcsPivot_mapped[i] in Table 42).

[0551] In one embodiment, the encoding device may derive the LMCS Δ codeword based on at least one LMCS codeword and the original codeword (OrgCW) included in the LMCS-related information, and may derive the mapped luma prediction sample based on at least one LMCS codeword and the original codeword. In one example, the information about the linear mapping may include the information about the LMCS Δ codeword.

[0552] In one embodiment, at least one LMCS codeword may be derived based on the sum of the LMCS Δ codeword and OrgCW. For example, OrgCW is (1 << BitDepthY) / 16, where BitDepthY represents the luma bit depth. This embodiment may be based on Equation 14.

[0553] In one embodiment, at least one LMCS codeword may be derived based on the sum of the LMCSΔ codeword and OrgCW*(lmcs_max_bin_idx - lmcs_min_bin_idx + 1). For example, lmcs_max_bin_idx and lmcs_min_bin_idx are the maximum bin index and the minimum bin index respectively, and OrgCW may be (1<<BitDepthY) / 16. This embodiment may be based on Equation 15 and Equation 16.

[0554] In one embodiment, at least one LMCS codeword may be a multiple of 2.

[0555] In one embodiment, when the luminance bit depth (BitDepthY) of the reconstructed luminance samples is higher than 10, at least one LMCS codeword may be a multiple of 1<<(BitDepthY - 10).

[0556] In one embodiment, at least one LMCS codeword may be in the range from (OrgCW>>1) to (OrgCW<<1)-1.

[0557] Figure 18 and Figure 19 Schematically shows an example of an image / video decoding method and related components according to an embodiment of this document. Figure 18 The method disclosed in Figure 3 may be executed by the Figure 18 disclosed decoding device. Specifically, for example, Figure 18 S1800 of

[0558] may be executed by the entropy decoder 310 of the decoding device, S1810 may be executed by the predictor 330 of the decoding device, S1820 may be executed by the residual processor 320, the predictor 330 and / or the adder 340 of the decoding device, and S1830 may be executed by the adder 340 of the decoding device. Figure 18 The method disclosed in

[0558] may include the above embodiments in this document.

[0558] Referring to Figure 18 , the decoding device may receive / acquire video / image information (S1800). The video / image information may include LMCS-related information. For example, the LMCS-related information may include information about luminance mapping (i.e., forward mapping, inverse mapping, linear mapping), information about chrominance residual scaling, and / or indices related to LMCS (or shaping, shaper) (i.e., maximum bin index, minimum bin index, mapping index). The decoding device may receive / acquire image / video information through a bitstream.

[0559] According to an embodiment of this document, the image / video information may include various information. For example, the image / video information may include the information disclosed in at least one of Tables 1 to 60 above.

[0560] The decoding device can generate predicted brightness samples (S1810). The decoding device can derive predicted brightness samples of the current block in the current frame based on the prediction mode. In this case, various prediction methods disclosed in this document, such as inter-frame prediction or intra-frame prediction, can be applied.

[0561] Image information may include residual information. The decoding device can generate residual chromaticity samples based on the residual information. Specifically, the decoding device can derive quantized transform coefficients based on the residual information. The quantized transform coefficients may have a one-dimensional vector form based on the coefficient scanning order. The decoding device can derive the transform coefficients based on the dequantization process of the quantized transform coefficients. The decoding device can derive residual chromaticity samples and / or residual luminance samples based on the transform coefficients.

[0562] The decoding device can derive LMCS codewords (S1820). The decoding device can derive LMCS codewords based on LMCS-related information. The decoding device can derive information about multiple modules and information about LMCS codewords based on LMCS-related information. The decoding device can derive individual LMCS codewords for multiple modules. For example, the above lmcsCW[i] can correspond to a codeword derived based on LMCS-related information.

[0563] The decoding device can generate a predicted luminance sample of the mapping (S1830). The decoding device can generate the predicted luminance sample of the mapping based on the predicted luminance sample and the LMCS codeword. For example, the decoding device can derive the input value and mapping value (output value) of the pivot point of the luminance mapping, and can generate the predicted luminance sample of the mapping based on the input value and mapping value. In one example, the decoding device can derive the (forward) mapping index (idxY) based on the first predicted luminance sample, and can generate the predicted luminance sample of the first mapping based on the input value and mapping value of the pivot point corresponding to the mapping index. In other examples, a linear mapping (linear integer, linear LMCS) can be used, and the predicted luminance sample of the mapping can be generated based on the forward mapping scaling factor derived from the two pivot points in the linear mapping; therefore, due to the linear mapping, the index derivation process can be omitted.

[0564] The decoding device can generate reconstructed luminance samples (S1840). The decoding device can generate reconstructed luminance samples for the current block based on the mapped predicted luminance samples. Specifically, the decoding device can sum the residual luminance samples with the mapped predicted luminance samples, and can generate reconstructed luminance samples based on the summation result.

[0565] The decoding device can generate scaled residual chroma samples. Specifically, the decoding device can derive a chroma residual scaling factor and generate scaled residual chroma samples based on the chroma residual scaling factor. Here, the chroma residual scaling on the decoding side, in contrast to the encoding side, can be referred to as inverse chroma residual scaling. Therefore, the chroma residual scaling factor derived by the decoding device can be referred to as the inverse chroma residual scaling factor, and inversely scaled residual chroma samples can be generated.

[0566] The decoding device can generate reconstructed chromaticity samples. The decoding device can generate reconstructed chromaticity samples based on scaled residual chromaticity samples. Specifically, the decoding device can perform prediction processing on the chromaticity components and generate predicted chromaticity samples. The decoding device can generate reconstructed chromaticity samples based on the sum of the predicted chromaticity samples and the scaled residual chromaticity samples.

[0567] In this implementation, the image information may include an Image Processing Position (APS). Each APS may include type information indicating whether the APS is an LMCS APS. The LMCS data field included in the APS indicating that the type information is an LMCS APS may include LMCS-related information. The LMCS codeword can be derived based on the LMCS data field. In an example, the ID value of the APS indicating that the type information is an LMCS APS may be within a predetermined range. For example, the predetermined range may be from 0 to 3. That is, the ID value of the APS indicating that the type information is an LMCS APS may be 0, 1, 2, or 3. In another example, the maximum number of LMCS APSs in the APS may be a predetermined value. For example, the maximum number of LMCS APSs (the predetermined value) may be 4.

[0568] In an implementation, based on a value of 1 for the type information, the APS may include an LMCS data field, which includes LMCS parameters.

[0569] In this implementation, the image information may include header information. The header information may include LMCS-related APS ID information. The LMCS-related APS ID information may represent the ID of the LMCS APS for the current frame or current block. For example, the header information may be a frame header (or slice header). In this example, the value of the LMCS-related APS ID information may be within a predetermined range related to the ID value of the APS representing the type information as an LMCS APS. For example, the value of the LMCS-related APS ID information may be in the range of 0 to 3. That is, the value of the LMCS-related APS ID information may be 0, 1, 2, or 3.

[0570] In this implementation, the image information may include an SPS (Special Status Program). The SPS may include a first LMCS availability flag indicating whether an LMCS data field is available.

[0571] In one implementation, based on the value of the first LMCS availability flag being 1, the header information may include a second LMCS availability flag indicating whether the LMCS data field is available in the frame.

[0572] In this implementation, based on the value of the second LMCS available flag being 1, the LMCS-related APS ID information can be included in the header information.

[0573] In the implementation, when the current block has a single-tree structure or a dual-tree structure (when the current block has a separate tree structure, or when the current block is encoded by a separate tree), a signal can be used to indicate whether chroma residual scaling is applied to the current block, indicating whether chroma residual scaling is applicable. When chroma residual scaling is applied to the current frame, the current slice, and / or the current block, the value of the chroma residual scaling applicable flag can be 1.

[0574] In implementation, the minimum bin index (e.g., lmcs_min_bin_idx) and / or the maximum bin index (e.g., LmcsMaxBinIdx) can be derived based on LMCS-related information. Based on the minimum bin index, a first mapping value LmcsPivot[lmcs_min_bin_idx] can be derived. Based on the maximum bin index, a second mapping value LmcsPivot[LmcsMaxBinIdx] or LmcsPivot[LmcsMaxBinIdx+1] can be derived. The values ​​of the reconstructed luminance samples (e.g., lumaSample from Table 51 or Table 52) can be within the range of the first mapping value to the second mapping value. In one example, the values ​​of all reconstructed luminance samples are within the range of the first mapping value to the second mapping value. In another example, some sample values ​​of the reconstructed luminance samples are within the range of the first mapping value to the second mapping value.

[0575] In this implementation, the image information may include a sequence parameter set (SPS). The SPS may include a linear LMCS availability flag indicating whether a linear LMCS is available.

[0576] In this implementation, the segment index (e.g., idxYInv of Tables 35, 36, or 37) can be identified based on LMCS-related information. The decoding device can derive the chroma residual scaling factor based on the segment index. The decoding device can generate scaled residual chroma samples based on the residual chroma samples and the chroma residual scaling factor.

[0577] In an implementation, the chromaticity residual scaling factor can be a single chromaticity residual scaling factor.

[0578] In this implementation, LMCS-related information may include an LMCS data field and information about linear LMCS. The information about linear LMCS may be referred to as information about linear mapping. The LMCS data field may include a linear LMCS flag indicating whether linear LMCS is applied. If the value of the linear LMCS flag is 1, then a predicted brightness sample for the mapping can be generated based on the information about linear LMCS.

[0579] In an implementation, information about the linear LMCS may include information about the first pivot point (e.g., Figure 12 Information about P1 and about the second pivot point (e.g., Figure 12 The information from P2). For example, the input value and mapped value of the first pivot point can be the minimum input value and the minimum mapped value, respectively. The input value and mapped value of the second pivot point can be the maximum input value and the maximum mapped value, respectively. The input values ​​between the minimum and maximum input values ​​can be linearly mapped.

[0580] In this implementation, the image information may include information about the maximum input value and information about the maximum mapped value. The maximum input value may be equal to the value of the information about the maximum input value (e.g., lmcs_max_input in Table 33). The maximum mapped value may be equal to the value of the information about the maximum mapped value (e.g., lmcs_max_mapped in Table 33).

[0581] In an implementation, information about the linear mapping may include information about the input Δ value at the second pivot point (e.g., lmcs_max_input_delta in Table 35) and information about the mapping Δ value at the second pivot point (e.g., lmcs_max_mapped_delta in Table 35). The maximum input value can be derived based on the input Δ value at the second pivot point, and the maximum mapping value can be derived based on the mapping Δ value at the second pivot point.

[0582] In an implementation, the maximum input value and the maximum mapping value can be derived based on at least one formula included in Table 36 as described above.

[0583] In one implementation, generating a predicted brightness sample for the map includes: deriving a forward mapping scaling factor (i.e., ScaleCoeffSingle) for the predicted brightness sample; and generating the predicted brightness sample for the map based on the forward mapping scaling factor. The forward mapping scaling factor can be a single factor of the predicted brightness sample.

[0584] In one implementation, a segmented index derived from reconstructed luminance samples can be used to derive the inverse mapping scaling factor.

[0585] In one implementation, the segmented index can be derived based on Table 51 above. That is, the comparison process (lumaSample) included in Table 51 can be performed iteratively from the segmented index as the minimum warehouse index to the segmented index as the maximum warehouse index. <LmcsPivot[idxYInv+1])。

[0586] In one implementation, the forward mapping scaling factor may be derived based on at least one formula included in Tables 36 and / or 38 above.

[0587] In one implementation, the predicted brightness samples of the mapping can be derived based on at least one formula included in Table 38 above.

[0588] In one implementation, the decoding device may derive an inverse mapping scaling factor (i.e., InvScaleCoeffSingle) for the reconstructed luminance sample (i.e., lumaSample). Alternatively, the decoding device may generate a modified reconstructed luminance sample (i.e., invSample) based on the reconstructed luminance sample and the inverse mapping scaling factor. The inverse mapping scaling factor may be a single factor of the reconstructed luminance sample.

[0589] In one implementation, the reverse mapping scaling factor may be derived based on at least one of the formulas included in Tables 33, 34, 35 and 36 above, or Formula 11 or Formula 12.

[0590] In one implementation, the modified reconstructed luminance sample can be derived based on Equations 20, 21, Table 39 and / or Table 40 above.

[0591] In one implementation, LMCS-related information may include information about the number of bins used to derive the predicted luminance samples for mapping (i.e., lmcs_num_bins_minus1 in Table 41). For example, the number of pivot points for the luminance mapping may be set to be equal to the number of bins. In one example, the decoding device may generate the Δinput value and Δmapped value of the pivot point based on the number of bins, respectively. In one example, the input value and mapped value of the pivot point are derived based on the Δinput value (i.e., lmcs_delta_input_cw[i] in Table 41) and the Δmapped value (i.e., lmcs_delta_mapped_cw[i] in Table 41), and the predicted luminance samples for the mapping may be generated based on the input value (i.e., LmcsPivot_input[i] in Table 42) and the mapped value (i.e., LmcsPivot_mapped[i] in Table 42).

[0592] In one embodiment, the decoding device may derive an LMCSΔ codeword based on at least one LMCS codeword and an original codeword (OrgCW) included in the LMCS-related information, and may derive a mapped luminance prediction sample based on at least one LMCS codeword and the original codeword. In one example, the information about the linear mapping may include the information about the LMCSΔ codeword.

[0593] In one embodiment, at least one LMCS codeword may be derived based on the sum of the LMCSΔ codeword and OrgCW. For example, OrgCW is (1<<BitDepthY) / 16, where BitDepthY represents the luminance bit depth. This embodiment may be based on Equation 14.

[0594] In one embodiment, at least one LMCS codeword may be derived based on the sum of the LMCSΔ codeword and OrgCW*(lmcs_max_bin_idx - lmcs_min_bin_idx + 1). For example, lmcs_max_bin_idx and lmcs_min_bin_idx are the maximum bin index and the minimum bin index respectively, and OrgCW may be (1<<BitDepthY) / 16. This embodiment may be based on Equation 15 and Equation 16.

[0595] In one embodiment, at least one LMCS codeword may be a multiple of 2.

[0596] In one embodiment, when the luminance bit depth (BitDepthY) of the reconstructed luminance sample is higher than 10, at least one LMCS codeword may be a multiple of 1<<(BitDepthY - 10).

[0597] In one embodiment, at least one LMCS codeword may be in the range from (OrgCW>>1) to (OrgCW<<1)-1.

[0598] In the above embodiments, the method is described based on a flowchart having a series of steps or blocks. This document is not limited to the order of the above steps or blocks. Some steps or blocks may occur simultaneously or in a different order from other steps or blocks as described above. In addition, those skilled in the art will understand that the steps shown in the above flowchart are not exclusive, may include additional steps, or one or more steps in the flowchart may be deleted without affecting the scope of this document.

[0599] The method according to the above embodiments of this document may be implemented in software form, and the encoding device and / or decoding device according to this document may be included, for example, in a device that performs image processing of a TV, a computer, a smart phone, a set-top box, a display device, etc.

[0600] When the embodiments described in this document are implemented in software, the above methods can be implemented as modules (processes, functions, etc.) that perform the above functions. Modules can be stored in memory and executed by a processor. The memory can be located internally or externally to the processor and can be connected to the processor by various well-known means. The processor may include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this document can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in the various figures can be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information about the instructions or algorithms used for implementation can be stored in a digital storage medium.

[0601] Furthermore, the decoding and encoding devices employing this document can be included in multimedia broadcasting transmitting / receiving devices, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, real-time communication devices (e.g., video communication), mobile streaming devices, storage media, cameras, VoD service providers, over-the-top (OTT) video devices, internet streaming service providers, three-dimensional (3D) video devices, teleconferencing video devices, transportation user equipment (i.e., vehicle user equipment, aircraft user equipment, ship user equipment, etc.), and medical video devices, and can be used to process video signals and data signals. For example, over-the-top (OTT) video devices may include game consoles, Blu-ray players, internet access TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.

[0602] Furthermore, the processing method described in this document can be generated in the form of a program to be executed by a computer and can be stored in a computer-readable recording medium. Multimedia data having the data structure according to this document can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices that store data readable by a computer system. For example, computer-readable recording media may include BD, Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. Additionally, computer-readable recording media include media implemented in the form of a carrier wave (i.e., transmission via the Internet). Furthermore, the bitstream generated by this encoding method can be stored in a computer-readable recording medium or transmitted via a wired / wireless communication network.

[0603] Furthermore, the embodiments described in this document can be implemented as a computer program product based on program code, and the program code can be executed in a computer using the embodiments described in this document. The program code can be stored on a computer-readable medium.

[0604] Figure 20 Examples of content streaming systems to which the implementation methods disclosed in this document can be applied are shown.

[0605] Reference Figure 20 The content streaming system that applies the implementation method of this document may mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.

[0606] An encoding server compresses content input from multimedia input devices (e.g., smartphones, cameras, camcorders, etc.) into digital data to generate a bitstream and sends the bitstream to a streaming server. As another example, when the multimedia input device (e.g., smartphone, camera, camcorder, etc.) generates the bitstream directly, the encoding server can be omitted.

[0607] Bitstreams can be generated using the encoding method or bitstream generation method described in this document, and the stream server can temporarily store the bitstreams during the sending or receiving of the bitstreams.

[0608] The streaming server sends multimedia data to the user's device based on the user's request via a web server, with the web server acting as a medium to inform the user of services. When a user requests a desired service from the web server, the web server forwards it to the streaming server, and the streaming server sends the multimedia data to the user. In this scenario, the content streaming system may include a separate control server. In this case, the control server is used to control the commands / responses between devices within the content streaming system.

[0609] A streaming server can receive content from media storage devices and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined period of time.

[0610] Examples of user devices may include mobile phones, smartphones, laptops, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays), digital TVs, desktop computers, digital signage, etc. Individual servers in a content streaming system can operate as distributed servers, in which case data received from each server can be distributed.

[0611] In a content streaming system, each server can operate as a distributed server, and in this case, the data received from each server can be distributed and processed.

[0612] The claims described herein can be combined in various ways. For example, the technical features of the method claims of this document can be combined and implemented as a device, and the technical features of the device claims of this document can be combined and implemented as a method. Furthermore, the technical features of the method claims and the device claims of this document can be combined to implement a device, and the technical features of the method claims and the device claims of this document can be combined and implemented as a method.

Claims

1. An image decoding method performed by a decoding device, the image decoding method comprising the following steps: Image information is obtained from the bitstream, including prediction-related information and luminance mapping and chrominance scaling (LMCS)-related information. Based on the prediction-related information, a predicted brightness sample is generated for the current block in the current image. Derive LMCS codewords based on the aforementioned LMCS-related information; Based on the predicted brightness samples and the LMCS codewords, a mapped predicted brightness sample is generated; and Based on the predicted brightness samples of the mapping, reconstructed brightness samples are generated for the current block. The image information includes an adaptive parameter set (APS). Each of the APSs includes type information specifying whether the APS is an LMCS APS. The type information specifies that the LMCS data fields included in the APS of the LMCS APS include LMCS-related information. Specifically, the LMCS codeword is derived based on the LMCS data field. The type information specifies that the ID value of the APS of the LMCS APS is within a predetermined range, and The predetermined range is 0 to 3.

2. An image encoding method performed by an encoding device, the image encoding method comprising the following steps: Generate a predicted brightness sample for the current block in the current image; Derive the LMCS codewords for luminance mapping and chrominance scaling; Based on the LMCS codeword, LMCS-related information is derived; Based on the LMCS codewords, generate the mapped predicted brightness samples; Based on the predicted brightness samples of the mapping, a residual brightness sample is generated for the current block; Residual information is derived based on the residual brightness samples; as well as The image information, including the LMCS-related information and the residual information, is encoded. The image information includes an adaptive parameter set (APS). Each of the APSs includes type information specifying whether the APS is an LMCS APS. The type information specifies that the LMCS data fields included in the APS of the LMCS APS include LMCS-related information. The LMCS data field includes information about the LMCS codeword. The type information specifies that the ID value of the APS of the LMCS APS is within a predetermined range, and The predetermined range is 0 to 3.

3. A method for transmitting data for image information, the method comprising the following steps: Obtain a bitstream for the image information, wherein the bitstream is generated by performing the following steps: generating a predicted luminance sample for the current block in the current frame, deriving luminance mapping and chroma scaling (LMCS) codewords, and deriving LMCS-related information based on the LMCS codewords, generating a predicted luminance sample of the mapping based on the LMCS codewords, generating a residual luminance sample for the current block based on the predicted luminance sample of the mapping, deriving residual information based on the residual luminance sample, and encoding image information including the LMCS-related information and the residual information; and Send the data including the bit stream. The image information includes an adaptive parameter set (APS). Each of the APSs includes type information specifying whether the APS is an LMCS APS. The type information specifies that the LMCS data fields included in the APS of the LMCS APS include LMCS-related information. The LMCS data field includes information about the LMCS codeword. The type information specifies that the ID value of the APS of the LMCS APS is within a predetermined range, and The predetermined range is 0 to 3.