Decoding device, encoding device, and data transmission device
Through brightness mapping and chromaticity scaling (LMCS) processing, the high cost problems of high resolution and high-quality images/videos are solved, the compression efficiency and visual quality are improved, and the calculation complexity is reduced, and it is suitable for the encoding process of various block tree structures.
Patent Information
- Application Number
- CN202510682224.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-20
- Filing Date
- 2020-06-22
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art has problems with high transmission and storage costs in the compression, transmission and storage of high resolution, high quality images/videos, especially when processing image/videos with different characteristics, compression efficiency and computational complexity are difficult to meet the needs.
The use of brightness mapping and chromaticity scaling (LMCS) processing is adopted to signal a single chroma residual scaling factor and flexible bin usage, combined with linear mapping and a simplified index derivation process, which is suitable for coding tree units with a dual-tree structure, limiting the number of LMCS to improve coding efficiency.
It improves image/video compression efficiency, improves subjective/objective visual quality, reduces computational complexity and resource requirements, and is suitable for the encoding process of various block tree structures.
Smart Images

Figure CN120343272A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the original application number 202080053827.2 (International Application No.: PCT / KR2020 / 008053, filing date: June 22, 2020, invention title: Video or Image Coding Based on Luminance Mapping and Chrominance Scaling). Technical Field
[0002] The technology of this document relates to video or image coding based on luminance mapping and chrominance scaling. Background Art
[0003] Recently, the demand for high-resolution and high-quality images / videos such as 4K or 8K or higher ultra-high definition (UHD) images / videos has increased in various fields. As image / video data has high resolution and quality, the amount of information or bits to be transmitted increases relative to existing image / video data. Therefore, transmitting image data using media such as existing wired / wireless broadband lines or existing storage media or storing image / video data using existing storage media increases the transmission cost and storage cost.
[0004] In addition, the interest and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms have recently increased, and the broadcasting of images / videos (e.g., game images) with characteristics different from real images has increased.
[0005] Therefore, very efficient image / video compression technologies are needed to effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality images / videos with various characteristics as described above.
[0006] In addition, luminance mapping and chrominance scaling (LMCS) processing is performed to improve compression efficiency and increase subjective / objective visual quality, and discussions have been made on reducing the computational complexity in LMCS processing. Summary of the Invention
[0007] Technical Solution
[0008] According to an embodiment of this document, a method and device for increasing image coding efficiency are provided.
[0009] According to an embodiment of this document, an efficient filtering application method and device are provided.
[0010] According to an embodiment of this document, an efficient LMCS application method and device are provided.
[0011] According to an embodiment of this document, LMCS codewords (or their ranges) can be constrained.
[0012] According to an embodiment of this document, a single chrominance residual scaling factor signaled directly in the chrominance scaling of LMCS can be used.
[0013] According to an embodiment of this document, a linear mapping (linear LMCS) can be used.
[0014] According to an embodiment of this document, information about pivot points required for the linear mapping can be signaled explicitly.
[0015] According to an embodiment of this document, a flexible number of bins can be used for luminance mapping.
[0016] According to an embodiment of this document, the index derivation process for inverse luminance mapping and / or chrominance residual scaling can be simplified.
[0017] According to an embodiment of this document, the LMCS process can be applied even when the luminance block and the chrominance block in a coding tree unit (CTU) have separate block tree structures (dual-tree structures).
[0018] According to an embodiment of this document, the number of LMCS APSs can be restricted.
[0019] According to an embodiment of this document, a video / image decoding method performed by a decoding device is provided.
[0020] According to an embodiment of this document, a decoding device for performing video / image decoding is provided.
[0021] According to an embodiment of this document, a video / image encoding method performed by an encoding device is provided.
[0022] According to an embodiment of this document, an encoding device for performing video / image encoding is provided.
[0023] According to an embodiment of this document, a computer-readable digital storage medium is provided, in which encoded video / image information generated according to the video / image encoding method disclosed in at least one embodiment of this document is stored.
[0024] According to an embodiment of this document, a computer-readable digital storage medium is provided, in which encoded information or encoded video / image information that causes a decoding device to perform the video / image decoding method disclosed in at least one embodiment of this document is stored.
[0025] Advantageous Effects
[0026] According to an embodiment of this document, the overall image / video compression efficiency can be improved.
[0027] According to an embodiment of this document, subjective / objective visual quality can be improved through efficient filtering.
[0028] According to an embodiment of this document, the LMCS process for image / video coding can be efficiently performed.
[0029] According to an embodiment of this document, the resources / costs (software or hardware) required for the LMCS process can be minimized.
[0030] According to an embodiment of this document, the hardware implementation of the LMCS process can be facilitated.
[0031] According to an embodiment of this document, the division operations required for the derivation of LMCS codewords in mapping (shaping) can be removed or minimized by constraining the LMCS codewords (or their ranges).
[0032] According to an embodiment of this document, a single chrominance residual scaling factor can be used to remove the latency identified according to the segment index.
[0033] According to an embodiment of this document, linear mapping in LMCS can be used to perform chrominance residual scaling processing without relying on the reconstruction of the luminance block, so the latency in scaling can be removed.
[0034] According to an embodiment of this document, the mapping efficiency in LMCS can be increased.
[0035] According to an embodiment of this document, by simplifying the index derivation process for inverse luminance mapping and / or chrominance residual scaling, the complexity of LMCS can be reduced, and thus the video / image coding efficiency can be increased.
[0036] According to an embodiment of this document, the LMCS process can even be performed on blocks with a dual-tree structure, so the efficiency of LMCS can be increased. In addition, the coding performance (e.g., objective / subjective image quality) of blocks with a dual-tree structure can be improved.
[0037] According to an embodiment of this document, since the number of LMCS APSs is limited, the complexity of LMCS can be reduced, and LMCS can consume (use) fewer resources (e.g., memory). Description of the Drawings
[0038] Figure 1 An example of a video / image coding system to which the embodiments of this document can be applied is shown.
[0039] Figure 2 It is a diagram schematically showing the configuration of a video / image coding device to which the embodiments of this document can be applied.
[0040] Figure 3It is a diagram schematically showing the configuration of a video / image decoding device to which the embodiments of this document can be applied.
[0041] Figure 4 An exemplary block tree structure is shown.
[0042] Figure 5 Exemplarily represents the hierarchical structure of an encoded image / video.
[0043] Figure 6 Exemplarily shows the hierarchical structure of CVS according to the embodiments of this document.
[0044] Figure 7 Exemplarily shows the hierarchical structure of CVS according to the embodiments of this document.
[0045] Figure 8 Exemplarily shows the hierarchical structure of CVS according to another embodiment of this document.
[0046] Figure 9 Shows an exemplary LMCS structure according to the embodiments of this document.
[0047] Figure 10 Shows an LMCS structure according to another embodiment of this document.
[0048] Figure 11 Shows a graph representing an exemplary forward mapping.
[0049] Figure 12 It is a flowchart showing a method for deriving a chrominance residual scaling index according to the embodiments of this document.
[0050] Figure 13 Shows the linear fitting of a pivot point according to the embodiments of this document.
[0051] Figure 14 Shows an example of linear shaping (or linear shaping, linear mapping) according to the embodiments of this document.
[0052] Figure 15 Shows an example of linear forward mapping in the embodiments of this document.
[0053] Figure 16 Shows an example of inverse forward mapping in the embodiments of this document.
[0054] Figure 17 and Figure 18 Schematically shows an example of a video / image encoding method and related components according to the embodiments of this document.
[0055] Figure 19 and Figure 20An example of an image / video decoding method and related components according to an embodiment of this document is schematically shown.
[0056] Figure 21 An example of a content stream system to which the embodiments disclosed in this document can be applied is shown. Detailed Embodiments
[0057] This document can be modified in various forms, and its specific embodiments will be described and shown in the drawings. However, these embodiments are not intended to limit this document. The terms used in the following description are only for describing specific embodiments and are not intended to limit this document. Singular expressions include plural expressions as long as they are clearly different in reading. Terms such as "including" and "having" are intended to indicate the presence of the features, quantities, steps, operations, elements, components, or combinations thereof used in the following description, and thus it should be understood that the possibility of the presence or addition of one or more different features, quantities, steps, operations, elements, components, or combinations thereof is not excluded.
[0058] In addition, the respective configurations in the drawings described in this document are shown independently for convenience in describing different characteristic functions, and it does not mean that each configuration is implemented as a separate hardware or separate software. For example, two or more of the respective components can be combined to form one component, or one component can be divided into multiple components. Embodiments in which the respective components are integrated and / or separated are also included in the scope of the disclosure of this document.
[0059] Hereinafter, examples of this embodiment will be described in detail with reference to the drawings. In addition, similar reference numerals are used throughout the drawings to indicate similar elements, and the same description of similar elements will be omitted.
[0060] Figure 1 An example of a video / image encoding system to which the embodiments of this document can be applied is shown.
[0061] Refer to Figure 1 , the video / image encoding system may include a first device (source device) and a second device (receiving device). The source device may send the encoded video / image information or data in the form of a file or a stream to the receiving device through a digital storage medium or a network.
[0062] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0063] The video source can obtain video / images through processes such as capturing, synthesizing, or generating video / images. The video source can include a video / image capturing device and / or a video / image generating device. For example, the video / image capturing device can include one or more cameras, a video / image archive including previously captured video / images, etc. For example, the video / image generating device can include a computer, a tablet computer, and a smart phone, and can (electronically) generate video / images. For example, virtual video / images can be generated by a computer, etc. In this case, the video / image capturing process can be replaced by a process of generating relevant data.
[0064] The encoding device can encode the input video / image. For compression and encoding efficiency, the encoding device can perform a series of processes such as prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0065] The transmitter can send the encoded image / image information or data output in the form of a bitstream to the receiver of the receiving device in the form of a file or a stream through a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating a media file in a predetermined file format and can include elements for transmitting through a broadcast / communication network. The receiver can receive / extract the bitstream and send the received bitstream to the decoding device.
[0066] The decoding device can decode the video / image by performing a series of processes such as dequantization, inverse transformation, and prediction corresponding to the operations of the encoding device.
[0067] The renderer can render the decoded video / image. The rendered video / image can be displayed through a display.
[0068] This document relates to video / image encoding. For example, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the General Video Coding (VVC) standard, the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the Second Generation Audio Video Coding Standard (AVS2), or the next-generation video / image coding standards (e.g., H.267, H.268, etc.).
[0069] This document presents various embodiments of video / image encoding, and unless otherwise specified, the above embodiments can also be executed in combination with each other.
[0070] In this document, video may refer to a series of images over time. A picture generally refers to a unit of an image representing a specific time range, and a slice / tile refers to a unit that forms part of a picture in terms of coding. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles. A tile may represent a rectangular area of CTU rows within a tile in a picture. A tile may be divided into multiple tiles, and each tile may be composed of one or more CTU rows within the tile. A tile that is not divided into multiple tiles may also be referred to as a tile. Tile scanning may represent a specific sequential ordering of CTUs that divides a picture, where CTUs may be ordered in a raster scan within a tile, and tiles within a tile of a tile group may be sequentially ordered in a raster scan of the tiles of the tile group, and tiles in a picture may be sequentially ordered in a raster scan of the tiles of the picture. A tile is a rectangular area of CTUs within a specific tile column and a specific tile row in a picture. A tile column is a rectangular area of CTUs with a height equal to the height of the picture and a width specified by a syntax element in the picture parameter set. A tile row is a rectangular area of CTUs with a height specified by a syntax element in the picture parameter set and a width equal to the width of the picture. Tile scanning is a specific sequential ordering of CTUs that divides a picture, where CTUs are sequentially ordered in a raster scan of the CTUs of a tile, and tiles in a picture are sequentially ordered in a raster scan of the tiles of the picture. A slice includes an integral number of tiles of a picture that are exclusively contained within a single NAL unit. A slice may be composed of a continuous sequence of multiple complete tiles or complete tiles of only one tile. In this document, tile groups and slices may be used interchangeably. For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.
[0071] In addition, a picture may be divided into two or more sub-pictures. A sub-picture may be a rectangular area of one or more slices within a picture.
[0072] A pixel or pel may mean the smallest unit that constitutes a picture (or image). Additionally, "sample" may be used as a term corresponding to a pixel. A sample generally may represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0073] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to the area. A unit may include one luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, a unit may be used interchangeably with terms such as a block or a region. In general, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients in M columns and N rows. Alternatively, a sample may mean a pixel value in the spatial domain, and when such a pixel value is transformed to the frequency domain, it may mean a transform coefficient in the frequency domain.
[0074] In this document, "A or B" may mean "only A", "only B", or "both A and B". In other words, "A or B" in this document may be interpreted as "A and / or B". For example, in this document, "A, B, or C (A, B, or C)" means "only A", "only B", "only C", or "any combination of A, B, and C".
[0075] The slashes ( / ) or commas (、) used in this document may mean "and / or". For example, "A / B" may mean "A and / or B". Therefore, "A / B" may mean "only A", "only B", or "both A and B". For example, "A, B, C" may mean "A, B, or C".
[0076] In this document, "at least one of A and B" may mean "only A", "only B", or "both A and B". Additionally, in this document, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted in the same way as "at least one of A and B".
[0077] Additionally, in this document, "at least one of A, B, and C" means "only A", "only B", "only C", or "any combination of A, B, and C". Additionally, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".
[0078] Additionally, the parentheses used in this document may mean "for example". Specifically, when indicating "prediction (intra prediction)", "intra prediction" may be proposed as an example of "prediction". In other words, "prediction" in this document is not limited to "intra prediction", and "intra prediction" may be proposed as an example of "prediction". Additionally, even when indicating "prediction (i.e., intra prediction)", "intra prediction" may be proposed as an example of "prediction".
[0079] In this document, the technical characteristics described separately in a figure may be implemented separately or simultaneously.
[0080] Figure 2FIG. is a diagram schematically showing a configuration of a video / image encoding device to which the present document can be applied. Hereinafter, the so-called video encoding device may include an image encoding device.
[0081] Referring to Figure 2 , the encoding device 200 includes an image splitter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image splitter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or a processor). Additionally, the memory 270 may include a decoded picture buffer (DPB), or may be configured by a digital storage medium. The hardware component may further include the memory 270 as an internal / external component.
[0082] The image splitter 210 may split an input image (or a picture or a frame) input to the encoding device 200 into one or more processors. For example, a processor may be referred to as a coding unit (CU). In this case, the coding unit may be recursively split from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quadtree binary tree ternary tree (QTBTTT) structure. For example, one coding unit may be split into multiple coding units with a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quadtree structure may be applied first, and later the binary tree structure and / or the ternary structure may be applied. Alternatively, the binary tree structure may be applied first. The encoding process according to the present disclosure may be performed based on the final coding unit that is no longer split. In this case, based on the coding efficiency according to the image characteristics, the largest coding unit may be used as the final coding unit, or if necessary, the coding unit may be recursively split into coding units with a deeper depth and the coding unit with an optimal size may be used as the final coding unit. Here, the encoding process may include processes of prediction, transformation, and reconstruction (to be described later). As another example, the processor may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may be split or divided from the above-mentioned final coding unit. The prediction unit may be a unit for sample prediction, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from the transformation coefficients.
[0083] In some cases, a unit may be used interchangeably with terms such as a block or a region. In general, an M×N block may represent a set of samples or transform coefficients consisting of M columns and N rows. Samples may generally represent pixels or pixel values, which may represent only the pixels / pixel values of the luminance component or only the pixels / pixel values of the chrominance component. A sample may be used as a term corresponding to a picture (or image) of pixels or picture elements.
[0084] In the encoding device 200, a prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 is subtracted from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, as shown, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) in the encoding device 200 may be referred to as the subtractor 231. The predictor may perform prediction on a block to be processed (hereinafter referred to as the current block) and generate a prediction block including the prediction samples of the current block. The predictor may determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or the CU. As will be described later in the description of each prediction mode, the predictor may generate various information related to the prediction (e.g., prediction mode information) and send the generated information to the entropy encoder 240. The information about the prediction may be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0085] The intra-frame predictor 222 may predict the current block by referring to the samples in the current picture. Depending on the prediction mode, the samples referred to may be located near the current block or may be separated. In intra-frame prediction, the prediction mode may include a plurality of non-directional modes and a plurality of directional modes. For example, the non-directional modes may include the DC mode and the planar mode. For example, depending on the level of detail of the prediction direction, the directional modes may include 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes may be used according to the settings. The intra-frame predictor 222 may use the prediction mode applied to the neighboring blocks to determine the prediction mode applied to the current block.
[0086] The inter-frame predictor 221 may derive a predicted block of a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. Here, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring blocks may be the same or different. The temporal neighboring blocks may be referred to as co-located reference blocks, co-located CUs (colCUs), etc., and the reference picture including the temporal neighboring blocks may be referred to as a co-located picture (colPic). For example, the inter-frame predictor 221 may configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter-frame prediction may be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter-frame predictor 221 may use the motion information of neighboring blocks as the motion information of the current block. In the skip mode, different from the merge mode, the residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of a neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be indicated by signaling a motion vector difference.
[0087] The predictor 220 may generate a prediction signal based on various prediction methods described below. For example, the predictor may apply not only intra-frame prediction or inter-frame prediction to predict a block, but also both intra-frame prediction and inter-frame prediction at the same time. This may be referred to as combined inter-frame and intra-frame prediction (CIIP). Additionally, the predictor may predict a block based on the intra-block copy (IBC) prediction mode or the palette mode. The IBC prediction mode or the palette mode may be used for content image / video coding such as games, e.g., screen content coding (SCC). IBC basically performs prediction in the current picture, but may be performed similarly to inter-frame prediction such that a reference block is derived in the current picture. That is, IBC may use at least one inter-frame prediction technique described in the present disclosure. The palette mode may be regarded as an example of intra-frame coding or intra-frame prediction. When the palette mode is applied, the sample values within the picture may be signaled based on information regarding the palette table and the palette index.
[0088] The prediction signal generated by a predictor (including the inter-frame predictor 221 and / or the intra-frame predictor 222) can be used to generate a reconstructed signal or to generate a residual signal. The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen - Loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, GBT means a transform obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to a transform generated based on a prediction signal generated using all previously reconstructed pixels. Additionally, the transform process can be applied to square pixel blocks of the same size or can be applied to blocks of variable size other than square.
[0089] The quantizer 233 can quantize the transform coefficients and send them to the entropy encoder 240, and the entropy encoder 240 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 233 can rearrange the block-type quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. Information about the transform coefficients can be generated. The entropy encoder 240 can perform various coding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 240 can encode together or separately the information required for video / image reconstruction other than the quantized transform coefficients (e.g., the values of syntax elements, etc.). The encoded information (e.g., the encoded video / image information) can be sent or stored in the form of a bitstream in units of NAL (network abstraction layer). The video / image information can also include information about various parameter sets, such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information can also include general constraint information. In the present disclosure, the information and / or syntax elements sent / signaled from the encoding device to the decoding device can be included in the video / picture information. The video / image information can be encoded through the above encoding process and included in the bitstream. The bitstream can be sent via a network or can be stored in a digital storage medium. The network can include a broadcast network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for sending the signal output from the entropy encoder 240 and / or a storage unit (not shown) for storing the signal can be included as internal / external elements of the encoding device 200, and alternatively, the transmitter can be included in the entropy encoder 240.
[0090] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transformation to the quantized transform coefficients via the dequantizer 234 and the inverse transform unit 235. The adder 250 adds the reconstructed residual signal and the prediction signal output from the inter-prediction unit 221 or the intra-prediction unit 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). If there is no residual in the block to be processed (e.g., in the case of applying the skip mode), the predicted block can be used as the reconstructed block. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. As described below, the generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture and can be used for inter-prediction of the next picture through filtering.
[0091] In addition, a luminance mapping with chroma scaling (LMCS) can be applied during picture encoding and / or reconstruction.
[0092] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 270 (specifically, the DPB of the memory 270). For example, various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various information related to filtering and send the generated information to the entropy encoder 240, as will be described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoder 240 and output in the form of a bitstream.
[0093] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter-prediction unit 221. When inter-prediction is applied by the encoding device, a prediction mismatch between the encoding device 200 and the decoding device 300 can be avoided and the encoding efficiency can be improved.
[0094] The DPB of the memory 270 can store the modified reconstructed picture used as a reference picture in the inter-prediction unit 221. The memory 270 can store the motion information of the blocks for deriving (or encoding) the motion information in the current picture and / or the motion information of the blocks that have been reconstructed in the picture. The stored motion information can be sent to the inter-prediction unit 221 and used as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 can store the reconstructed samples of the reconstructed blocks in the current picture and can transfer the reconstructed samples to the intra-prediction unit 222.
[0095] Figure 3 is a schematic diagram showing the configuration of a video / image decoding device to which embodiments of the present disclosure can be applied.
[0096] Reference Figure 3 As shown in Figure 3 , the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be configured by hardware components (e.g., a decoder chipset or a processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0097] When the input is a bitstream including video / image information, the decoding device 300 may reconstruct an image corresponding to the processing of the video / image information in the Figure 2 encoding device. For example, the decoding device 300 may derive units / blocks based on block segmentation related information obtained from the bitstream. The decoding device 300 may use the processor applied in the encoding device to perform decoding. Thus, for example, the decoding processor may be an encoding unit, and the encoding unit may be segmented from a coding tree unit or a largest coding unit according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. The reconstructed image signal decoded and output by the decoding device 300 may be reproduced by a reproduction device.
[0098] The decoding device 300 may receive from Figure 2The signal output by the encoding device in the form of a bitstream, and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets, such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may also include general constraint information. The decoding device may also decode the picture based on the information about the parameter set and / or the general constraint information. The information and / or syntax elements signaled / received described later in the present disclosure may be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 310 decodes the information in the bitstream based on an encoding method such as exponential Golomb coding, CAVLC, or CABAC, and outputs the syntax elements required for image reconstruction and the quantization values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method may receive bins corresponding to the respective syntax elements in the bitstream, determine a context model using the information of the syntax element to be decoded, the decoding information of the block to be decoded, or the information of the symbols / bins decoded in the previous stage, and perform arithmetic decoding on the bins by predicting the probability of the bin occurrence according to the determined context model, and generate symbols corresponding to the values of the respective syntax elements. In this case, the CABAC entropy decoding method may update the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin after determining the context model. Among the information decoded by the entropy decoder 310, the information related to prediction may be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values (i.e., the quantized transform coefficients and related parameter information) for which entropy decoding has been performed in the entropy decoder 310 may be input to the residual processor 320. The residual processor 320 may derive a residual signal (residual block, residual sample, residual sample array). Additionally, the information about filtering among the information decoded by the entropy decoder 310 may be provided to the filter 350. Furthermore, a receiver (not shown) for receiving the signal output by the encoding device may also be configured as an internal / external component of the decoding device 300, or the receiver may be a component of the entropy decoder 310. Additionally, the decoding device according to the present disclosure may be referred to as a video / image / picture decoding device, and the decoding device may be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0099] The dequantizer 321 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 may rearrange the quantized transform coefficients in a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scan order executed in the encoding device. The dequantizer 321 may dequantize the quantized transform coefficients using quantization parameters (e.g., quantization step information) and obtain the transform coefficients.
[0100] The inverse transformer 322 inversely transforms the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0101] The predictor may perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 310 and may determine a specific intra / inter prediction mode.
[0102] The predictor 330 may generate a prediction signal based on various prediction methods described below. For example, the predictor may not only apply intra prediction or inter prediction to predict a block, but may also apply intra prediction and inter prediction simultaneously. This may be referred to as combined inter and intra prediction (CIIP). Additionally, the predictor may predict a block based on an intra block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or the palette mode may be used for content image / video coding such as games, e.g., screen content coding (SCC). IBC basically performs prediction in the current picture, but may be performed similarly to inter prediction such that a reference block is derived in the current picture. That is, IBC may use at least one of the inter prediction techniques described in the present disclosure. The palette mode may be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture may be signaled based on the information about the palette table and the palette index.
[0103] The intra predictor 331 may predict the current block by referring to samples in the current picture. Depending on the prediction mode, the samples referred to may be located near the current block or may be separated. In intra prediction, the prediction mode may include a plurality of non-directional modes and a plurality of directional modes. The intra predictor 331 may use the prediction mode applied to neighboring blocks to determine the prediction mode applied to the current block.
[0104] The inter - frame predictor 332 can derive a predicted block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information can be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include inter - frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter - frame prediction, neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter - frame predictor 332 can configure a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter - frame prediction can be performed based on various prediction modes, and the information regarding the prediction can include information indicating the inter - frame prediction mode of the current block.
[0105] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to a prediction signal (predicted block, predicted sample array) output from a predictor (including the inter - frame predictor 332 and / or the intra - frame predictor 331). If there is no residual for the block to be processed, for example, when the skip mode is applied, the predicted block can be used as the reconstructed block.
[0106] The adder 340 can be referred to as a reconstructor or a reconstructed - block generator. The generated reconstructed signal can be used for intra - frame prediction of the next block to be processed in the current picture, can be output through filtering as described below, or can be used for inter - frame prediction of the next picture.
[0107] In addition, luminance mapping with chroma scaling (LMCS) can be applied in the picture decoding process.
[0108] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 360 (specifically, the DPB of the memory 360). For example, various filtering methods can include de - blocking filtering, sample - adaptive offset, adaptive loop filter, bilateral filter, etc.
[0109] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store the motion information of the blocks that derive (or decode) the motion information in the current picture and / or the motion information of the blocks that have been reconstructed in the picture. The stored motion information can be sent to the inter-frame predictor 332 to be used as the motion information of spatially neighboring blocks or temporally neighboring blocks. The memory 360 can store the reconstructed samples of the reconstructed blocks in the current picture and transfer the reconstructed samples to the intra-frame predictor 331.
[0110] In this document, the embodiments described in the filter 260, the inter-frame predictor 221, and the intra-frame predictor 222 of the encoding device 200 can be the same as or respectively correspond to the filter 350, the inter-frame predictor 332, and the intra-frame predictor 331 of the decoding device 300. This also applies to the inter-frame predictor 332 and the intra-frame predictor 331.
[0111] As described above, in video coding, prediction is performed to increase the compression efficiency. Thus, a prediction block including prediction samples of a current block (a block to be encoded) can be generated. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived identically from the encoding device and the decoding device, and the encoding device decodes information about the residual between the original block and the prediction block (residual information) rather than the original sample values of the original block itself. By signaling the device, the image coding efficiency can be increased. The decoding device can derive a residual block including residual samples based on the residual information, and generate a reconstructed block including reconstructed samples by summing the residual block and the prediction block, and generate a reconstructed picture including the reconstructed block.
[0112] The residual information can be generated through transform processing and quantization processing. For example, the encoding device can derive a residual block between the original block and the prediction block, and perform transform processing on the residual samples (residual sample array) included in the residual block to derive transform coefficients, and then, by performing quantization processing on the transform coefficients, derive quantized transform coefficients to signal the residual-related information (via the bitstream) to the decoding device. Here, the residual information can include position information, transform technology, transform kernel, quantization parameters, value information of the quantized transform coefficients, etc. The decoding device can perform dequantization / inverse transform processing based on the residual information and derive residual samples (or a residual block). The decoding device can generate a reconstructed picture based on the prediction block and the residual block. The encoding device can also dequantize / inverse transform the quantized transform coefficients for inter-frame prediction reference of a later picture to derive a residual block, and generate a reconstructed picture based on it.
[0113] In this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or may still be referred to as transform coefficients for the sake of consistency in expression.
[0114] In this document, the quantized transform coefficients and the transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, the residual information may include information about the transform coefficients, and the information about the transform coefficients may be signaled by the residual coding syntax. The transform coefficients may be derived based on the residual information (or the information about the transform coefficients), and the scaled transform coefficients may be derived by inverse transform (scaling) of the transform coefficients. The residual samples may be derived based on the inverse transform (transform) of the scaled transform coefficients. This may also be applied / expressed in other parts of this document.
[0115] Intra prediction may refer to a prediction for generating prediction samples of a current block based on reference samples in a picture (hereinafter referred to as the current picture) to which the current block belongs. When intra prediction is applied to the current block, neighboring reference samples to be used for intra prediction of the current block may be derived. The neighboring reference samples of the current block may include samples adjacent to the left boundary of the current block having a size of nW×nH and a total of 2×nH samples neighboring the lower left, samples adjacent to the upper boundary of the current block and a total of 2×nW samples neighboring the upper right, and one sample neighboring the upper left of the current block. Alternatively, the neighboring reference samples of the current block may include a plurality of upper neighboring samples and a plurality of left neighboring samples. Additionally, the neighboring reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block having a size of nW×nH, a total of nW samples adjacent to the lower boundary of the current block, and one sample neighboring the lower right of the current block.
[0116] However, some of the neighboring reference samples of the current block may not have been decoded or may not be available yet. In this case, the decoder may configure the neighboring reference samples to be used for prediction by replacing the unavailable samples with available samples. Alternatively, the neighboring reference samples to be used for prediction may be configured by interpolation of the available samples.
[0117] When deriving the neighboring reference samples, (i) the prediction samples may be derived based on the average or interpolation of the neighboring reference samples of the current block, and (ii) the prediction samples may be derived based on the reference samples among the peripheral reference samples of the current block that exist in a specific (prediction) direction of the prediction samples. The case of (i) may be referred to as a non-directional mode or a non-angle mode, and the case of (ii) may be referred to as a directional mode or an angle mode.
[0118] In addition, the predicted sample can also be generated by interpolating between a second neighboring sample and a first neighboring sample among the neighboring reference samples where the predicted sample based on the current block is in a direction opposite to the prediction direction of the intra prediction mode of the current block. The above situation can be referred to as linear interpolation intra prediction (LIP). Additionally, a chrominance prediction sample can be generated based on a linear model using luminance samples. This situation can be referred to as the LM mode.
[0119] Furthermore, a temporary prediction sample of the current block can be derived based on filtered neighboring reference samples, and at least one reference sample (i.e., unfiltered neighboring reference sample) derived according to the intra prediction mode among the existing neighboring reference samples and the temporary prediction sample can be weighted and summed to derive the prediction sample of the current block. The above situation can be referred to as position-dependent intra prediction (PDPC).
[0120] Moreover, the reference sample row with the highest prediction accuracy among the neighboring multi-reference sample rows of the current block can be selected to derive the prediction sample using the reference samples located in the prediction direction on the corresponding row, and then the reference sample row used herein can be indicated (signaled) to the decoding device to perform intra prediction coding. The above situation can be referred to as multi-reference row (MRL) intra prediction or MRL-based intra prediction.
[0121] Also, intra prediction can be performed by dividing the current block into vertical or horizontal sub-partitions based on the same intra prediction mode, and the neighboring reference samples can be derived and used on a sub-partition basis. That is, in this case, the intra prediction mode of the current block also applies to the sub-partitions, and in some cases, the intra prediction performance can be improved by deriving and using the neighboring reference samples on a sub-partition basis. This prediction method can be referred to as intra-sub-partition (ISP) or ISP-based intra prediction.
[0122] The above intra prediction methods can be called intra prediction types separately from the intra prediction mode. Intra prediction types can be referred to by various terms such as intra prediction techniques or additional intra prediction modes. For example, the intra prediction type (or additional intra prediction mode) can include at least one of the above LIP, PDPC, MRL, and ISP. The general intra prediction method other than specific intra prediction types such as LIP, PDPC, MRL, or ISP can be called the normal intra prediction type. The normal intra prediction type is usually applied when no specific intra prediction type is applied and can perform prediction based on the above intra prediction mode. Additionally, post-filtering can be performed on the derived prediction samples as needed.
[0123] Specifically, the intra prediction process can include an intra prediction mode / type determination step, a neighboring reference sample derivation step, and a prediction sample derivation step based on the intra prediction mode / type. Additionally, a post-filtering step can be performed on the derived prediction samples as needed.
[0124] When intra prediction is applied, the intra prediction mode applied to a current block may be determined using the intra prediction modes of neighboring blocks. For example, a decoding device may select one of the mpm candidates of an mpm list derived from the intra prediction modes of neighboring blocks (e.g., left and / or upper neighboring blocks) of the current block based on the received most probable mode (mpm) index, and select one of the other remaining intra prediction modes not included in the mpm candidates (and the planar mode) based on the remaining intra prediction mode information. The mpm list may be configured to include or not include the planar mode as a candidate. For example, if the mpm list includes the planar mode as a candidate, the mpm list may have six candidates. If the mpm list does not include the planar mode as a candidate, the mpm list may have three candidates. When the mpm list does not include the planar mode as a candidate, a non-planar flag (e.g., intra_luma_not_planar_flag) indicating whether the intra prediction mode of the current block is not the planar mode may be signaled. For example, the mpm flag may be signaled first, and when the value of the mpm flag is 1, the mpm index and the non-planar flag may be signaled. Additionally, when the value of the non-planar flag is 1, the mpm index may be signaled. Here, the mpm list is configured not to include the planar mode as a candidate, and the non-planar flag is not signaled first to check whether it is the planar mode because the planar mode is always considered an mpm.
[0125] For example, whether the intra prediction mode applied to the current block is among the MPM candidates (and the planar mode) or in the residual mode can be indicated based on an MPM flag (e.g., Intra_luma_mpm_flag). A value of 1 for the MPM flag can indicate that the intra prediction mode of the current block is within the MPM candidates (and the planar mode), and a value of 0 for the MPM flag can indicate that the intra prediction mode of the current block is not among the MPM candidates (and the planar mode). A value of 0 for the non-planar flag (e.g., Intra_luma_not_planar_flag) can indicate that the intra prediction mode of the current block is the planar mode, and a value of 1 for the non-planar flag can indicate that the intra prediction mode of the current block is not the planar mode. The MPM index can be signaled in the form of an mpm_idx or intra_luma_mpm_idx syntax element, and the residual intra prediction mode information can be signaled in the form of a rem_intra_luma_pred_mode or intra_luma_mpm_remainder syntax element. For example, the residual intra prediction mode information can index the residual intra prediction modes not included in the MPM candidates (and the planar mode) among all the intra prediction modes in the order of the prediction mode numbers to indicate one of them. The intra prediction mode can be the intra prediction mode of the luminance component (samples). Hereinafter, the intra prediction mode information can include at least one of an MPM flag (e.g., Intra_luma_mpm_flag), a non-planar flag (e.g., Intra_luma_not_planar_flag), an MPM index (e.g., mpm_idx or intra_luma_mpm_idx), and residual intra prediction mode information (rem_intra_luma_pred_mode or intra_luma_mpm_remainder). In this document, the MPM list can be referred to by various terms such as the MPM candidate list and candModeList. When MIP is applied to the current block, a separate MPM flag (e.g., intra_mip_mpm_flag), an MPM index (e.g., intra_mip_mpm_idx), and residual intra prediction mode information (e.g., intra_mip_mpm_remainder) for MIP can be signaled, and the non-planar flag is not signaled.
[0126] In other words, generally, when block splitting is performed on an image, the current block to be encoded and the neighboring blocks have similar image characteristics. Therefore, the probability that the current block and the neighboring blocks have the same or similar intra prediction modes is high. Thus, the encoder can use the intra prediction mode of the neighboring blocks to encode the intra prediction mode of the current block.
[0127] For example, an encoder / decoder may configure a list of the most probable modes (MPMs) for a current block. The MPM list may also be referred to as an MPM candidate list. In this document, an MPM may refer to a mode that improves coding efficiency by considering the similarity between a current block and neighboring blocks in intra prediction mode coding. As described above, the MPM list may be configured to include the planar mode, or may be configured not to include the planar mode. For example, when the MPM list includes the planar mode, the number of candidates in the MPM list may be 6. And if the MPM list does not include the planar mode, the number of candidates in the MPM list may be 5.
[0128] The encoder / decoder may configure an MPM list that includes 5 or 6 MPMs.
[0129] To configure the MPM list, three types of modes may be considered: a default intra mode, a neighbor intra mode, and a derived intra mode.
[0130] For the neighbor intra mode, two neighboring blocks may be considered, namely, a left neighboring block and an upper neighboring block.
[0131] As described above, if the MPM list is configured not to include the planar mode, the planar mode is excluded from the list, and the number of MPM list candidates may be set to 5.
[0132] In addition, non - directional modes (or non - angular modes) among the intra prediction modes may include a DC mode based on the average of neighboring reference samples of the current block or an interpolated planar mode.
[0133] When inter - frame prediction is applied, the predictor of an encoding device / decoding device may derive a predicted sample by performing inter - frame prediction in units of blocks. Inter - frame prediction may be a prediction derived in a manner depending on data elements (e.g., sample values or motion information) of a picture other than the current picture. When inter - frame prediction is applied to a current block, a predicted block (predicted sample array) of the current block may be derived based on a reference block (reference sample array) specified by a motion vector on a reference picture indicated by a reference picture index. Here, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information of the current block may be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include inter - frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter - frame prediction, neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block may be referred to as a collocated picture (colPic). For example, a motion information candidate list may be configured based on neighboring blocks of the current block, and a flag or index information indicating which candidate is selected (used) may be signaled to derive the motion vector and / or reference picture index of the current block. Inter - frame prediction may be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the motion information of the current block may be the same as that of the neighboring block. In the skip mode, different from the merge mode, no residual signal may be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of the selected neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be signaled. In this case, the motion vector of the current block may be derived using the sum of the motion vector predictor and the motion vector difference.
[0134] Depending on the inter - frame prediction type (L0 prediction, L1 prediction, Bi - prediction, etc.), the motion information may include L0 motion information and / or L1 motion information. The motion vector in the L0 direction may be referred to as the L0 motion vector or MVL0, and the motion vector in the L1 direction may be referred to as the L1 motion vector or MVL1. The prediction based on the L0 motion vector may be called L0 prediction, the prediction based on the L1 motion vector may be called L1 prediction, and the prediction based on both the L0 motion vector and the L1 motion vector may be called bi - prediction. Here, the L0 motion vector may indicate the motion vector associated with the reference picture list L0 (L0), and the L1 motion vector may indicate the motion vector associated with the reference picture list L1 (L1). The reference picture list L0 may include pictures earlier than the current picture in the output order as reference pictures, and the reference picture list L1 may include pictures later than the current picture in the output order. The previous picture may be called the forward (reference) picture, and the subsequent picture may be called the backward (reference) picture. The reference picture list L0 may also include pictures later than the current picture in the output order as reference pictures. In this case, the previous picture may be indexed first in the reference picture list L0, and the subsequent picture may be indexed later. The reference picture list L1 may also include the previous picture earlier than the current picture in the output order as a reference picture. In this case, the subsequent picture may be indexed first in the reference picture list 1, and the previous picture may be indexed later. The output order may correspond to the picture order count (POC) order.
[0135] Figure 4 Shows an exemplary block - tree structure. Figure 4 Illustratively shows that a CTU is divided into multiple CUs based on a quadtree and a nested multi - type tree structure.
[0136] The bold block edges represent quadtree segmentation, and the remaining edges represent multi - type tree segmentation. The quadtree segmentation using the nested multi - type tree can provide a content - adaptive coding tree structure. A CU may correspond to a coding block (CB). Alternatively, a CU may include a coding block for luminance samples and two coding blocks for corresponding chrominance samples. The size of a CU may be as large as that of a CTU or as small as 4×4 of a luminance sample unit. For example, in the case of the 4:2:0 color format (or chrominance format), the maximum chrominance CB size may be 64×64, and the minimum chrominance CB size may be 2×2.
[0137] In this document, for example, the maximum allowable luminance TB size may be 64×64, and the maximum allowable chrominance TB size may be 32×32. When the width or height of a CB divided according to the tree structure is greater than the maximum transform width or the maximum transform height, the corresponding CB may be automatically (or implicitly) divided until the TB size limits in the horizontal and vertical directions are met.
[0138] In this document, the coding tree scheme can support that the luma (component) blocks and chroma (component) blocks have separate block tree structures. Blocks with separate block tree structures can be blocks that have been coded as separate trees. The case where the luma blocks and chroma blocks in a CTU have the same block tree structure can be represented as SINGLE_TREE (single tree structure). The case where the luma blocks and chroma blocks in a CTU have separate block tree structures can be represented as DUAL_TREE (dual tree structure). In this case, the block tree type of the luma component can be called DUAL_TREE_LUMA, and the block tree type of the chroma component can be called DUAL_TREE_CHROMA. Blocks with a dual tree structure can be blocks that have been coded as a dual tree. For P and B slices / tile groups, the luma and chroma CTBs in a CTU can be restricted to have the same coding tree structure. However, for I slices / tile groups, the luma blocks and chroma blocks can have separate block tree structures respectively. When applying the separate block tree mode, the luma CTB can be divided into CUs based on a specific coding tree structure, and the chroma CTB can be divided into chroma CUs based on another coding tree structure. This can mean that the CUs in I slices / tile groups are composed of coded blocks of the luma component or coded blocks of the two chroma components, and the CUs of P or B slices / tile groups are composed of blocks of the three color components. In this document, a slice can be called a tile / tile group, and a tile / tile group can be called a slice.
[0139] Although the quadtree coding tree structure accompanied by multiple types of trees has been described, the structure for partitioning CUs is not limited to this. For example, the BT structure and TT structure can be interpreted as including the concept of a multi-partition tree (MPT) structure, and CUs can be interpreted as being partitioned by the QT structure and MPT structure. In an example of partitioning CUs by the QT structure and MPT structure, the splitting structure can be determined by signaling a syntax element (e.g., MPT_split_type) that includes information on how many blocks the leaf nodes of the QT structure are divided into, and a syntax element (e.g., MPT_split_mode) that includes information on whether the leaf nodes of the QT structure are divided in the vertical direction or the horizontal direction.
[0140] Furthermore, in another example, CUs can be partitioned by any other method other than the QT structure, BT structure, or TT structure. That is, different from dividing a CU at a lower depth into CUs of 1 / 4 size at a higher depth according to the QT structure, or dividing a CU at a lower depth into CUs of 1 / 2 size at a higher depth according to the BT structure, or dividing a CU at a lower depth into CUs of 1 / 4 or 1 / 2 size at a higher depth according to the TT structure, a CU at a lower depth can be divided into CUs of 1 / 5, 1 / 3, 3 / 8, 3 / 5, 2 / 3, or 5 / 8 size at a higher depth according to the situation, and the CU splitting method is not limited to this.
[0141] Figure 5 Exemplarily shows the hierarchical structure of the encoded image / video.
[0142] Referring to Figure 5 , the encoded image / video is divided into a video coding layer (VCL) that processes the image / video and its own decoding process, a subsystem that transmits and stores the encoded information, and a network abstraction layer (NAL) that is responsible for functions and exists between the VCL and the subsystem.
[0143] In the VCL, VCL data including compressed image data (slice data) is generated, or a parameter set including a picture parameter set (PSP), a sequence parameter set (SPS), and a video parameter set (VPS), or a supplementary enhancement information (SEI) message required for image decoding processing may be generated.
[0144] In the NAL, a NAL unit can be generated by adding header information (NAL unit header) to the raw byte sequence payload (RBSP) generated in the VCL. In this case, the RBSP refers to slice data, parameter sets, SEI messages, etc. generated in the VCL. The NAL unit header may include NAL unit type information specified according to the RBSP data included in the corresponding NAL unit.
[0145] As shown in the figure, the NAL unit can be classified into a VCL NAL unit and a non-VCL NAL unit according to the RBSP generated in the VCL. The VCL NAL unit may mean a NAL unit including information about an image (slice data), and the non-VCL NAL unit may mean a NAL unit including information (parameter set or SEI message) required for decoding the image.
[0146] The above VCL NAL units and non-VCL NAL units can be sent over the network by attaching header information according to the data standard of the subsystem. For example, the NAL unit can be transformed into a data format of a predetermined standard such as the H.266 / VVC file format, the real-time transport protocol (RTP), the transport stream (TS), etc. and sent over various networks.
[0147] As described above, the NAL unit can be specified by the NAL unit type according to the RBSP data structure included in the corresponding NAL unit, and the information about the NAL unit type can be stored in the NAL unit header and signaled.
[0148] For example, NAL units can be classified into VCL NAL unit types and non-VCL NAL unit types according to whether the NAL unit includes information about an image (slice data). The VCL NAL unit types can be classified according to the nature and type of the pictures included in the VCL NAL units, and the non-VCL NAL unit types can be classified according to the type of parameter sets.
[0149] The following are examples of NAL unit types specified according to the type of parameter sets included in the non-VCL NAL unit types.
[0150] - APS (Adaptive Parameter Set) NAL unit: The type of NAL unit including APS
[0151] - DPS (Decoding Parameter Set) NAL unit: The type of NAL unit including DPS
[0152] - VPS (Video Parameter Set) NAL unit: The type of NAL unit including VPS
[0153] - SPS (Sequence Parameter Set) NAL unit: The type of NAL unit including SPS
[0154] - PPS (Picture Parameter Set) NAL unit: The type of NAL unit including PPS
[0155] - PH (Picture Header) NAL unit: The type of NAL unit including PH
[0156] The above NAL unit types may have syntax information of the NAL unit type, and this syntax information can be stored in the NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified by the nal_unit_type value.
[0157] In addition, as described above, a picture can include multiple slices, and a slice can include a slice header and slice data. In this case, a picture header can be further added to multiple slices (slice header and slice data sets) in a picture. The picture header (picture header syntax) can include information / parameters generally applicable to the picture. In this document, slices can be mixed with tile groups or replaced by tile groups. Additionally, in this document, the slice header can be mixed with the tile group header or replaced by the tile group header.
[0158] The slice header (slice header syntax) may include information / parameters that are generally applicable to a slice. The APS (APS syntax) or PPS (PPS syntax) may include information / parameters that are generally applicable to one or more slices or pictures. The SPS (SPS syntax) may include information / parameters that are generally applicable to one or more sequences. The VPS (VPS syntax) may include information / parameters that are generally applicable to multiple layers. The DPS (DPS syntax) may include information / parameters that are generally applicable to the overall video. The DPS may include information / parameters related to the concatenation of coded video sequences (CVS). The high-level syntax (HLS) in this document may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, and slice header syntax.
[0159] In this document, the image / image information encoded by an encoding device and signaled to a decoding device in the form of a bitstream includes not only information related to segmentation in a picture, intra / inter prediction information, residual information, loop filter information, etc., but also information included in the slice header, information included in the APS, information included in the PPS, information included in the SPS, and / or information included in the VPS.
[0160] In addition, in order to compensate for the difference between the original image and the reconstructed image caused by errors occurring in compression encoding processing such as quantization, loop filter processing may be performed on the reconstructed samples or reconstructed pictures as described above. As described above, the loop filter may be performed by the filter of the encoding device and the filter of the decoding device, and a deblocking filter, SAO, and / or an adaptive loop filter (ALF) may be applied. For example, the ALF process may be performed after the deblocking filter process and / or the SAO process. However, even in this case, the deblocking filter process and / or the SAO process may be omitted.
[0161] In addition, in order to increase the encoding efficiency, luminance mapping and chrominance scaling (LMCS) may be applied as described above. LMCS may be referred to as a loop shaper (shaping). In order to increase the encoding efficiency, the control of LMCS and / or the signaling of LMCS-related information may be performed in a hierarchical manner.
[0162] Figure 6 An exemplary hierarchical structure of a CVS according to an embodiment of this document is shown.
[0163] Referring to Figure 6 , a coded video sequence (CVS) may include an SPS, one or more PPSs (picture parameter sets), and one or more subsequent coded pictures. Each coded picture may be divided into rectangular regions. The rectangular regions may be referred to as tiles. One or more tiles collected may form a tile group or a slice. In this case, the tile group header may be linked to the picture parameter set (PPS), and the PPS may be linked to the SPS.
[0164] Figure 7 Exemplarily shown is the hierarchical structure of the CVS according to an embodiment of this document. Figure 8 Exemplarily shown is the hierarchical structure of the CVS according to another embodiment of this document.
[0165] Referring to Figure 8 , the coded video sequence (CVS) may include SPS, PPS, tile group headers, tile data, and / or CTUs. Here, the tile group header and tile data may be referred to as slice header and slice data, respectively.
[0166] The SPS may include local flags to enable tools to be used in the CVS. Additionally, the SPS may be referenced by the PPS including information on parameters changed for each picture. Each coded picture may include one or more coded rectangular domain tiles. The tiles may be grouped to form a raster scan of tile groups. Each tile group is encapsulated with header information called a tile group header. Each tile is composed of a CTU including coded data. Here, the data may include original sample values, predicted sample values, and their luminance and chrominance components (luminance predicted sample values and chrominance predicted sample values).
[0167] According to the existing method, ALF data (ALF parameters) or LMCS data (LMCS parameters) are incorporated into the tile group header. Since a video is composed of multiple pictures and one picture includes multiple tiles, frequently signaling the ALF data (ALF parameters) or LMCS data (LMCS parameters) in units of tile groups results in a problem of reduced coding efficiency.
[0168] According to the embodiment proposed in this document, the ALF parameters or LMCS data (LMCS parameters) may be incorporated into the APS and signaled as follows.
[0169] Referring to Figure 7 , the APS is defined, and the APS may carry necessary ALF data (ALF parameters). In addition, the APS may have self-identification parameters, ALF data, and / or LMCS data. The self-identification parameters of the APS may include the APS ID. That is, the APS may include information representing the APS ID. The tile group header or slice header may use the APS index information to reference the APS. In other words, the tile group header or slice header may include the APS index information, and the ALF process of the target block may be performed based on the LMCS data (LMCS parameters) included in the APS having the APS ID indicated by the APS index information, or the ALF process of the target block may be performed based on the ALF data (ALF parameters) included in the APS having the APS ID indicated by the APS index information. Here, the APS index information may be referred to as APS ID information.
[0170] In an example, the SPS may include a flag that allows the use of the ALF. For example, when the CVS starts, the SPS may be checked, and the flag in the SPS may be checked. For example, the SPS may include the syntax of Table 1 below. The syntax in Table 1 may be part of the SPS.
[0171] [Table 1]
[0172]
[0173] For example, the semantics of the syntax elements included in the syntax of Table 1 may be represented as shown in the table below.
[0174] [Table 2]
[0175]
[0176] That is, the sps_alf_enabled_flag syntax element may indicate whether the ALF is enabled based on whether its value is 0 or 1. The sps_alf_enabled_flag syntax element may be referred to as the ALF enable flag (which may be referred to as the first ALF enable flag), and may be included in the SPS. That is, the ALF enable flag may be signaled in the SPS (or at the SPS level). When the value of the ALF enable flag signaled in the SPS is 1, the ALF may be determined to be substantially enabled for the picture that references the SPS in the CVS. In addition, as described above, the ALF may be individually turned on / off by signaling an additional enable flag at a level lower than the SPS.
[0177] For example, if the ALF tool is enabled for the CVS, an additional enable flag (which may be referred to as the second ALF enable flag) may be signaled in the tile group header or slice header. For example, when the ALF is enabled at the SPS level, the second ALF enable flag may be parsed / signaled. When the value of the second ALF enable flag is 1, the ALF data may be parsed through the tile group header or slice header. For example, the second ALF enable flag may specify the ALF enable conditions for the luminance component and the chrominance component. The ALF data may be accessed through the APS ID information.
[0178] [Table 3]
[0179]
[0180] [Table 4]
[0181]
[0182] For example, the semantics of the syntax elements included in the syntax of Table 3 or Table 4 above may be represented as shown in the table below.
[0183] [Table 5]
[0184]
[0185] [Table 6]
[0186]
[0187]
[0188] The second ALF enable flag may include the tile_group_alf_enabled_flag syntax element or the slice_alf_enabled_flag syntax element.
[0189] The APS corresponding to the tile group or the corresponding slice reference may be identified based on the APS ID information (e.g., the tile_group_aps_id syntax element or the slice_aps_id syntax element). The APS may include ALF data.
[0190] In addition, for example, the structure of the APS including ALF data may be described based on the following syntax and semantics. The syntax in Table 7 may be part of the APS.
[0191] [Table 7]
[0192]
[0193] [Table 8]
[0194]
[0195] As described above, the adaptation_parameter_set_id syntax element may indicate the identifier of the corresponding APS. That is, the APS may be identified based on the adaptation_parameter_set_id syntax element. The adaptation_parameter_set_id syntax element may be referred to as APS ID information. In addition, the APS may include an ALF data field. The ALF data field may be parsed / signaled after the adaptation_parameter_set_id syntax element.
[0196] In addition, for example, an APS extension flag (e.g., the aps_extension_flag syntax element) may be parsed / signaled in the APS. The APS extension flag may indicate whether there is an APS extension data flag (aps_extension_data_flag) syntax element. For example, the APS extension flag may be used to provide an extension point for a later version of the VVC standard.
[0197] Figure 9Shows an exemplary LMCS structure according to an embodiment of this document. Figure 9 The LMCS structure 900 includes a loop mapping unit 910 for the luminance component based on an adaptive piecewise linear (adaptive PWL) model and a luminance-related chrominance residual scaling unit 920 for the chrominance component. The dequantization and inverse transform 911, reconstruction 912, and intra prediction 913 blocks of the loop mapping unit 910 represent the processing applied in the mapped (shaped) domain. The loop filter 915, motion compensation or inter prediction 917 blocks of the loop mapping unit 910 and the reconstruction 922, intra prediction 923, motion compensation or inter prediction 924, loop filter 925 blocks of the chrominance residual scaling unit 920 represent the processing applied in the original (unmapped, unshaped) domain.
[0198] As Figure 9 shown, when LMCS is enabled, at least one of inverse mapping (shaping) processing 914, forward mapping (shaping) processing 918, and chrominance scaling processing 921 can be applied. For example, inverse mapping processing can be applied to the (reconstructed) luminance samples (or luminance samples or array of luminance samples) in the reconstructed picture. The inverse mapping processing can be performed based on the piecewise function (inverse) index of the luminance samples. The piecewise function (inverse) index can identify the segment to which the luminance samples belong. The output of the inverse mapping processing is the modified (reconstructed) luminance samples (or modified luminance samples or array of modified luminance samples). LMCS can be enabled or disabled at the tile group (or slice), picture, or higher level.
[0199] Forward mapping processing and / or chrominance scaling processing can be applied to generate the reconstructed picture. The picture can include luminance samples and chrominance samples. The reconstructed picture with luminance samples can be referred to as the reconstructed luminance picture, and the reconstructed picture with chrominance samples can be referred to as the reconstructed chrominance picture. The combination of the reconstructed luminance picture and the reconstructed chrominance picture can be referred to as the reconstructed picture. The reconstructed luminance picture can be generated based on forward mapping processing. For example, if inter prediction is applied to the current block, forward mapping is applied to the luminance prediction samples derived from the (reconstructed) luminance samples in the reference picture. Since the (reconstructed) luminance samples in the reference picture are generated based on inverse mapping processing, forward mapping can be applied to the luminance prediction samples, and thus the mapped (shaped) luminance prediction samples can be derived. The forward mapping processing can be performed based on the piecewise function index of the luminance prediction samples. The piecewise function index can be derived based on the value of the luminance prediction samples or the value of the luminance samples in the reference picture used for inter prediction. If intra prediction (or intra block copy (IBC)) is applied to the current block, forward mapping is not required because inverse mapping processing has not been applied to the reconstructed samples in the current picture. The (reconstructed) luminance samples in the reconstructed luminance picture are generated based on the mapped luminance prediction samples and the corresponding luminance residual samples.
[0200] The reconstructed chrominance picture can be generated based on chrominance scaling processing. For example, the (reconstructed) chrominance samples in the reconstructed chrominance picture can be derived based on the chrominance residual samples (c res ) and chrominance prediction samples in the current block. The chrominance residual samples (c resScale ) are derived based on the (scaled) chrominance residual samples in the current block and the chrominance residual scaling factor (cScaleInv can be referred to as varScale). The chrominance residual samples (c res ) can be calculated based on the rounded luma prediction sample values of the current block. For example, the scaling factor can be calculated based on the average luma value ave(Y' pred ) of the rounded luma prediction samples. For reference, the (scaled) chrominance residual samples derived based on inverse transform / dequantization can be referred to as c pred resScale , and the chrominance residual samples derived by performing (inverse) scaling processing on the (scaled) chrominance residual samples can be referred to as c res .
[0201] Figure 10 FIG. shows an LMCS structure according to another embodiment of this document. Figure 10 Refer to Figure 9 for description. Here, the differences between the LMCS structure of Figure 10 and the LMCS structure 900 of Figure 9 are mainly described. Figure 10 The loop mapping unit and the luma-related chrominance residual scaling unit of Figure 9 operate in the same (similarly) way as the loop mapping unit 910 and the luma-related chrominance residual scaling unit 920 of
[0202] Refer to Figure 10 . The chrominance residual scaling factor can be derived based on the luma reconstruction samples. In this case, the average luma value (avgYr) can be obtained (derived) based on the neighboring luma reconstruction samples outside the reconstruction block, rather than the internal luma reconstruction samples of the reconstruction block, and the chrominance residual scaling factor can be derived based on the average luma value (avgYr). Here, the neighboring luma reconstruction samples can be the neighboring luma reconstruction samples of the current block, or can be the neighboring luma reconstruction samples of the virtual pipeline data unit (VPDU) including the current block. For example, when intra prediction is applied to the target block, the reconstruction samples can be derived based on the prediction samples derived based on intra prediction. In another example, when inter prediction is applied to the target block, forward mapping is applied to the prediction samples derived based on inter prediction, and the reconstruction samples are generated (derived) based on the rounded (or forward-mapped) luma prediction samples.
[0203] Video / image information signaled via a bitstream may include LMCS parameters (information about LMCS). The LMCS parameters may be configured as high-level syntax (HLS, including slice header syntax), etc. A detailed description and configuration of the LMCS parameters will be described later. As described above, the syntax tables described in this document (and the following embodiments) may be configured / encoded at the encoder side and signaled via the bitstream to the decoder side. The decoder may parse / decode information about LMCS in the syntax table (in the form of syntax components). One or more of the following embodiments to be described may be combined. The encoder may encode the current picture based on the information about LMCS, and the decoder may decode the current picture based on the information about LMCS.
[0204] Loop mapping of the luminance component may adjust the dynamic range of the input signal by redistributing the codewords across the dynamic range to improve the compression efficiency. For luminance mapping, a forward mapping (shaping) function (FwdMap) and a corresponding inverse mapping (shaping) function (InvMap) to the forward mapping function (FwdMap) may be used. The FwdMap function may be signaled using a piecewise linear model. For example, the piecewise linear model may have 16 segments or bins, and the segments may have equal lengths. In one example, the InvMap function does not need to be signaled but is derived from the FwdMap function. That is, the inverse mapping may be the forward mapping function. For example, the inverse mapping function may be mathematically constructed as a symmetric function of the forward mapping, as reflected by the line y = x.
[0205] Loop (luminance) shaping may be used to map the input luminance values (samples) to modified values in the shaping domain. The shaped values may be encoded and then mapped back to the original (unmapped, unshaped) domain after reconstruction. To compensate for the interaction between the luminance signal and the chrominance signal, chrominance residual scaling may be applied. Loop shaping is done by specifying the high-level syntax of the shaper model. The shaper model syntax may signal a piecewise linear model (PWL model). For example, the shaper model syntax may signal a PWL model with 16 bins or segments of the same length. The forward lookup table (FwdLUT) and / or the inverse lookup table (InvLUT) may be derived based on the piecewise linear model. For example, the PWL model pre-computes 1024 current forward (FwdLUT) and inverse (InvLUT) lookup tables (LUTs). As an example, when deriving the forward lookup table FwdLUT, the inverse lookup table InvLUT may be derived based on the forward lookup table FwdLUT. The forward lookup table FwdLUT may map the input luminance value Yi to a modified value Yr, and the inverse lookup table InvLUT may map the modified value Yr to a reconstructed value Y'i. The reconstructed value Y′i may be derived based on the input luminance value Yi.
[0206] In one example, the SPS may include the syntax of Table 9 below. The syntax of Table 9 may include sps_reshaper_enabled_flag as a tool enable flag. Here, sps_reshaper_enabled_flag may be used to specify whether to use a reshaper in an encoded video sequence (CVS). That is, sps_reshaper_enabled_flag may be a flag to enable reshaping in the SPS. In one example, the syntax of Table 9 may be part of the SPS.
[0207] [Table 9]
[0208]
[0209] In one example, the semantics of the syntax elements sps_seq_parameter_set_id and sps_reshaper_enabled_flag may be as shown in Table 10 below.
[0210] [Table 10]
[0211]
[0212] In one example, a tile group header or a slice header may include the syntax of Table 11 or Table 12 below.
[0213] [Table 11]
[0214]
[0215] [Table 12]
[0216]
[0217] The semantics of the syntax elements included in the syntax of Table 11 or Table 12 may include, for example, the matters disclosed in the following table.
[0218] [Table 13]
[0219]
[0220] [Table 14]
[0221]
[0222] As an example, once the flag enabling shaping (i.e., sps_reshaper_enabled_flag) is parsed in the SPS, the tile group header can parse additional data (i.e., the information included in Table 13 or Table 14 above) for constructing look-up tables (FwdLUT and / or InvLUT). To this end, the status of the SPS reshaper flag (sps_reshaper_enabled_flag) can be first checked in the slice header or the tile group header. When sps_reshaper_enabled_flag is true (or 1), additional flags can be parsed, i.e., tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag). The purpose of tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag) can be to indicate the presence of a shaping model. For example, when tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag) is true (or 1), it can indicate that a reshaper exists for the current tile group (or the current slice). When tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag) is false (or 0), it can indicate that no reshaper exists for the current tile group (or the current slice).
[0223] If there is a reshaper and the reshaper is enabled in the current tile group (or current slice), the reshaper model (i.e., tile_group_reshaper_model() or slice_reshaper_model()) can be processed. Further, the additional flag tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag) can also be parsed. The tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag) can indicate whether the reshaping model is used for the current tile group (or slice). For example, if the tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag) is 0 (or false), it can indicate that the reshaping model is not used for the current tile group (or current slice). If the tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag) is 1 (or true), it can indicate that the reshaping model is used for the current tile group (or slice).
[0224] As an example, the tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag) can be true (or 1), and the tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag) can be false (or 0). This means that the reshaping model exists but is not used in the current tile group (or slice). In this case, the reshaping model can be used in future tile groups (or slices). As another example, the tile_group_reshaper_enable_flag can be true (or 1) and the tile_group_reshaper_model_present_flag can be false (or 0). In this case, the decoder uses the reshaper initialized from the previous one.
[0225] When parsing the integer model (i.e., tile_group_reshaper_model() or slice_reshaper_model()) and tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag), it is possible to determine (evaluate) whether the conditions required for chroma scaling exist. The above conditions include Condition 1 (the current tile group / slice has not been intra-coded yet) and / or Condition 2 (the current tile group / slice has not been split into two separate coding quadtree structures for luma and chroma, i.e., the block structure of the current tile group / slice is not a dual-tree structure). If Condition 1 and / or Condition 2 is true and / or tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag) is true (or 1), then tile_group_reshaper_chroma_residual_scale_flag (or slice_reshaper_chroma_residual_scale_flag) can be parsed. When tile_group_reshaper_chroma_residual_scale_flag (or slice_reshaper_chroma_residual_scale_flag) is enabled (if it is 1 or true), it can indicate that chroma residual scaling is enabled for the current tile group (or slice). When tile_group_reshaper_chroma_residual_scale_flag (or slice_reshaper_chroma_residual_scale_flag) is disabled (if it is 0 or false), it can indicate that chroma residual scaling is disabled for the current tile group (or slice).
[0226] The purpose of the tile group reshaping model is to parse the data required to construct a lookup table (LUT). The construction concept of these LUTs is that the distribution of the allowable range of luma values can be divided into multiple bins (e.g., 16 bins), which can be represented using a set of 16 piecewise linear (PWL) equations. Therefore, any luma value falling within a given bin can be mapped to a modified luma value.
[0227] Figure 11 A figure showing an exemplary forward mapping is shown. In Figure 11 five bins are exemplarily shown.
[0228] Referring to Figure 11, the x-axis represents the input luminance value, and the y-axis represents the changed output luminance value. The x-axis is divided into 5 bins or slices, each bin having a length of L. That is, the five bins mapped to the changed luminance values have the same length. The forward lookup table (FwdLUT) can be constructed using the data available from the tile group header (i.e., the reshaper data), so mapping can be facilitated.
[0229] In one embodiment, the output pivot points associated with the bin indices can be calculated. The output pivot points can set (mark) the minimum and maximum boundaries of the output range of the luminance codeword shaping. The calculation process of the output pivot points can be performed by calculating the piecewise cumulative distribution function (CDF) of the number of codewords. The output pivot range can be sliced based on the maximum number of bins to be used and the size of the lookup table (FwdLUT or InvLUT). As an example, the output pivot range can be sliced based on the product between the maximum number of bins and the size of the lookup table (size of the LUT * maximum number of bin indices). For example, if the product between the maximum number of bins and the size of the lookup table is 1024, the output pivot range can be sliced into 1024 entries. This slicing of the output pivot range can be performed (applied or implemented) based on (using) a scaling factor. In one example, the scaling factor can be derived based on Equation 1 below.
[0230] [Equation 1]
[0231] SF = (y2 - y1) * (1 << FP_PREC) + c
[0232] In Equation 1, SF represents the scaling factor, and y1 and y2 represent the output pivot points corresponding to each bin. Additionally, FP_PREC and c can be predetermined constants. The scaling factor determined based on Equation 1 can be referred to as the scaling factor for forward shaping.
[0233] In another embodiment, regarding reverse shaping (reverse mapping), for the defined range of bins to be used (i.e., from reshaper_model_min_bin_idx to reshape_model_max_bin_idx), the input shaping pivot points and the mapped reverse output pivot points corresponding to the mapping pivot points of the forward LUT are obtained (given by the considered bin index * the initial number of codewords). In another example, the scaling factor SF can be derived based on Equation 2 below.
[0234] [Equation 2]
[0235] In Equation 2, SF = (y2 - y1) * (1 << FP_PREC) / (x2 - x1), where SF represents the scaling factor, x1 and x2 represent the input pivot points, and y1 and y2 represent the output pivot points (the output pivot points of inverse mapping) corresponding to each segment (bin). Here, the input pivot points can be pivot points mapped based on a forward look-up table (FwdLUT), and the output pivot points can be pivot points inversely mapped based on an inverse look-up table (InvLUT). Additionally, FP_PREC can be a predetermined constant value. The FP_PREC in Equation 2 can be the same as or different from the FP_PREC in Equation 1. The scaling factor determined based on Equation 2 can be referred to as the scaling factor for inverse shaping. During inverse shaping, the segmentation of the input pivot points can be performed based on the scaling factor of Equation 2. The scaling factor SF is used to slice the range of the input pivot points. Based on the segmented input pivot points, bin indices in the range from 0 to the minimum bin index (reshaper_model_min_bin_idx) and / or from the minimum bin index (reshaper_model_min_bin_idx) to the maximum bin index (reshape_model_max_bin_idx) are assigned pivot values corresponding to the minimum bin value and the maximum bin value.
[0236] In one example, LMCS data (lmcs_data) can be included in the APS. For example, the semantics of the APS can be 32 APSs signaled for encoding.
[0237] The following table shows the syntax and semantics of an exemplary APS according to an embodiment of this document.
[0238] [Table 15]
[0239]
[0240] [Table 16]
[0241]
[0242] Referring to Table 15, type information of APS parameters (e.g., aps_params_type) can be parsed / signaled in the APS. The type information of APS parameters can be parsed / signaled after the adaptation_parameter_set_id.
[0243] The APS_params_type, ALF_APS, and LMCS_APS included in Table 15 above can be described according to the following table. That is, according to the APS_params_type included in Table 15, the types of APS parameters applied to the APS can be set as shown in Table 3.2 included in Table 16.
[0244] Referring to Table 16, for example, aps_params_type can be a syntactic element for classifying the type of corresponding APS parameters. The types of APS parameters can include ALF parameters and LMCS parameters. Referring to Table 16, when the value of the type information (aps_params_type) is 0, the name of aps_params_type can be determined as ALF_APS (or ALF APS), and the type of APS parameters can be determined as ALF parameters (the APS parameters can represent ALF parameters). In this case, the ALF data field (i.e., alf_data()) can be parsed / signaled to APS. When the value of the type information (aps_params_type) is 1, the name of aps_params_type can be determined as LMCS_APS (or LMCS APS), and the type of APS parameters can be determined as LMCS parameters (the APS parameters can represent LMCS parameters). In this case, the LMCS (reshaper model, reshaper) data (i.e., lmcs_data()) can be parsed / signaled to APS.
[0245] Table 17 and / or Table 18 below show the syntax of the reshaper model according to an embodiment. The reshaper model can be referred to as the LMCS model. Although the reshaper model is exemplarily described here as a tile group reshaper, this specification is not necessarily limited to this embodiment. For example, the reshaper model can be included in APS, or the tile group reshaper model can be referred to as a slice reshaper model or LMCS data (LMCS data field). Additionally, the prefix "reshaper_model" or "Rsp" can be used interchangeably with "lmcs". For example, in the following tables and the following description, reshaper_model_min_bin_idx, reshaper_model_delta_max_bin_idx, reshaper_model_max_bin_idx, RspCW, RsepDeltaCW can be used interchangeably with lmcs_min_bin_idx, lmcs_delta_cs_bin_idx, lmx mixed, lmcs_delta_csDcs_bin_idx, Wlm_idx respectively.
[0246] The LMCS data (lmcs_data()) or the reshaper model (tile group reshaper or slice reshaper) included in Table 15 above can be represented as the syntax included in the following table.
[0247] [Table 17]
[0248]
[0249] [Table 18]
[0250]
[0251] The semantics of syntactic elements included in the syntax of Table 17 and / or Table 18 may include, for example, the matters disclosed in the following table.
[0252] [Table 19]
[0253]
[0254]
[0255] [Table 20]
[0256]
[0257]
[0258] The inverse mapping process of the luminance samples according to this document can be described in the form of a standard document as shown in the following table.
[0259] [Table 21]
[0260]
[0261] The identification of the piecewise function indexing process of the luminance samples according to this document can be described in the form of a standard document as shown in the following table. In Table 22, idxYInv can be referred to as the inverse mapping index, and the inverse mapping index can be derived based on the reconstructed luminance samples (lumaSample).
[0262] [Table 22]
[0263]
[0264] The luminance mapping can be performed based on the above embodiments and examples, and the above syntax and the components included therein may be merely exemplary representations. The embodiments in this document are not limited to the above tables or formulas. Hereinafter, a method for performing chrominance residual scaling (scaling of the chrominance components of the residual samples) based on the luminance mapping is described.
[0265] (Luminance - related) Chrominance residual scaling is designed to compensate for the interaction between the luminance signal and its corresponding chrominance signals. For example, it is also signaled at the tile group level whether chrominance residual scaling is enabled. In one example, if luminance mapping is enabled and if dual - tree partitioning (also known as separate chrominance tree) is not applied to the current tile group, an additional flag is signaled to indicate whether luminance - related chrominance residual scaling is enabled. In other examples, when luminance mapping is not used, or when dual - tree partitioning is used in the current tile group, luminance - related chrominance residual scaling is disabled. In another example, for chrominance blocks with an area less than or equal to 4, luminance - related chrominance residual scaling is always disabled.
[0266] Chrominance residual scaling can be based on the average value of the corresponding luminance prediction block (the luminance component of the prediction block to which the intra - prediction mode and / or inter - prediction mode is applied). The scaling operation at the encoder side and / or decoder side can be implemented using fixed - point integer arithmetic based on Equation 3 below.
[0267] [Equation 3]
[0268] c’ = sign(c)*((abs(c)*s + 2CSCALE_FP_PREC - 1) >> CSCALE_FP_PREC)
[0269] In Equation 3, c' represents the scaled chrominance residual sample (the scaled chrominance component of the residual sample), c represents the chrominance residual sample (the chrominance residual sample, the chrominance component of the residual sample), s represents the chrominance residual scaling factor, and CSCALE_FP_PREC represents a (pre - defined) constant value to specify the precision. For example, CSCALE_FP_PREC can be 11.
[0270] Figure 12 is a flowchart showing a method for deriving a chrominance residual scaling index according to an embodiment of this document. Figure 12 The method in Figure 9 can be based on Figure 9 and the tables, equations, variables, arrays, and functions included in the description related to
[0271] In step S1210, it can be determined whether the prediction mode of the current block is an intra - prediction mode or an inter - prediction mode based on the prediction mode information. If the prediction mode is an intra - prediction mode, the current block or the prediction samples of the current block are considered to be in the shaping (mapping) region. If the prediction mode is an inter - prediction mode, the current block or the prediction samples of the current block are considered to be in the original (un - mapped, un - shaped) region.
[0272] In step S1220, when the prediction mode is an intra prediction mode, the average of the current block (or the luminance prediction samples of the current block) can be calculated (derived). That is, the average of the current block in the already shaped region is directly calculated. The average can also be referred to as the average value or the mean value.
[0273] In step S1221, when the prediction mode is an inter prediction mode, forward shaping (forward mapping) can be performed (applied) to the luminance prediction samples of the current block. Through forward shaping, the luminance prediction samples based on the inter prediction mode can be mapped from the original region to the shaped region. In one example, the forward shaping of the luminance prediction samples can be performed based on the shaping models described in Table 17 and / or Table 18 above.
[0274] In step S1222, the average of the forward shaped (forward mapped) luminance prediction samples can be calculated (derived). That is, the averaging process of the forward shaping result can be performed.
[0275] In step S1230, the chroma residual scaling index can be calculated. When the prediction mode is an intra prediction mode, the chroma residual scaling index can be calculated based on the average of the luminance prediction samples. When the prediction mode is an inter prediction mode, the chroma residual scaling index can be calculated based on the average of the forward shaped luminance prediction samples.
[0276] In an embodiment, the chroma residual scaling index can be calculated based on the for loop syntax. The following table shows an exemplary for loop syntax for deriving (calculating) the chroma residual scaling index.
[0277] [Table 23]
[0278]
[0279] In Table 23, idxS represents the chroma residual scaling index, idxFound represents an index indicating whether the chroma residual scaling index that satisfies the condition of the if statement is obtained, S represents a predetermined constant value, and MaxBinIdx represents the maximum allowable bin index. ReshapPivot[idxS + 1] (in other words, LmcsPivot[idxS + 1]) can be derived based on Table 19 and / or Table 20 above.
[0280] In an embodiment, the chroma residual scaling factor can be derived based on the chroma residual scaling index. Equation 4 is an example for deriving the chroma residual scaling factor.
[0281] [Equation 4]
[0282] s = ChromaScaleCoef[idxS]
[0283] In Equation 4, s represents a chroma residual scaling factor, and ChromaScaleCoef can be a variable (or array) derived based on Table 19 and / or Table 20 above.
[0284] As described above, an average luminance value of a reference sample can be obtained, and a chroma residual scaling factor can be derived based on the average luminance value. As described above, a chroma component residual sample can be scaled based on the chroma residual scaling factor, and a chroma component reconstruction sample can be generated based on the scaled chroma component residual sample.
[0285] In one embodiment of this document, a signaling structure for efficiently applying the above LMCS is proposed. According to this embodiment of this document, for example, LMCS data can be included in HLS (i.e., APS), and through lower-level header information of APS (i.e., picture header, slice header), an LMCS model (shaper model) can be adaptively derived by signaling the ID of APS (referred to as header information). The LMCS model can be derived based on LMCS parameters. Additionally, for example, multiple APS IDs can be signaled through header information, and thus, different LMCS models can be applied in units of blocks within the same picture / slice.
[0286] In one embodiment according to this document, a method for efficiently performing operations required for LMCS is proposed. According to the semantics described in Table 19 and / or Table 20 above, it is necessary to perform a division operation of the segment length lmcsCW[i] (also referred to as RspCW[i] in this document) to derive InvScaleCoeff[i]. The segment length of the inverse mapping may not be a power of 2, which means that the division cannot be performed by bit shifting.
[0287] For example, calculating InvScaleCoeff may require up to 16 divisions per slice. According to Table 19 and / or Table 20 above, for 10-bit coding, the range of lmcsCW[i] is from 8 to 511. Therefore, in order to implement the division operation according to lmcsCW[i] using a LUT, the size of the LUT must be 504. Additionally, for 12-bit coding, the range of lmcsCW[i] is from 32 to 2047. Therefore, the LUT size needs to be 2016 to implement the division operation according to lmcsCW[i] using a LUT. That is, division is expensive in terms of hardware implementation. Therefore, it is desirable to avoid division as much as possible.
[0288] In one aspect of this embodiment, lmcsCW[i] can be constrained to be a multiple of a fixed number (or a predetermined number). Therefore, the lookup table (LUT) (the capacity or size of the LUT) for division can be reduced. For example, if lmcsCW[i] becomes a multiple of 2, the size of the LUT replacing the division process can be reduced by half.
[0289] In another aspect of this embodiment, for encoding with a higher internal bit depth, on top of the existing constraint that "the value of lmcsCW[i] should be in the range of (OrgCW>>3) to (OrgCW<<3 - 1)", if the encoding bit depth is higher than 10, lmcsCW[i] is further constrained to be a multiple of 1<<(BitDepthY - 10). Here, BitDepthY can be the luminance bit depth. Thus, the possible number of lmcsCW[i] does not change with the encoding bit depth, and the size of the LUT required to calculate InvScaleCoeff does not increase due to the higher encoding bit depth. For example, for a 12 - bit internal encoding bit depth, the limit value of lmcsCW[i] is a multiple of 4, and the LUT replacing the division process will be the same as that for 10 - bit encoding. This aspect can be implemented separately, but can also be implemented in combination with the above - mentioned aspect.
[0290] In another aspect of this embodiment, lmcsCW[i] can be constrained to a narrower range. For example, lmcsCW[i] can be constrained in the range from (OrgCW>>1) to (OrgCW<<1)-1. Then for 10 - bit encoding, the range of lmcsCW[i] can be [32, 127], and thus, only a LUT of size 96 is needed to calculate InvScaleCoeff.
[0291] In another aspect of this embodiment, lmcsCW[i] can be approximated to the number closest to a power of 2 and used in the shaper design. Thus, the division in the inverse mapping can be performed (and replaced) by bit shifting.
[0292] In an embodiment according to this document, a constraint on the LMCS codeword range is proposed. According to Table 8 above, the value of the LMCS codeword is in the range from (OrgCW>>3) to (OrgCW<<3)-1. This codeword range is too wide. When there is a large difference between RspCW[i] and OrgCW, it may cause visual artifact problems.
[0293] According to an embodiment of this document, a constraint of mapping the LMCS PWL codeword to a narrow range is proposed. For example, the range of lmcsCW[i] can be within the range (OrgCW>>1) to (OrgCW<<1)-1.
[0294] In one embodiment according to this document, a single chroma residual scaling factor is proposed for chroma residual scaling in LMCS. Existing methods for deriving the chroma residual scaling factor use the average value of the corresponding luminance block and derive the slope of each segment of the inverse luminance mapping as the corresponding scaling factor. Additionally, the process of identifying the segment index requires the availability of the corresponding luminance block, which results in a latency issue. This is not desirable for hardware implementation. According to this embodiment of this document, the scaling in the chroma block may not depend on the luminance block value and may not require identifying the segment index. Therefore, the chroma residual scaling process in LMCS can be performed without latency issues.
[0295] In one embodiment according to this document, a single chroma scaling factor can be derived from both the encoder and the decoder based on the luminance LMCS information. When the LMCS luminance model is received, the chroma residual scaling factor can be updated. For example, when the LMCS model is updated, the single chroma residual scaling factor can be updated.
[0296] The following table shows an example of obtaining a single chroma scaling factor according to this embodiment.
[0297] [Table 24]
[0298]
[0299] Referring to Table 24, a single chroma scaling factor (e.g., ChromaScaleCoeff or ChromaScaleCoeffSingle) can be obtained by averaging the inverse luminance mapping slopes of all segments within LMCS_min_bin_idx and lmcs_max_bin_idx.
[0300] Figure 13 Shows a linear fit of the pivot points according to an embodiment of this document. In Figure 13 pivot points P1, Ps, and P2 are shown. The following embodiments or their examples will be described in Figure 13 terms of
[0301] In an example of this embodiment, a single chroma scaling factor can be obtained based on the linear approximation of the luminance PWL mapping between the pivot points lmcs_min_bin_idx and lmcs_max_bin_idx + 1 (LmcsMaxBinIdx + 1). That is, the inverse slope of the linear mapping can be used as the chroma residual scaling factor. For example, Figure 13 the linear line 1 of Figure 13, in P1, the input value is x1 and the mapped value is 0, and in P2, the input value is x2 and the mapped value is y2. The inverse slope (inverse scale) of the linear line 1 is (x2 - x1) / y2, and a single chroma scaling factor ChromaScaleCoeffSingle can be calculated based on the input values and mapped values of the pivot points P1 and P2 and the following formula.
[0302] [Equation 5]
[0303] ChromaScaleCoeffSingle = (x2 - x1)*(1 << CSCALE_FP_PREC) / y2
[0304] In Equation 5, CSCALE_FP_PREC represents a shift factor. For example, CSCALE_FP_PREC can be a predetermined constant value. In one example, CSCALE_FP_PREC can be 11.
[0305] In another example according to this embodiment, referring to Figure 13 , the input value at the pivot point Ps is min_bin_idx + 1, and the mapped value at the pivot point Ps is ys. Therefore, the inverse slope (inverse scale) of the linear line 1 can be calculated as (xs - x1) / ys, and a single chroma scaling factor ChromaScaleCoeffSingle can be calculated based on the input values and mapped values of the pivot points P1 and Ps and the following formula.
[0306] [Equation 6]
[0307] ChromaScaleCoeffSingle = (xs - x1)*(1 << CSCALE_FP_PREC) / ys
[0308] In Equation 6, CSCALE_FP_PREC represents a shift factor (bit shift factor). For example, CSCALE_FP_PREC can be a predetermined constant value. In one example, CSCALE_FP_PREC can be 11, and the bit shift for the inverse scale can be performed based on CSCALE_FP_PREC.
[0309] In another example according to this embodiment, a single chroma residual scaling factor can be derived based on a linear approximation line. Examples of deriving the linear approximation line can include a linear connection of pivot points (i.e., lmcs_min_bin_idx, lmcs_max_bin_idx + 1). For example, the linear approximation result can be represented by the codewords of the PWL mapping. The mapped value y2 at P2 can be the sum of the codewords of all bins (shards), and the difference between the input value at P2 and the input value at P1 (x2 - x1) is OrgCW * (lmcs_max_bin_idx - lmcs_min_bin_idx + 1) (for OrgCW, see Table 19 and / or Table 20 above). The following table shows an example of obtaining a single chroma scaling factor according to the above embodiment.
[0310] [Table 25]
[0311]
[0312] Referring to Table 25, a single chroma scaling factor (e.g., ChromaScaleCoeffSingle) can be obtained from two pivot points (i.e., lmcs_min_bin_idx, lmcs_max_bin_idx). For example, the inverse slope of the linear mapping can be used as the chroma scaling factor.
[0313] In another example of this embodiment, a single chroma scaling factor can be obtained by linear fitting of pivot points to minimize the error (or mean square error) between the linear fitting and the existing PWL mapping. This example can be more accurate than simply connecting the two pivot points at lmcs_min_bin_idx and lmcs_max_bin_idx. There are many ways to find the optimal linear mapping, and examples are described below.
[0314] In one example, the parameters b1 and b0 of the linear fitting formula y = b1 * x + b0 for minimizing the sum of least squares errors can be calculated based on Equation 7 and / or Equation 8 below.
[0315] [Equation 7]
[0316]
[0317] [Equation 8]
[0318]
[0319] In Equation 7 and Equation 8, x is the original luminance value, y is the quantized luminance value, and are the means of x and y, x i and y i represent the values of the i-th pivot point.
[0320] Refer to Figure 13 , another simple approximation for identifying the linear mapping is given as follows:
[0321] - Obtain the linear line 1 by connecting the pivot points of the PWL mapping at lmcs_min_bin_idx and lmcs_max_bin_idx + 1, and calculate lmcs_pivots_linear[i] of this linear line with input values that are multiples of OrgW
[0322] - Use the linear line 1 and sum the differences between the pivot point mapping values using the PWL mapping
[0323] - Obtain the average difference avgDiff
[0324] - Adjust the last pivot point of the linear line according to the average difference, for example, 2 * avgDiff
[0325] - Use the negative slope of the adjusted linear line as the chroma residual scale
[0326] According to the above linear fitting, the chroma scaling factor (i.e., the negative slope of the forward mapping) can be derived (obtained) based on Equation 9 or Equation 10 below
[0327] [Equation 9]
[0328] ChromaScaleCoeffSingle = OrgCW * (1 << CSCALE_FP_PREC)
[0329] / lmcs_pivots_linear[lmcs_min_bin_idx + 1]
[0330] [Equation 10]
[0331] ChromaScaleCoeffSingle = OrgCW * (lmcs_max_bin_idx - lmcs_max_bin_idx + 1)
[0332] * (1 << CSCALE_FP_PREC)) / lmcs_pivots_linear[lmcs_max_bin_idx + 1]
[0333] In the above formula, lmcs_pivots_lienar[i] can be the mapping value of a linear mapping. For a linear mapping, all segments of the PWL mapping between the minimum bin index and the maximum bin index can have the same LMCS codeword (lmcsCW). That is, lmcs_pivots_linear[lmcs_min_bin_idx + 1] can be the same as lmcsCW[lmcs_min_bin_idx].
[0334] In addition, in Equations 9 and 10, CSCALE_FP_PREC represents a shift factor (bit shift factor). For example, CSCALE_FP_PREC can be a predetermined constant value. In one example, CSCALE_FP_PREC can be 11.
[0335] Using a single chroma residual scaling factor (ChromaScaleCoeffSingle), it is no longer necessary to calculate the average of the corresponding luma block and find the index in the PWL linear mapping to obtain the chroma residual scaling factor. Therefore, the coding efficiency using chroma residual scaling can be increased. This not only eliminates the dependence on the corresponding luma block, solves the latency problem, but also reduces the complexity.
[0336] The luminance-related chroma residual scaling process related to LMCS data according to the above-described embodiment can be described in the standard document format as shown in the following table.
[0337] [Table 26]
[0338]
[0339]
[0340] [Table 27]
[0341]
[0342] In another embodiment of this document, the encoder can determine parameters related to a single chroma scaling factor and signal these parameters to the decoder. Using signaling, the encoder can derive the chroma scaling factor using other information available at the encoder. This embodiment aims to eliminate the chroma residual scaling latency problem.
[0343] For example, another example of identifying the linear mapping to be used to determine the chroma residual scaling factor is given as follows:
[0344] - Calculate the lmcs_pivots_linear[i] of this linear line with the input value that is a multiple of OrgW by connecting the pivot points of the PWL mapping at lmcs_min_bin_idx and lmcs_max_bin_idx + 1
[0345] - Use those of the linear line 1 and the luminance PWL mapping to obtain the weighted sum of the differences in the mapped values of the pivot points. The weights can be based on encoder statistics (e.g., the histogram of bins).
[0346] - Obtain the weighted average difference avgDiff.
[0347] - Adjust the last pivot point of the linear line 1 according to the weighted average difference, e.g., 2 * avgDiff
[0348] - Calculate the chroma residual scale using the negative slope of the adjusted linear line.
[0349] The following table shows an example of the syntax for signaling the y value used for chroma scaling factor derivation.
[0350] [Table 28]
[0351]
[0352] In Table 28, the syntax element lmcs_chroma_scale can specify a single chroma (residual) scaling factor (ChromaScaleCoeffSingle = lmcs_chroma_scale) for LMCS chroma residual scaling. That is, the information about the chroma residual scaling factor can be directly signaled, and the signaled information can be derived as the chroma residual scaling factor. In other words, the value of the signaled information about the chroma residual scaling factor can be (directly) derived as the value of a single chroma residual scaling factor. Here, the syntax element lmcs_chroma_scale can be signaled together with other LMCS data (i.e., syntax elements related to the absolute value and sign of the codeword, etc.).
[0353] Alternatively, the encoder can only signal the necessary parameters to derive the chroma residual scaling factor at the decoder. To derive the chroma residual scaling factor at the decoder, the input value x and the mapped value y are required. Since the x value is the bin length and is a known number, it does not need to be signaled. After all, only the y value needs to be signaled to derive the chroma residual scaling factor. Here, the y value can be the mapped value of any pivot point in the linear mapping (i.e., Figure 13 the mapped value of P2 or Ps in).
[0354] The following table shows an example of the syntax for signaling the mapped value used for deriving the chroma residual scaling factor.
[0355] [Table 29]
[0356]
[0357] [Table 30]
[0358]
[0359] One of the syntaxes of the above Table 29 and Table 30 can be used to signal the y value at any linear pivot point specified to the encoder and decoder. That is, the encoder and decoder can use the same syntax to derive the y value.
[0360] First, describe the embodiment according to Table 29. In Table 29, lmcs_cw_linear can represent the mapping value at Ps or P2. That is, in the embodiment according to Table 29, a fixed number can be signaled by lmcs_cw_linear.
[0361] In an example according to this embodiment, if lmcs_cw_linear represents the mapping value of a bin (i.e., Figure 13 lmcs_pivots_linear[lmcs_min_bin_idx + 1] in Ps of
[0362] [Equation 11]
[0363] ChromaScaleCoeffSingle = OrgCW * (1 << CSCALE_FP_PREC) / lmcs_cw_linear
[0364] In another example according to this embodiment, if lmcs_cw_linear represents lmcs_max_bin_idx + 1 (i.e., Figure 13 lmcs_pivots_linear[lmcs_max_bin_idx + 1] in P2 of
[0365] [Equation 12]
[0366] ChromaScaleCoeffSingle = OrgCW * (lmcs_max_bin_idx - lmcs_max_bin_idx + 1)
[0367] * (1 << CSCALE_FP_PREC) / lmcs_cw_linear
[0368] In the above formula, CSCALE_FP_PREC represents a shift factor (bit shift factor). For example, CSCALE_FP_PREC can be a predetermined constant value. In one example, CSCALE_FP_PREC can be 11.
[0369] Next, an embodiment according to Table 30 will be described. In this embodiment, the available signal notifies lmcs_cw_linear as the Δ value relative to a fixed number (i.e., lmcs_delta_abs_cw_linear, lmcs_delta_sign_cw_linear_flag). In an example of this embodiment, when lmcs_cw_linear represents the mapped value in lmcs_pivots_linear[lmcs_min_bin_idx + 1] (i.e., Figure 13 Ps), lmcs_cw_linear_delta and lmcs_cw_linear can be derived based on the following formula.
[0370] [Equation 13]
[0371] lmcs_cw_linear_delta = (1 - 2 * lmcs_delta_sign_cw_linear_flag) * lmcs_delta_abs_linear_cw
[0372] [Equation 14]
[0373] lmcs_cw_linear = lmcs_cw_linear_delta + OrgCW
[0374] In another example of this embodiment, when lmcs_cw_linear represents the mapped value in lmcs_pivots_linear[lmcs_max_bin_idx + 1] (i.e., Figure 13 P2), lmcs_cw_linear_delta and lmcs_cw_linear can be derived based on the following formula.
[0375] [Equation 15]
[0376] lmcs_cw_linear_delta = (1 - 2 * lmcs_delta_sign_cw_linear_flag) * lmcs_delta_abs_linear_cw
[0377] [Equation 16]
[0378] lmcs_cw_linear = lmcs_cw_linear_delta
[0379] +OrgCW * (lmcs_max - bin_idx - lmcs_max_bin_idx + 1)
[0380] In the above formula, OrgCW can be a value derived based on the above Table 19 and / or Table 20.
[0381] The luminance - related chrominance residual scaling process for semantics and / or chrominance samples related to LMCS data according to the present embodiment above can be described in the standard document format as shown in the following table.
[0382] [Table 31]
[0383]
[0384] [Table 32]
[0385]
[0386]
[0387] Figure 14 Shows an example of linear shaping (or linear shaping, linear mapping) according to an embodiment of this document. That is, in this embodiment, a linear shaper is proposed to be used in LMCS. For example, Figure 14 This example in
[0388] may relate to forward linear shaping (mapping). Figure 14 Referring to Figure 14 , the linear shaper may include two pivot points, namely, P1 and P2. P1 and P2 may represent the input value and the mapped value. For example, P1 may be (min_input, 0) and P2 may be (max_input, max_mapped). Here, min_input represents the minimum input value, and max_input represents the maximum input value. Any input value less than or equal to min_input is mapped to 0, and any input value greater than max_input is mapped to max_mapped. Any input luminance value within min_input and max_input is linearly mapped to other values.
[0389] In another embodiment according to this document, another example of a method for signaling the linear shaper may be proposed. The pivot points P1, P2 of the linear shaper model may be explicitly signaled. The following table shows an example of the syntax and semantics for explicitly signaling the linear shaper model according to this example.
[0390] [Table 33]
[0391]
[0392] [Table 34]
[0393]
[0394]
[0395] Referring to Table 33 and Table 34, the input value of the first pivot point can be derived based on the syntactic element lmcs_min_input, and the input value of the second pivot point can be derived based on the syntactic element lmcs_max_input. The mapped value of the first pivot point can be a predetermined value (a value known to both the encoder and the decoder), for example, the mapped value of the first pivot point is 0. The mapped value of the second pivot point can be derived based on the syntactic element lmcs_max_mapped. That is, the linear shaping model can be signaled explicitly (directly) based on the information signaled in the syntax of Table 33.
[0396] Alternatively, lmcs_max_input and lmcs_max_mapped can be signaled as Δ values. The following table shows examples of the syntax and semantics of signaling the linear shaping model as Δ values.
[0397] [Table 35]
[0398]
[0399] [Table 36]
[0400]
[0401] Referring to Table 36, the input value of the first pivot point can be derived based on the syntactic element lmcs_min_input. For example, lmcs_min_input can have a mapped value of 0. lmcs_max_input_delta can specify the difference between the input value of the second pivot point and the maximum luminance value (i.e., (1<<bitdepthY)-1). lmcs_max_mapped_delta can specify the difference between the mapped value of the second pivot point and the maximum luminance value (i.e., (1<<bitdepthY)-1).
[0402] According to the embodiments of this document, forward mapping of luminance prediction samples, inverse mapping of luminance reconstruction samples, and chrominance residual scaling can be performed based on the above examples of the linear shaper. In one example, in the inverse mapping based on the linear shaper, the inverse scaling for luminance (reconstruction) samples (pixels) may only require one inverse scaling factor. The same applies to forward mapping and chrominance residual scaling. That is, the steps of determining ScaleCoeff[i], InvScaleCoeff[i], and ChromaScaleCoeff[i] (where i is the bin index) can be replaced by only one single factor. Here, a single factor is the fixed-point representation of the (positive) slope or inverse slope of the linear mapping. In one example, the inverse luminance mapping scaling factor (the inverse scaling factor in the inverse mapping of luminance reconstruction samples) can be derived based on at least one of the following equations.
[0403] [Equation 17]
[0404] InvScaleCoeffSingle = OrgCW / lmcsCWLinear
[0405] [Equation 18]
[0406] InvScaleCoeffSingle = OrgCW * (lmcs_max_bin_idx - lmcs_max_bin_idx + 1)
[0407] / lmcsCWL_inearAll
[0408] [Equation 19]
[0409] InvScaleCoeffSingle = (lmcs_max_input - lmcs_min_input) / lmcsCWLinearAll
[0410] The lmcsCWLinear in Equation 17 can be derived from Table 31 above. The lmcsCWLinearALL in Equation 18 and Equation 19 can be derived from at least one of Tables 33 to 36 above. In Equation 17 or Equation 18, OrgCW can be derived from Table 19 and / or Table 20.
[0411] The following table describes the equations and syntax (conditional statements) indicating the forward mapping process for luminance samples (i.e., luminance prediction samples) in picture reconstruction. In the following tables and equations, FP PREC is a constant value for bit shift and can be a predetermined value. For example, FP PREC can be 11 or 15.
[0412] [Table 37]
[0413]
[0414] [Table 38]
[0415]
[0416] Table 37 can be used to derive the luminance samples of the forward mapping in the luminance mapping process based on Tables 17 to 20 above. That is, Table 37 can be described together with Tables 19 and 20. In Table 37, the luminance (prediction) samples PredMapPSamples[i][j] of the forward mapping as the output can be derived from the luminance (prediction) samples predSamples[i][j] as the input. The idxY of Table 37 can be called the (forward) mapping index, and the mapping index can be derived based on the predicted luminance samples.
[0417] Table 38 can be used to derive the luminance samples of the forward mapping in the luminance mapping based on the linear shaper. For example, lmcs_min_input, lmcs_max_input, lmcs_max_mapped, and ScaleCoeffSingle of Table 38 can be derived from at least one of Tables 33 to 36. In Table 38, when "lmcs_min_input < predSamples[i][j] < lmcs_max_input", the luminance (prediction) samples PredMapSamples[i][j] of the forward mapping can be derived from the input luminance (prediction) samples predSamples[i][j] as the output. Comparing between Table 37 and Table 38, the change with respect to the existing LMCS according to the application of the linear shaper can be seen from the perspective of the forward mapping.
[0418] The following formula and table describe the inverse mapping process of the luminance samples (i.e., the luminance reconstruction samples). In the following formula and table, the "lumaSample" as the input can be the luminance reconstruction sample before the inverse mapping (before modification). The "invSample" as the output can be the inverse mapped (modified) luminance reconstruction sample. In other cases, the clipped invSample can be called the modified luminance reconstruction sample.
[0419] [Equation 20]
[0420] invSample = InputPivot[idxYInv] + (InvScaleCoeff[idxYInv] *
[0421] (lumaSample - LmcsPivot[idxYInv]) + (1 << (FP - PREC - 1))) >> FP - PREC
[0422] [Equation 21]
[0423] invSample = lmcs_min_input
[0424] +(InvScaleCoeffSingle * (lumaSample - lmcs_min_input)+(1 << (FP_PREC - 1)))
[0425] >> FP_PREC
[0426] [Table 39]
[0427]
[0428] [Table 40]
[0429]
[0430] Equation 21 can be used to derive the luminance samples of the inverse mapping in the luminance mapping according to this document. In Equation 20, the index idxInv can be derived based on Table 50, Table 51, or Table 52 described later.
[0431] Equation 21 can be used to derive the luminance samples of the inverse mapping from the luminance mapping according to the application of the linear shaper. For example, lmcs_min_input of Equation 21 can be derived from at least one of Tables 33 to 36. Through the comparison between Equation 20 and Equation 21, the change with respect to the existing LMCS according to the application of the linear shaper can be seen from the perspective of the forward mapping.
[0432] Table 39 may include examples of equations for deriving the luminance samples of the inverse mapping in the luminance mapping. For example, the index idxInv can be derived based on Table 50, Table 51, or Table 52 described later.
[0433] Table 40 may include other examples of equations for deriving the luminance samples of the inverse mapping in the luminance mapping. For example, lmcs_min_input and / or lmcs_max_mapped of Table 40 can be derived from at least one of Tables 33 to 36, and / or InvScaleCoeffSingle of Table 40 is at least Tables 33 to 36, and / or Equations 17 to 19 can be derived by one.
[0434] Based on the above example of the linear shaper, the segmented index identification process can be omitted. That is, in this example, since there is only one segment with valid shaped luminance pixels, the segmented index identification process for the inverse luminance mapping and chrominance residual scaling can be removed. Therefore, the complexity of the inverse luminance mapping can be reduced. In addition, the delay problem caused by relying on the luminance segmented index identification during chrominance residual scaling can be eliminated.
[0435] According to the embodiments using the above linear shaper, the following advantages can be provided for LMCS: i) The encoder shaper design can be simplified, thereby preventing possible artifacts caused by sudden changes between piecewise linear segments; ii) The decoder inverse mapping process that can remove the piecewise index identification process can be simplified by eliminating the piecewise index identification process; iii) By removing the piecewise index identification process, the latency problem caused by depending on the corresponding luma block in chroma residual scaling can be removed; iv) The signaling overhead can be reduced, and more frequent updates of the shaper can be made more feasible; v) For many places where a loop with 16 segments was required in the past, the loop can be eliminated. For example, in order to derive InvScaleCoeff[i], the number of division operations according to lmcsCW[i] can be reduced to 1.
[0436] In another embodiment according to this document, a flexible-bin based LMCS is proposed. Here, a flexible bin may refer to a bin whose number is not fixed to a predetermined (predefined, specific) number. In the existing embodiments, the number of bins in LMCS is fixed to 16, and for the input sample values, these 16 bins are equally distributed. In this embodiment, a flexible number of bins is proposed, and these segments (bins) may not be equally distributed in terms of the original pixel values.
[0437] The following table exemplarily shows the syntax of the LMCS data (data field) according to this embodiment and the semantics of the syntax elements included therein.
[0438] [Table 41]
[0439]
[0440] [Table 42]
[0441]
[0442] Referring to Table 41, the information lmcs_num_bins_minus1 regarding the number of bins can be signaled. Referring to Table 42, lmcs_num_bins_minus1 + 1 can be equal to the number of bins, and the number of bins can be in the range from 1 to (1 << BitDepthY) - 1. For example, lmcs_num_bins_minus1 or lmcs_num_bins_minus1 + 1 can be a multiple of a power of 2.
[0443] In the embodiments described with Tables 41 and 42, regardless of whether the shaper is linear (signaling of lmcs_num_bins_minus1), the number of pivot points can be derived based on lmcs_num_bins_minus1 (information about the number of bins), and the input value and mapped value of the pivot points (LmcsPivot_input[i], LmcsPivot_mapped[i]) can be derived based on the sum of the signaled codeword values (lmcs_delta_input_cw[i], lmcs_delta_mapped_cw[i]) (where the initial input value LmcsPivot_input[0] and the initial output value LmcsPivot_mapped[0] are 0).
[0444] Figure 15 Shows an example of linear forward mapping in the embodiments of this document. Figure 16 Shows an example of inverse forward mapping in the embodiments of this document.
[0445] In accordance with Figure 15 and Figure 16 embodiments, a method for supporting both conventional LMCS and linear LMCS is proposed. In an example according to this embodiment, conventional LMCS and / or linear LMCS can be indicated based on the syntax element lmcs_is_linear. In the encoder, after determining the linear LMCS line, the mapped values (i.e., the mapped values in pL of Figure 15 and Figure 16 ) can be divided into equal segments (i.e., LmcsMaxBinIdx - lmcs_min_bin_idx + 1). The codewords in bin LmcsMaxBinIdx can be signaled using the above LMCS data or the syntax of the shaper mode.
[0446] The following table exemplarily shows the syntax of the LMCS data (data field) according to an example of this embodiment and the semantics of the syntax elements included therein.
[0447] [Table 43]
[0448]
[0449] [Table 44]
[0450]
[0451]
[0452] The following table exemplarily shows the syntax of the LMCS data (data field) according to another example of this embodiment and the semantics of the syntax elements included therein.
[0453] [Table 45]
[0454]
[0455] [Table 46]
[0456]
[0457]
[0458] Referring to Tables 43 to 46, when lmcs_is_linear_flag is true, all LMCSDeltaCW[i] between lmcs_min_bin_idx and LmcsMaxBinIdx may have the same value. That is, lmcsCW[i] for all segments between lmcs_min_bin_idx and LmcsMaxBinIdx may have the same value. The scaling, inverse scaling, and chroma scaling for all segments between lmcs_min_bin_idx and lmcsMaxBinIdx may be the same. Then if the linear shaper is true, the segment index does not need to be derived, and it can use the scaling and inverse scaling from only one segment.
[0459] The following table exemplarily shows the identification process of the segment index according to the present embodiment.
[0460] [Table 47]
[0461]
[0462] According to another embodiment of this document, the application of the conventional 16-segment PWL LMCS and the linear LMCS may depend on the higher-level syntax (i.e., the sequence level).
[0463] The following table exemplarily shows the syntax of the SPS according to the present embodiment and the semantics of the syntax elements included therein.
[0464] [Table 48]
[0465]
[0466] [Table 49]
[0467]
[0468] Referring to Tables 48 and 49, the enabling of the conventional LMCS and / or the linear LMCS can be determined (signaled) by the syntax elements included in the SPS. Referring to Table 48, based on the syntax element sps_linear_lmcs_enabled_flag, one of the conventional LMCS or the linear LMCS can be used on a sequence-by-sequence basis.
[0469] In addition, whether only linear LMCS or conventional LMCS or both are enabled may also depend on the profile level. In one example, for a specific profile (i.e., the SDR profile), only linear LMCS may be allowed, for another profile (i.e., the HDR profile), only conventional LMCS may be allowed, and for another profile, both conventional LMCS and / or linear LMCS may be allowed.
[0470] According to another embodiment of this document, the LMCS segment index identification process can be used in inverse luminance mapping and chrominance residue scaling. In this embodiment, the identification process of the segment index can be used for the blocks where chrominance residue scaling is enabled, and is also called for all luminance samples in the shaping (mapping) domain. This embodiment aims to keep its complexity low.
[0471] The following table shows the identification process (derivation process) of the existing segment function index.
[0472] [Table 50]
[0473]
[0474] In an example, in the segment index identification process, the input samples can be classified into at least two categories. For example, the input samples can be classified into three categories, the first, second, and third categories. For example, the first category can represent samples (their values) less than LmcsPivot[lmcs_min_bin_idx + 1], the second category can represent samples (their values) greater than or equal to LmcsPivot[LmcsMaxBinIdx], and the third category can indicate samples (their values) between LmcsPivot[lmcs_min_bin_idx + 1] and LmcsPivot[LmcsMaxBinIdx].
[0475] In this embodiment, an optimization of the identification process by eliminating category classification is proposed. This is because the input of the segment index identification process is the luminance value in the shaping (mapping) domain, and there should be no value exceeding the mapping values at the pivot points lmcs_min_bin_idx and LmcsMaxBinIdx + 1. Therefore, the conditional processing of classifying samples into categories in the existing segment index identification process is unnecessary. For more details, specific examples will be described in the table below.
[0476] In an example according to this embodiment, the identification process included in Table 50 can be replaced by one of the following Table 51 or Table 52. Referring to Table 51 and Table 52, the first two categories of Table 50 can be removed, and for the last category, the boundary value (the second boundary value or the end point) in the iterative for loop is changed from LmcsMaxBinIdx to LmcsMaxBinIdx + 1. That is, the identification process can be simplified, and the complexity of segment index derivation can be reduced. Therefore, the LMCS-related encoding can be efficiently performed according to this embodiment.
[0477] [Table 51]
[0478]
[0479] [Table 52]
[0480]
[0481] Referring to Table 52, the comparison process (the formula corresponding to the condition of the if statement) corresponding to the condition of the if statement can be iteratively performed for all bin indices from the minimum bin index to the maximum bin index. When the formula corresponding to the condition of the if statement is true, the bin index can be derived as the inverse mapping index of the inverse luminance mapping (or the inverse scaling index of the chrominance residual scaling). Based on the inverse mapping index, the modified reconstructed luminance sample (or the scaled chrominance residual sample) can be derived.
[0482] In the existing embodiment, for a slice encoded as a separate block tree (e.g., an intra slice), the delay caused by the chrominance residual scaling dependency in the corresponding luminance block is higher than that of a slice encoded as a dual tree. Therefore, in the existing embodiment, the LMCS chrominance residual scaling is not applied to a slice encoded as a separate block tree.
[0483] In this embodiment, the chrominance residual scaling can also be applied to a slice encoded as a separate tree. When using the above single chrominance residual scaling factor, since there is no dependency between the chrominance residual scalings in the corresponding luminance block, there may be no delay caused by the application of the chrominance residual scaling.
[0484] The following table shows the syntax and semantics of the slice header according to this embodiment.
[0485] [Table 53]
[0486]
[0487] [Table 54]
[0488]
[0489] Referring to Table 54, regardless of the conditional clause (or its flag) indicating whether the current block has a dual-tree structure or a single-tree structure, the chrominance residual scaling flag can be signaled.
[0490] In an embodiment according to this document, ALF data and / or LMCS data can be signaled in the APS. For example, 32 APSs can be used. In one example, if all APSs are used for ALF and / or LMCS, buffering of the APSs may require approximately 10 KB of on-chip memory. To limit (reduce) the memory required to store ALF / LMCS parameters and the computational complexity required for LMCS, this embodiment proposes a method of limiting the number of ALF and / or LMCS APSs.
[0491] In an example according to this embodiment, regardless of the number of slices or tiles included in the picture, one LMCS model per picture can be used (allowed). The number of APSs for LMCS can be less than 32. For example, the number of APSs for LMCS can be 4.
[0492] The following table shows the semantics related to APS according to this embodiment.
[0493] [Table 55]
[0494]
[0495] Referring to Table 55, the maximum number of LMCS APSs can be determined in advance. For example, the maximum number of LMCS APSs can be four. Referring to Tables 15 and 55, multiple APSs can include LMCS APSs. The (maximum) number of LMCS APSs can be four. In one example, the LMCS data field included in one of the LMCS APSs included in the LMCS APSs can be used in the LMCS process of the current block in the current picture.
[0496] The following table shows the semantics of the syntax elements included in the slice header (or picture header).
[0497] [Table 56]
[0498]
[0499] The syntax elements described in Table 56 can be described with reference to Table 54. In the example, the syntax element slice_lmcs_aps_id can be included in the slice header. In another example, the syntax element slice_lmcs_aps_id in Table 56 can be included in the picture header, and in this case, slice_lmcs_aps_id can be modified to ph_lmcs_aps_id.
[0500] The following drawings are created to illustrate specific examples of this specification. Since the names of specific devices or the names of specific signals / messages / fields described in the drawings are presented as examples, the technical features of this specification are not limited to the specific names used in the following drawings.
[0501] Figure 17 and Figure 18 schematically shows examples of a video / image encoding method and related components according to an embodiment of this document. Figure 17 The method disclosed in Figure 2 can be executed by the encoding device disclosed in Figure 17 Specifically, for example, S1700 and S1710 of Figure 17 can be executed by the predictor 220 of the encoding device, S1720 can be executed by the residual processor 230 of the encoding device, S1730 can be executed by the predictor 220 or the residual processor 230 of the encoding device, S1740 can be executed by the residual processor 230 or the adder 250 of the encoding device, S1750 can be executed by the residual processor 230 of the encoding device, and S1760 can be executed by the entropy encoder 240 of the encoding device. Figure 17 The method disclosed in
[0502] may include the embodiments described above in this document.
[0502] Referring to Figure 17 , the encoding device can derive an inter-frame prediction mode (S1700) for the current block in the current picture. The encoding device can derive at least one of various modes disclosed in this document among the inter-frame prediction modes.
[0503] The encoding device can generate predicted luminance samples based on the inter-frame prediction mode (S1710). The encoding device can generate predicted luminance samples by performing prediction on the original samples included in the current block.
[0504] The encoding device can derive predicted chrominance samples. The encoding device can derive residual chrominance samples based on the original chrominance samples and the predicted chrominance samples of the current block. For example, the encoding device can derive residual chrominance samples based on the difference between the predicted chrominance samples and the original chrominance samples.
[0505] The encoding device can derive bins and LMCS codewords for luminance mapping. The encoding device can derive bins and / or LMCS codewords based on SDR or HDR.
[0506] The encoding device can derive LMCS-related information (S1720). The LMCS-related information may include bins and LMCS codewords for luminance mapping.
[0507] The encoding device may generate mapped predicted luminance samples based on the mapping process of luminance samples (S1730). The encoding device may generate mapped predicted luminance samples based on the bins for luminance mapping and / or the LMCS codewords. For example, the encoding device may derive the input values and mapped values (output values) of the pivot points for luminance mapping, and may generate mapped predicted luminance samples based on the input values and the mapped values. In one example, the encoding device may derive a mapping index (idxY) based on the first predicted luminance sample, and may generate the first mapped predicted luminance sample based on the input value and the mapped value of the pivot point corresponding to the mapping index. In other examples, a linear mapping (linear shaping, linear LMCS) may be used, and the mapped predicted luminance samples may be generated based on the forward mapping scaling factors derived from two pivot points in the linear mapping. Therefore, due to the linear mapping, the index derivation process may be omitted.
[0508] The encoding device may generate scaled residual chrominance samples. Specifically, the encoding device may derive a chrominance residual scaling factor and generate scaled residual chrominance samples based on the chrominance residual scaling factor. Here, the chrominance residual scaling at the encoding stage may be referred to as forward chrominance residual scaling. Therefore, the chrominance residual scaling factor derived by the encoding device may be referred to as the forward chrominance residual scaling factor, and forward scaled residual chrominance samples may be generated.
[0509] The encoding device may generate reconstructed luminance samples. The encoding device may generate reconstructed luminance samples based on the mapped predicted luminance samples. Specifically, the encoding device may sum the above-mentioned residual luminance samples and the mapped predicted luminance samples, and may generate reconstructed luminance samples based on the result of the summation.
[0510] The encoding device may generate modified reconstructed luminance samples based on the reverse mapping process of luminance samples (S1740). The encoding device may generate modified reconstructed luminance samples based on the bins for luminance mapping, the LMCS codewords, and the reconstructed luminance samples. The encoding device may generate modified reconstructed luminance samples through the reverse mapping process of the reconstructed luminance samples. For example, the encoding device may derive a reverse mapping index (i.e., invYIdx) based on the reconstructed luminance samples and / or the mapped values assigned to each bin index (i.e., LmcsPivot[i], i = lmcs_min_bin_idx...LmcsMaxBinIdx+1) in the reverse mapping process. The encoding device may generate modified reconstructed luminance samples based on the mapped value (LmcsPivot[invYIdx]) assigned to the reverse mapping index.
[0511] The encoding device may generate residual luminance samples based on the mapped predicted luminance samples. For example, the encoding device may derive residual luminance samples based on the difference between the mapped predicted luminance samples and the original luminance samples.
[0512] The encoding device can derive residual information (S1750). In one example, the encoding device can generate residual information based on the mapped predicted luminance samples and the modified reconstructed luminance samples. For example, the encoding device can derive residual information based on the scaled residual chrominance samples and / or the residual luminance samples. The encoding device can derive transform coefficients based on the transform processing of the scaled residual chrominance samples and / or the residual luminance samples. For example, the transform processing can include at least one of DCT, DST, GBT, or CNT. The encoding device can derive quantized transform coefficients based on the quantization processing of the transform coefficients. The quantized transform coefficients can have a one-dimensional vector form based on the coefficient scan order. The encoding device can generate residual information specifying the quantized transform coefficients. The residual information can be generated by various encoding methods such as exponential Golomb, CAVLC, CABAC, etc.
[0513] The encoding device can encode image / video information (S1760). The image information can include information about the LMCS data and / or the residual information. For example, the LMCS-related information can include information about the linear LMCS. In one example, at least one LMCS codeword can be derived based on the information about the linear LMCS. The encoded video / image information can be output in the form of a bitstream. The bitstream can be sent to the decoding device via a network or a storage medium.
[0514] According to the embodiments of this document, the image / video information can include various information. For example, the image / video information can include the information disclosed in at least one of Tables 1 to 56 above.
[0515] In an embodiment, the image information can include the LMCS APS. The LMCS APS can include an LMCS data field. The LMCS data field can include LMCS-related information. The LMCS codewords used in the mapping process and the inverse mapping process can be derived based on the LMCS data field. The maximum number of LMCS APSs can be a predetermined value. For example, the maximum number (predetermined value) of LMCS APSs can be four.
[0516] In an embodiment, during the mapping process of the luminance samples, a mapping index can be derived based on the predicted luminance samples, and a mapped predicted luminance sample can be generated using a first mapping value based on the mapping index.
[0517] In an embodiment, the LMCS-related information can include information about the bins for inverse mapping. Based on the information about the bins, a minimum bin index and a maximum bin index are derived. During the inverse mapping process of the luminance samples, an inverse mapping index is derived based on the bin indices from the minimum bin index to the maximum bin index based on the mapping value, and a modified reconstructed luminance sample is generated using a second mapping value based on the inverse mapping index.
[0518] In an embodiment, an encoding device may generate a segmented index for chrominance residual scaling. The encoding device may derive a chrominance residual scaling factor based on the segmented index. The encoding device may generate scaled residual chrominance samples based on the residual chrominance samples and the chrominance residual scaling factor.
[0519] In one embodiment, when the current block has a single-tree structure or a dual-tree structure (when the current block has a separate tree structure, when the current block is encoded as a separate tree), a chrominance residual scaling enable flag indicating whether chrominance residual scaling is applied to the current block may be generated by the encoding device. Alternatively, regardless of the block tree structure of the current block, a chrominance residual scaling enable flag indicating whether chrominance residual scaling is applied to the current block may be generated. When chrominance residual scaling is applied to the current picture, current slice, and / or current block, the value of the chrominance residual scaling enable flag may be 1.
[0520] In an embodiment, the chrominance residual scaling factor may be a single chrominance residual scaling factor.
[0521] In an embodiment, based on the value of the type information being 1, the APS may include an LMCS data field that includes LMCS parameters.
[0522] In an embodiment, the image information may include header information. The header information may include LMCS-related APS ID information. The LMCS-related APS ID information may represent the ID of the LMCS APS of the current picture or current block. For example, the header information may be a picture header (or slice header).
[0523] In an embodiment, the image information may include a sequence parameter set (SPS). The SPS may include a linear LMCS enable flag indicating whether linear LMCS is enabled.
[0524] In an embodiment, a minimum bin index (e.g., lmcs_min_bin_idx) and / or a maximum bin index (e.g., LmcsMaxBinIdx) may be derived based on information about the LMCS data. A first mapping value (LmcsPivot[lmcs_min_bin_idx]) may be derived based on the minimum bin index. A second mapping value (LmcsPivot[LmcsMaxBinIdx] or LmcsPivot[LmcsMaxBinIdx+1]) may be derived based on the maximum bin index. The value of the reconstructed luma samples (e.g., lumaSample in Tables 36 or 37) may be within the range of the first mapping value to the second mapping value. In one example, the values of all reconstructed luma samples may be within the range of the first mapping value to the second mapping value. In another example, the values of some samples of the reconstructed luma samples may be within the range of the first mapping value to the second mapping value.
[0525] In an embodiment, information about the LMCS data may include information about the linear LMCS and the LMCS data fields. The information about the linear LMCS may be referred to as information about the linear mapping. The LMCS data fields may include a linear LMCS flag indicating whether the linear LMCS is applied. When the value of the linear LMCS flag is 1, the predicted luminance samples of the mapping may be generated based on the information about the linear LMCS.
[0526] In an embodiment, the information about the linear LMCS may include information about a first pivot point (e.g., Figure 13 P1 in Figure 13 ) and information about a second pivot point (e.g.,
[0527] P2 in
[0528] ). For example, the input value and the mapped value of the first pivot point may be the minimum input value and the minimum mapped value, respectively. The input value and the mapped value of the second pivot point may be the maximum input value and the maximum mapped value, respectively. The input values between the minimum input value and the maximum input value may be linearly mapped.
[0529] In one embodiment, the image information includes information about the maximum input value and information about the maximum mapped value. The maximum input value is equal to the value of the information about the maximum input value (i.e., lmcs_max_input in Table 33). The maximum mapped value is equal to the value of the information about the maximum mapped value (i.e., lmcs_max_mapped in Table 33).
[0530] In one embodiment, generating the predicted luminance samples of the mapping includes: deriving a forward mapping scaling factor for the predicted luminance samples (i.e., ScaleCoeffSingle); and generating the predicted luminance samples of the mapping based on the forward mapping scaling factor. The forward mapping scaling factor may be a single factor for the predicted luminance samples.
[0531] In one embodiment, the forward mapping scaling factor may be derived based on at least one equation included in Table 36 and / or Table 38 above.
[0532] In one embodiment, the mapped predicted luma samples may be derived based on at least one equation included in Table 38 above.
[0533] In one embodiment, the encoding device may derive an inverse mapped scaling factor (i.e., InvScaleCoeffSingle) of the reconstructed luma samples (i.e., lumaSample). Additionally, the encoding device may generate modified reconstructed luma samples (i.e., invSample) based on the reconstructed luma samples and the inverse mapped scaling factor. The inverse mapped scaling factor may be a single factor of the reconstructed luma samples.
[0534] In one embodiment, the inverse mapped scaling factor may be derived using a piecewise index derived based on the reconstructed luma samples.
[0535] In one embodiment, the piecewise index may be derived based on Table 51 above. That is, the comparison process (lumaSample < LmcsPivot[idxYInv+1]) included in Table 51 may be iteratively performed from the piecewise index as the minimum bin index to the piecewise index as the maximum bin index.
[0536] In one embodiment, the inverse mapped scaling factor may be derived based on at least one equation included in Table 33, Table 34, Table 35, and Table 36 or Equation 11 or Equation 12 above.
[0537] In one embodiment, the modified reconstructed luma samples may be derived based on Equation 20, Equation 21, Table 39, and / or Table 40 above.
[0538] In one embodiment, the LMCS related information may include information about the number of bins (i.e., lmcs_num_bins_minus1 in Table 41) for deriving the predicted luma samples of the mapping. For example, the number of pivot points of the luma mapping may be set to be equal to the number of bins. In one example, the encoding device may generate the Δ input value and the Δ mapped value of the pivot points respectively according to the number of bins. In one example, the input value and the mapped value of the pivot points are derived based on the Δ input value (i.e., lmcs_delta_input_cw[i] in Table 41) and the Δ mapped value (i.e., lmcs_delta_mapped_cw[i] in Table 41), and the mapped predicted luma samples may be generated based on the input value (i.e., LmcsPivot_input[i] in Table 42) and the mapped value (i.e., LmcsPivot_mapped[i] in Table 42).
[0539] In one embodiment, the encoding device may derive an LMCSΔ codeword based on at least one LMCS codeword and an original codeword (OrgCW) included in the LMCS-related information, and may derive the mapped luminance prediction samples based on at least one LMCS codeword and the original codeword. In one example, the information regarding the linear mapping may include the information regarding the LMCSΔ codeword.
[0540] In one embodiment, at least one LMCS codeword may be derived based on the sum of the LMCSΔ codeword and the OrgCW. For example, OrgCW is (1 << BitDepthY) / 16, where BitDepthY represents the luminance bit depth. This embodiment may be based on Equation 12.
[0541] In one embodiment, at least one LMCS codeword may be derived based on the sum of the LMCSΔ codeword and OrgCW * (lmcs_max_bin_idx - lmcs_min_bin_idx + 1). For example, lmcs_max_bin_idx and lmcs_min_bin_idx are the maximum bin index and the minimum bin index respectively, and OrgCW may be (1 << BitDepthY) / 16. This embodiment may be based on Equation 15 and Equation 16.
[0542] In one embodiment, at least one LMCS codeword may be a multiple of 2.
[0543] In one embodiment, when the luminance bit depth (BitDepthY) of the reconstructed luminance samples is higher than 10, at least one LMCS codeword may be a multiple of 1 << (BitDepthY - 10).
[0544] In one embodiment, at least one LMCS codeword may be in the range from (OrgCW >> 1) to (OrgCW << 1) - 1.
[0545] In the above paragraphs, the information regarding the LMCS data may be the same as the information regarding the LMCS.
[0546] Figure 19 and Figure 20 Schematically shows an example of an image / video decoding method and related components according to an embodiment of this document. Figure 19 The method disclosed in Figure 3 may be executed by the decoding device shown in Figure 19S1900 can be executed by the entropy decoder 310 of the decoding device, S1910 and S1920 can be executed by the predictor 330 of the decoding device, S1930 can be executed by the residual processor 320 or the predictor 330 of the decoding device, and S1940 can be executed by the residual processor 320, the predictor 330, and / or the adder 340 of the decoding device. Figure 19 The method disclosed in
[0547] With reference to Figure 19 , the decoding device may receive / acquire video / image information (S1900). The video / image information may include information about the LMCS data and / or residual information. For example, the information about the LMCS data may include information about the luminance mapping (i.e., forward mapping, inverse mapping, linear mapping), information about the chrominance residual scaling, and / or an index related to the LMCS (or shaping, shaper) (i.e., the maximum bin index, the minimum bin index, the mapping index). The decoding device may receive / acquire the image / video information through the bitstream.
[0548] According to an embodiment of this document, the image / video information may include various information. For example, the image / video information may include the information disclosed in at least one of Tables 1 to 56 above.
[0549] The decoding device may derive the prediction mode of the current block in the current picture based on the prediction mode information (S1910). The decoding device may derive at least one of various modes disclosed in this document among the inter-frame prediction modes.
[0550] The decoding device may generate predicted luminance samples. The decoding device may derive the predicted luminance samples of the current block based on the prediction mode. In this case, various prediction methods disclosed in this document, such as inter-frame prediction or intra-frame prediction, may be applied.
[0551] The decoding device may generate predicted luminance samples (S1920). The decoding device may derive the predicted luminance samples of the current block based on the prediction mode. The decoding device may generate the predicted luminance samples by performing prediction on the original samples included in the current block.
[0552] The image information may include residual information. The decoding device may generate residual chrominance samples based on the residual information. Specifically, the decoding device may derive the quantized transform coefficients based on the residual information. The quantized transform coefficients may have a one-dimensional vector form based on the coefficient scan order. The decoding device may derive the transform coefficients based on the inverse quantization process of the quantized transform coefficients. The decoding device may derive the residual chrominance samples and / or the residual luminance samples based on the transform coefficients.
[0553] The decoding device may generate mapped predicted luminance samples (S1930). The decoding device may generate mapped predicted luminance samples based on the mapping process of the luminance samples. For example, the decoding device may derive the input value and the mapped value (output value) of the pivot point of the luminance mapping, and may generate mapped predicted luminance samples based on the input value and the mapped value. In one example, the decoding device may derive a (forward) mapping index (idxY) based on the first predicted luminance sample, and may generate the predicted luminance sample of the first mapping based on the input value and the mapped value of the pivot point corresponding to the mapping index. In other examples, a linear mapping (linear shaping, linear LMCS) may be used, and the predicted luminance sample of the mapping may be generated based on the forward mapping scaling factor derived from two pivot points in the linear mapping. Therefore, due to the linear mapping, the index derivation process may be omitted.
[0554] The decoding device may generate residual luminance samples based on the residual information. For example, the decoding device may derive the quantized transform coefficients based on the residual information. The quantized transform coefficients may have a one-dimensional vector form based on the coefficient scan order. The decoding device may derive the transform coefficients based on the dequantization process of the quantized transform coefficients. The decoding device may derive the residual samples based on the inverse transform process of the transform coefficients. The residual samples may include residual luminance samples and / or residual chrominance samples.
[0555] The decoding device may generate reconstructed luminance samples. The decoding device may generate reconstructed luminance samples based on the mapped predicted luminance samples. Specifically, the decoding device may sum the residual luminance samples and the mapped predicted luminance samples, and may generate reconstructed luminance samples based on the result of the summation.
[0556] The decoding device may generate modified reconstructed luminance samples. The decoding device may generate modified reconstructed luminance samples based on the reverse mapping process of the luminance samples. The decoding device may generate modified reconstructed luminance samples based on the information about the LMCS data and the reconstructed luminance samples (S1940). The decoding device may generate modified reconstructed luminance samples through the reverse mapping process of the reconstructed luminance samples.
[0557] The decoding device may generate scaled residual chrominance samples. Specifically, the decoding device may derive the chrominance residual scaling factor and generate scaled residual chrominance samples based on the chrominance residual scaling factor. Here, contrary to the encoding side, the chrominance residual scaling on the decoding side may be referred to as reverse chrominance residual scaling. Therefore, the chrominance residual scaling factor derived by the decoding device may be referred to as the reverse chrominance residual scaling factor, and reverse-scaled residual chrominance samples may be generated.
[0558] The decoding device may generate reconstructed chrominance samples. The decoding device may generate the reconstructed chrominance samples based on the scaled residual chrominance samples. Specifically, the decoding device may perform prediction processing on the chrominance component and may generate predicted chrominance samples. The decoding device may generate the reconstructed chrominance samples based on the sum of the predicted chrominance samples and the scaled residual chrominance samples.
[0559] In an embodiment, the image information may include LMCS APS. The LMCS APS may include an LMCS data field. The LMCS data field may include LMCS-related information. The LMCS codewords used in the mapping process and the inverse mapping process may be derived based on the LMCS data field. The maximum number of LMCS APSs may be a predetermined value. For example, the maximum number (predetermined value) of LMCS APSs may be four.
[0560] In an embodiment, during the mapping process of the luminance samples, a mapping index may be derived based on the predicted luminance samples, and a mapped predicted luminance sample may be generated using a first mapping value based on the mapping index.
[0561] In an embodiment, the LMCS-related information may include information about the bins for inverse mapping. Based on the information about the bins, a minimum bin index and a maximum bin index are derived. During the inverse mapping process of the luminance samples, an inverse mapping index is derived based on the bin indices from the minimum bin index to the maximum bin index based on the mapping values, and a modified reconstructed luminance sample is generated using a second mapping value based on the inverse mapping index.
[0562] In an embodiment, a segmentation index (e.g., idxYInv in Table 35, Table 36, or Table 37) may be identified based on information about the LMCS data. The decoding device may derive a chrominance residual scaling factor based on the segmentation index. The decoding device may generate scaled residual chrominance samples based on the residual chrominance samples and the chrominance residual scaling factor.
[0563] In one embodiment, when the current block has a single-tree structure or a dual-tree structure (when the current block has a separate tree structure, when the current block is encoded as a separate tree), a chrominance residual scaling enable flag indicating whether chrominance residual scaling is applied to the current block may be signaled. Alternatively, regardless of the block tree structure of the current block, a chrominance residual scaling enable flag indicating whether chrominance residual scaling is applied to the current block may be signaled. When chrominance residual scaling is applied to the current picture, current slice, and / or current block, the value of the chrominance residual scaling enable flag may be 1.
[0564] In an embodiment, the chrominance residual scaling factor may be a single chrominance residual scaling factor.
[0565] In an embodiment, based on the type information, the value is 1, and the APS may include an LMCS data field that includes LMCS parameters.
[0566] In an embodiment, the image information may include header information. The header information may include LMCS-related APS ID information. The LMCS-related APS ID information may represent the ID of the LMCS APS of the current picture or the current block. For example, the header information may be a picture header (or a slice header).
[0567] In an embodiment, the minimum bin index (e.g., lmcs_min_bin_idx) and / or the maximum bin index (e.g., LmcsMaxBinIdx) may be derived based on information about the LMCS data. The first mapping value (LmcsPivot[lmcs_min_bin_idx]) may be derived based on the minimum bin index. The second mapping value (LmcsPivot[LmcsMaxBinIdx] or LmcsPivot[LmcsMaxBinIdx + 1]) may be derived based on the maximum bin index. The values of the reconstructed luminance samples (e.g., lumaSample in Table 51 or Table 52) may be within the range from the first mapping value to the second mapping value. In one example, the values of all the reconstructed luminance samples may be within the range from the first mapping value to the second mapping value. In another example, the values of some of the reconstructed luminance samples may be within the range from the first mapping value to the second mapping value.
[0568] In an embodiment, the image information may include a sequence parameter set (SPS). The SPS may include a linear LMCS enable flag indicating whether linear LMCS is enabled.
[0569] In an embodiment, the chrominance residual scaling factor may be a single chrominance residual scaling factor.
[0570] In an embodiment, the information about the LMCS data may include information about linear LMCS and an LMCS data field. The information about linear LMCS may be referred to as information about linear mapping. The LMCS data field may include a linear LMCS flag indicating whether linear LMCS is applied. When the value of the linear LMCS flag is 1, the predicted luminance samples of the mapping may be generated based on the information about linear LMCS.
[0571] In an embodiment, the information about linear LMCS may include information about a first pivot point (e.g., Figure 13 P1 in Figure 13The information of P2) in it. For example, the input value and the mapped value of the first pivot point may be the minimum input value and the minimum mapped value respectively. The input value and the mapped value of the second pivot point may be the maximum input value and the maximum mapped value respectively. The input values between the minimum input value and the maximum input value may be linearly mapped.
[0572] In an embodiment, the image information may include the information about the maximum input value and the information about the maximum mapped value. The maximum input value may be the same as the value of the information about the maximum input value (e.g., lmcs_max_input in Table 33). The maximum mapped value may be the same as the value of the information about the maximum mapped value (e.g., lmcs_max_mapped in Table 33).
[0573] In an embodiment, the information about the linear mapping may include the information about the input Δ value of the second pivot point (e.g., lmcs_max_input_delta in Table 35) and the information about the mapped Δ value of the second pivot point (e.g., lmcs_max_mapped_delta in Table 35). The maximum input value may be derived based on the input Δ value of the second pivot point, and the maximum mapped value may be derived based on the mapped Δ value of the second pivot point.
[0574] In an embodiment, the maximum input value and the maximum mapped value may be derived based on at least one formula included in Table 36 above.
[0575] In one embodiment, generating a mapped predicted luminance sample includes: deriving a forward mapping scaling factor (i.e., ScaleCoeffSingle) of the predicted luminance sample; and generating the mapped predicted luminance sample based on the forward mapping scaling factor. The forward mapping scaling factor may be a single factor of the predicted luminance sample.
[0576] In one embodiment, a reverse mapping scaling factor may be derived using a segment index derived based on a reconstructed luminance sample.
[0577] In one embodiment, the segment index may be derived based on Table 51 above. That is, the comparison process (lumaSample < LmcsPivot[idxYInv+1]) included in Table 51 may be iteratively executed from the segment index as the minimum bin index to the segment index as the maximum bin index.
[0578] In one embodiment, the forward mapping scaling factor may be derived based on at least one formula included in Table 36 and / or Table 38 above.
[0579] In one embodiment, the mapped predicted luminance sample may be derived based on at least one formula included in Table 38 above.
[0580] In one embodiment, the decoding device may derive an inverse mapping scaling factor (i.e., InvScaleCoeffSingle) of the reconstructed luma sample (i.e., lumaSample). Additionally, the decoding device may generate a modified reconstructed luma sample (i.e., invSample) based on the reconstructed luma sample and the inverse mapping scaling factor. The inverse mapping scaling factor may be a single factor of the reconstructed luma sample.
[0581] In one embodiment, the inverse mapping scaling factor may be derived based on at least one of the equations included in Table 33, Table 34, Table 35, and Table 36 or Equation 11 or Equation 12 above.
[0582] In one embodiment, the modified reconstructed luma sample may be derived based on Equation 20, Equation 21, Table 39, and / or Table 40 above.
[0583] In one embodiment, the LMCS-related information may include information about the number of bins of the predicted luma sample used for deriving the mapping (i.e., lmcs_num_bins_minus1 in Table 41). For example, the number of pivot points of the luma mapping may be set to be equal to the number of bins. In one example, the decoding device may generate the Δ input value and the Δ mapped value of the pivot points according to the number of bins, respectively. In one example, the input value and the mapped value of the pivot points are derived based on the Δ input value (i.e., lmcs_delta_input_cw[i] in Table 41) and the Δ mapped value (i.e., lmcs_delta_mapped_cw[i] in Table 41), and the mapped predicted luma sample may be generated based on the input value (i.e., LmcsPivot_input[i] in Table 42) and the mapped value (i.e., LmcsPivot_mapped[i] in Table 42).
[0584] In one embodiment, the decoding device may derive an LMCS Δ codeword based on at least one LMCS codeword and the original codeword (OrgCW) included in the LMCS-related information, and may derive the mapped luma prediction sample based on at least one LMCS codeword and the original codeword. In one example, the information about the linear mapping may include information about the LMCS Δ codeword.
[0585] In one embodiment, at least one LMCS codeword may be derived based on the sum of the LMCS Δ codeword and OrgCW. For example, OrgCW is (1<<BitDepthY) / 16, where BitDepthY represents the luma bit depth. This embodiment may be based on Equation 14.
[0586] In one embodiment, at least one LMCS codeword may be derived based on the sum of the LMCSΔ codeword and OrgCW*(lmcs_max_bin_idx - lmcs_min_bin_idx + 1), where, for example, lmcs_max_bin_idx and lmcs_min_bin_idx are the maximum bin index and the minimum bin index respectively, and OrgCW may be (1<<BitDepthY) / 16. This embodiment may be based on Equation 15 and Equation 16.
[0587] In one embodiment, at least one LMCS codeword may be a multiple of 2.
[0588] In one embodiment, when the luma bit depth (BitDepthY) of the reconstructed luma samples is higher than 10, at least one LMCS codeword may be a multiple of 1<<(BitDepthY - 10).
[0589] In one embodiment, at least one LMCS codeword may be in the range from (OrgCW>>1) to (OrgCW<<1)-1.
[0590] In the above embodiments, the method is described based on a flowchart having a series of steps or blocks. The present disclosure is not limited to the order of the above steps or blocks. Some steps or blocks may occur simultaneously or in a different order from other steps or blocks as described above. In addition, those skilled in the art will understand that the steps shown in the above flowchart are not exclusive, may include additional steps, or one or more steps in the flowchart may be deleted without affecting the scope of the present disclosure.
[0591] The method according to the above embodiments of this document may be implemented in software form, and the encoding device and / or decoding device according to this document may be included, for example, in a device that performs image processing of a TV, a computer, a smart phone, a set-top box, a display device, etc.
[0592] When the embodiments in this document are implemented in software, the above methods can be implemented as modules (processes, functions, etc.) that execute the above functions. The modules can be stored in a memory and executed by a processor. The memory can be located inside or outside the processor and can be connected to the processor by various well-known means. The processor can include an application-specific integrated circuit (ASIC), other chip sets, logic circuits, and / or data processing devices. The memory can include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in the respective drawings can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information about the instructions or algorithms for implementation can be stored in a digital storage medium.
[0593] In addition, the decoding device and the encoding device to which the present disclosure is applied can be included in a multimedia broadcast transmission / reception device, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real-time communication device (e.g., video communication), a mobile streaming device, a storage medium, a camera, a VoD service providing device, an over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a videoconference video device, a vehicle user device (i.e., a vehicle user device, an aircraft user device, a ship user device, etc.), and a medical video device, and can be used to process video signals and data signals. For example, the over-the-top (OTT) video device can include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.
[0594] In addition, the processing method to which this document is applied can be generated in the form of a program to be executed by a computer and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the present disclosure can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices that store data readable by a computer system. For example, the computer-readable recording medium can include a BD, a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (i.e., transmission via the Internet). In addition, the bit stream generated by this encoding method can be stored in a computer-readable recording medium or can be transmitted via a wired / wireless communication network.
[0595] In addition, the embodiments of this document can be implemented as a computer program product according to program code, and the program code can be executed in a computer by the embodiments of this document. The program code can be stored on a computer-readable carrier.
[0596] Figure 21 An example of a content stream system to which the embodiments disclosed in this document can be applied is shown.
[0597] Referring to Figure 21 , the content stream system to which the embodiments of this document are applied may mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.
[0598] The encoding server compresses the content input from a multimedia input device (e.g., a smart phone, a camera, a video camera, etc.) into digital data to generate a bitstream and sends the bitstream to the streaming server. As another example, when the multimedia input device (e.g., a smart phone, a camera, a video camera, etc.) directly generates a bitstream, the encoding server can be omitted.
[0599] The bitstream can be generated by an encoding method or a bitstream generation method to which the embodiments of the present disclosure are applied, and during the process of sending or receiving the bitstream, the streaming server can temporarily store the bitstream.
[0600] The streaming server sends multimedia data to the user device via the network server based on the user's request, and the network server serves as a medium for informing the user of the service. When the user requests a desired service from the network server, the network server transmits it to the streaming server, and the streaming server sends the multimedia data to the user. In this case, the content stream system may include a separate control server. In this case, the control server is used to control the commands / responses between the devices in the content stream system.
[0601] The streaming server can receive content from the media storage device and / or the encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined time.
[0602] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smart watches, smart glasses, head-mounted displays), digital TVs, desktop computers, digital signage, etc. Each server in the content stream system can operate as a distributed server, and in this case, the data received from each server can be distributed.
[0603] Each server in the content stream system can operate as a distributed server, and in this case, the data received from each server can be distributed and processed.
[0604] The claims described herein can be combined in various ways. For example, the technical features of the method claims in this document can be combined and implemented as a device, and the technical features of the device claims in this document can be combined and implemented as a method. Additionally, the technical features of the method claims in this document and the technical features of the device claims in this document can be combined to be implemented as a device, and the technical features of the method claims in this document and the technical features of the device claims in this document can be combined and implemented as a method.
Claims
1. An apparatus for decoding image information, the apparatus comprising a memory and at least one processor coupled to the memory, the at least one processor being configured to: obtain the image information from a bitstream, the image information including prediction mode information and luminance mapping and chrominance scaling LMCS-related information; derive an inter prediction mode for a current block in a current picture based on the prediction mode information; generate predicted luminance samples based on the inter prediction mode; generate mapped predicted luminance samples based on a mapping process for the luminance samples; and generate modified reconstructed luminance samples based on an inverse mapping process for the luminance samples, wherein, the image information includes an LMCS adaptive parameter set APS; the LMCS APS includes an LMCS data field; the LMCS data field includes the LMCS-related information; the LMCS codewords used in the mapping process and the inverse mapping process are derived based on the LMCS data field; the maximum number of the LMCS APS is a predetermined value; the predetermined value is 4.
2. An apparatus for encoding image information, the apparatus comprising a memory and at least one processor coupled to the memory, the at least one processor being configured to: derive an inter prediction mode; generate predicted luminance samples based on the inter prediction mode; generate luminance mapping and chrominance scaling LMCS-related information; generate mapped predicted luminance samples based on a mapping process for the luminance samples; generate modified reconstructed luminance samples based on an inverse mapping process for the luminance samples; generate residual information based on the mapped predicted luminance samples and the modified reconstructed luminance samples; and and encode the image information including the LMCS-related information and the residual information, wherein, the image information includes an LMCS adaptive parameter set APS; the LMCS APS includes an LMCS data field; the LMCS data field includes the LMCS-related information; the LMCS codewords used in the mapping process and the inverse mapping process are derived based on the LMCS data field; the maximum number of the LMCS APS is a predetermined value; and the predetermined value is 4.
3. An apparatus for transmitting data for image information, the apparatus comprising: at least one processor configured to obtain a bitstream of the image information including luminance mapping and chrominance scaling LMCS-related information and residual information, wherein the LMCS-related information is generated by deriving an inter prediction mode, generating predicted luminance samples based on the inter prediction mode, and generating the LMCS-related information, and the residual information is generated by generating mapped predicted luminance samples based on a mapping process for the luminance samples, generating modified reconstructed luminance samples based on an inverse mapping process for the luminance samples, and generating residual information based on the mapped predicted luminance samples and the modified reconstructed luminance samples; and A transmitter configured to transmit the data of the bitstream including the image information, the image information including the LMCS-related information and the residual information, wherein, the image information includes a LMCS Adaptive Parameter Set (APS); the LMCS APS includes a LMCS data field; the LMCS data field includes the LMCS-related information; the LMCS codewords used in the mapping process and the inverse mapping process are derived based on the LMCS data field; the maximum number of the LMCS APSs is a predetermined value; and the predetermined value is 4.