Video or image coding based on luma mapping and chroma scaling
Luma mapping with chroma scaling (LMCS) enhances video/image coding efficiency by employing efficient filtering and linear mapping techniques, addressing the need for cost-effective compression of high-resolution media formats.
Patent Information
- Application Number
- JP2024231593
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-06-17
- Filing Date
- 2024-12-27
- Publication Date
- 2025-10-15
- Estimated Expiration
- 2040-06-17
AI Technical Summary
The increasing demand for high-resolution and high-quality images/videos, particularly in immersive media formats like VR and AR, necessitates the development of highly efficient image/video compression technologies to reduce transmission and storage costs while maintaining subjective/objective visual quality.
The implementation of luma mapping with chroma scaling (LMCS) procedures, including efficient filtering methods, restriction of LMCS codewords, use of a single chroma residual scaling factor, and linear mapping with explicit signaling of pivot points, along with flexible bin usage and simplified inverse mapping index derivation, to enhance video/image coding efficiency.
This approach increases overall image/video compression efficiency, enhances subjective/objective visual quality, minimizes resource/costs, and reduces latency and complexity in the LMCS procedure, facilitating hardware implementation.
Smart Images

Figure 0007755039000064 
Figure 0007755039000065 
Figure 0007755039000066
Abstract
Description
[Technical Field]
[0001] The present technology relates to video or image coding based on luma mapping and chroma scaling. [Background technology]
[0002] In recent years, demand for high-resolution, high-quality images / videos, such as 4K or 8K or higher UHD (Ultra High Definition) images / videos, has been increasing in various fields. As the resolution and quality of image / video data increases, the amount of information or bits to be transmitted increases relatively compared to existing image / video data. Therefore, when transmitting image data using existing media such as wired or wireless broadband lines or storing image / video data using existing storage media, transmission costs and storage costs increase.
[0003] In addition, interest in and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing in recent years, and the broadcast of images / videos with different image characteristics from real images, such as game images, has been increasing.
[0004] This necessitates the development of highly efficient image / video compression technologies to effectively compress and transmit, store, and play back high-resolution, high-quality image / video information that has the various characteristics described above.
[0005] In addition, to improve compression efficiency and enhance subjective / objective visual quality, a luma mapping with chroma scaling (LMCS) procedure is used, and there are discussions to reduce the computational complexity of the LMCS procedure. Summary of the Invention [Problem to be solved by the invention]
[0006] According to one embodiment of the present document, a method and apparatus for improving image / video coding efficiency is provided. [Means for solving the problem]
[0007] According to one embodiment of the present document, an efficient method and apparatus for applying filtering is provided.
[0008] According to one embodiment of the present document, an efficient method and apparatus for applying LMCS is provided.
[0009] According to one embodiment of this document, the LMCS codeword (or its range) can be restricted.
[0010] According to one embodiment of this document, a single chroma residual scaling factor signaled directly in the chroma scaling of the LMCS may be used.
[0011] According to one embodiment of this document, a linear mapping (linear LMCS) may be used.
[0012] According to one embodiment of this document, information regarding the pivot points required for linear mapping can be explicitly signaled.
[0013] According to one embodiment of this document, a flexible number of bins can be used for luma mapping.
[0014] According to one embodiment of this document, the procedure for deriving the inverse mapping index for luma samples may be simplified.
[0015] According to one embodiment of the present document, there is provided a video / image decoding method performed by a decoding device.
[0016] According to one embodiment of the present document, there is provided a decoding device for performing video / image decoding.
[0017] According to one embodiment of the present document, there is provided a video / image encoding method performed by an encoding device.
[0018] According to one embodiment of the present document, an encoding device for video / image encoding is provided.
[0019] According to one embodiment of the present document, there is provided a computer-readable digital storage medium having encoded video / image information stored thereon, the encoded video / image information being generated by the video / image encoding method disclosed in at least one of the embodiments of the present document.
[0020] According to one embodiment of the present document, there is provided a computer-readable digital storage medium having stored thereon encoded information or encoded video / image information that causes a decoding device to perform the video / image decoding method disclosed in at least one of the embodiments of the present document. [Effects of the Invention]
[0021] According to one embodiment of this document, the overall image / video compression efficiency can be increased.
[0022] According to one embodiment of this document, subjective / objective visual quality can be enhanced through efficient filtering.
[0023] According to one embodiment of this document, the LMCS procedure for image / video coding can be performed efficiently.
[0024] According to one embodiment of this document, the resources / costs (software or hardware) required for the LMCS procedure can be minimized.
[0025] According to one embodiment of this document, hardware implementation for the LMCS procedure can be facilitated.
[0026] According to one embodiment of this document, the division operations required to derive the LMCS codewords in the mapping (reshaping) can be eliminated or minimized through the LMCS codeword (or its range) restriction.
[0027] According to one embodiment of this document, the latency due to piecewise index identification can be eliminated through the use of a single chroma-residual scaling factor.
[0028] According to one embodiment of this document, the chroma residual scaling procedure can be performed without relying on (restoring) the luma block through the use of linear mapping in LMCS, and therefore latency in scaling can be eliminated.
[0029] According to one embodiment of this document, mapping efficiency in LMCS can be improved.
[0030] According to one embodiment of this document, the complexity of LMCS can be reduced through simplification of the procedure for deriving inverse mapping indexes or inverse scaling indexes for inverse mapping or chroma residual scaling for luma samples, and therefore, the video / image coding efficiency can be increased. [Brief explanation of the drawings]
[0031] [Figure 1] 1 illustrates schematically an example of a video / image coding system to which embodiments of the present document may be applied. [Figure 2] 1 is a diagram illustrating a schematic configuration of a video / image encoding device to which embodiments of the present document can be applied; [Figure 3] 1 is a diagram illustrating the configuration of a video / image decoding device to which the embodiments of the present document can be applied; [Figure 4] 1 illustrates an example of a video / image encoding method based on inter prediction. [Figure 5]1 illustrates an example of a video / image decoding method based on inter prediction. [Figure 6] 1 illustrates an exemplary inter-prediction procedure. [Figure 7] 1 shows an exemplary hierarchical structure for coded images / videos. [Figure 8] 1 illustrates an exemplary hierarchical structure of a CVS according to one embodiment of the present document. [Figure 9] 1 illustrates an exemplary LMCS structure according to an embodiment of the present document. [Figure 10] 1 illustrates an LMCS structure according to another embodiment of the present document. [Figure 11] 1 shows a graph illustrating an exemplary forward mapping. [Figure 12] 1 is a flow diagram illustrating a method for deriving a chroma-residual scaling index according to one embodiment of the present document; [Figure 13] 1 illustrates a linear fitting of pivot points according to an embodiment of the present document. [Figure 14] 1 illustrates an example of a linear reshaper according to an embodiment of the present document. [Figure 15] 1 illustrates an example of a linear forward mapping in accordance with an embodiment of the present document. [Figure 16] 1 illustrates an example of an inverse forward mapping in accordance with an embodiment of the present document. [Figure 17] 1 illustrates a schematic diagram of an example video / image encoding method and associated components according to an embodiment (or others) of the present document; [Figure 18] 1 illustrates a schematic diagram of an example video / image encoding method and associated components according to an embodiment (or others) of the present document; [Figure 19] 1 illustrates a schematic diagram of an example of an image / video decoding method and related components according to embodiments of the present document. [Figure 20] 1 illustrates a schematic diagram of an example of an image / video decoding method and related components according to embodiments of the present document. [Figure 21] 1 illustrates an example of a content streaming system to which embodiments disclosed herein may be applied. DETAILED DESCRIPTION OF THE INVENTION
[0032] Although the disclosure of this document may be modified in various ways and may have various embodiments, specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit the disclosure to the specific embodiments. The terms used in this document are used merely to describe specific embodiments and are not intended to limit the technical ideas of the embodiments in this document. The singular expressions include the plural expressions unless the context clearly indicates otherwise. In this document, terms such as "comprise" or "have" are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the document, and should be understood not to preclude the possibility of the presence or addition of one or more different features, numbers, steps, operations, components, parts, or combinations thereof.
[0033] Meanwhile, each component in the drawings described in this document is shown independently for the convenience of explaining the different characteristic functions, and does not mean that each component is realized by separate hardware or software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included in the scope of the disclosure of this document.
[0034] Hereinafter, the embodiments of the present document will be described with reference to the accompanying drawings. Hereinafter, the same reference numerals may be used for the same components in the drawings, and duplicate descriptions of the same components may be omitted.
[0035] FIG. 1 illustrates schematically an example of a video / image coding system in which embodiments of the present document may be applied.
[0036] As shown in Figure 1, a video / image coding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or a network.
[0037] The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may be a separate device or an external component.
[0038] A video source can acquire video / images through a video / image capture, synthesis, or generation process. A video source can include a video / image capture device and / or a video / image generation device. A video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, a virtual video / image can be generated via a computer, etc., in which case the video / image capture process can be replaced by a process in which related data is generated.
[0039] An encoding device can encode input video / images. The encoding device can perform a series of steps such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0040] The transmitter may transmit the encoded video / image information or data output in the form of a bitstream to a receiver of a receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter may include elements for generating a media file in a predetermined file format and elements for transmission via a broadcasting / communication network. The receiver may receive / extract the bitstream and transmit it to a decoding device.
[0041] The decoding device can decode the video / image by performing a series of steps such as inverse quantization, inverse transform, prediction, etc., which correspond to the operations of the encoding device.
[0042] The renderer can render the decoded video / image, and the rendered video / image can be displayed via the display unit.
[0043] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document may be applied to methods disclosed in the versatile video coding (VVC) standard. The methods / embodiments disclosed in this document may also be applied to methods disclosed in the essential video coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation audio video coding standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268).
[0044] This document presents various embodiments relating to video / image coding, which may be combined with one another unless otherwise stated.
[0045] In this document, video may refer to a collection of a series of images over time. A picture generally refers to a unit that shows one image at a specific time, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile may include one or more coding tree units (CTUs). One picture may consist of one or more slices / tiles. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs, and the rectangular region has a height equal to the height of the picture, and the width may be specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the width of the picture.A tile scan may indicate a specific sequential ordering of CTUs partitioning a picture, in which the CTUs are ordered consecutively in CTU raster scan in a tile, whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice includes an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that may be exclusively contained in a single NAL unit.
[0046] On the other hand, a picture can be divided into two or more sub-pictures, each of which can be a rectangular region of one or more slices within a picture.
[0047] A pixel or a pel can refer to the smallest unit that constitutes one picture (or image). A term corresponding to a pixel can also be used: "sample." A sample can generally indicate a pixel or a pixel value, can indicate only a pixel / pixel value of a luma component, or can indicate only a pixel / pixel value of a chroma component.
[0048] A unit may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. The term unit may be used interchangeably with terms such as block or area. In general, an M×N block may include a set (or array) of samples or transform coefficients consisting of M columns and N rows.
[0049] In this document, "A or B" may mean "only A," "only B," or "both A and B." In other words, in this document, "A or B" may be interpreted as "A and / or B." For example, in this document, "A, B or C" may mean "only A," "only B," "only C," or "any combination of A, B and C."
[0050] A slash ( / ) or a comma (comma) used in this document can mean "and / or." For example, "A / B" can mean "A and / or B." This means that "A / B" can mean "only A," "only B," or "both A and B." For example, "A, B, C" can mean "A, B, or C."
[0051] In this document, "at least one of A and B" can mean "only A," "only B," or "both A and B." Also, in this document, the expressions "at least one of A or B" and "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B."
[0052] Also, in this document, "at least one of A, B and C" can mean "only A," "only B," "only C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" can mean "at least one of A, B and C."
[0053] Furthermore, parentheses used in this document may mean "for example." Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" may be suggested as an example of "prediction." In other words, "prediction" in this document is not limited to "intra prediction," and "intra prediction" may be suggested as an example of "prediction." Furthermore, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" may be suggested as an example of "prediction."
[0054] Technical features individually described in one drawing in this document may be implemented individually or simultaneously.
[0055] 2 is a diagram for schematically illustrating the configuration of a video / image encoding device to which the embodiments of this document can be applied. Hereinafter, the encoding device may include an image encoding device and / or a video encoding device.
[0056] As shown in FIG. 2, the encoding device 200 may include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter predictor 221 and an intra predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. The image dividing unit 210, the predicting unit 220, the residual processing unit 230, the entropy encoding unit 240, the adding unit 250, and the filtering unit 260 may be configured by one or more hardware components (e.g., an encoder chipset or a processor) depending on the embodiment. Also, the memory 270 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0057] The image division unit 210 may divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) using a quad-tree, binary-tree, ternary-tree (QTBTTT) structure. For example, one coding unit may be divided into multiple coding units of deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and then the binary tree structure and / or ternary structure may be applied. Alternatively, the binary tree structure may be applied first. The coding procedure according to the present disclosure may be performed based on a final coding unit that is not further divided. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to image characteristics, or the coding unit may be recursively divided into coding units of lower depths as needed, and a coding unit of an optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be divided or partitioned from the final coding unit.The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0058] The term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block can refer to a set of samples or transform coefficients consisting of M columns and N rows. A sample generally refers to a pixel or pixel value, and can refer to only a pixel / pixel value of a luma component or only a pixel / pixel value of a chroma component. A sample can also be used as a term corresponding to one pixel or pel of a picture (or image).
[0059] The encoding apparatus 200 may subtract a prediction signal (predicted block, prediction sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from an input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, a unit in the encoder 200 that subtracts a prediction signal (predicted block, prediction sample array) from an input image signal (original block, original sample array) may be referred to as a subtraction unit 231. The prediction unit may perform prediction on a current block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied in units of the current block or CU. The prediction unit may generate various information related to prediction, such as prediction mode information, and transmit the information to the entropy encoding unit 240, as will be described later in the description of each prediction mode. The prediction information can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.
[0060] The intra prediction unit 222 may predict the current block by referring to samples in the current picture. The referenced samples may be located adjacent to or distant from the current block depending on the prediction mode. In intra prediction, prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, DC mode and planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the granularity of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 222 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.
[0061] The inter prediction unit 221 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (such as L0 prediction, L1 prediction, or Bi prediction). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (col CU), or the like, and the reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter predictor 221 may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of a skip mode or a merge mode, the inter predictor 221 may use motion information of neighboring blocks as motion information of the current block. In the case of the skip mode, unlike the merge mode, a residual signal may not be transmitted.In the case of motion vector prediction (MVP) mode, the motion vector of the current block can be indicated by using the motion vector of the neighboring block as a motion vector predictor and signaling the motion vector difference.
[0062] The prediction unit 220 may generate a prediction signal based on various prediction methods, which will be described later. For example, the prediction unit may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as combined inter and intra prediction (CIIP). The prediction unit may also use an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content image / video coding, such as games, such as screen content coding (SCC). IBC basically performs prediction within a current picture, but may be similar to inter prediction in deriving a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, sample values within a picture may be signaled based on information about a palette table and a palette index.
[0063] The prediction signal generated by the prediction unit (including the inter prediction unit 221 and / or the intra prediction unit 222) may be used to generate a reconstructed signal or a residual signal. The transform unit 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, GBT refers to a transform obtained from a graph representing relationship information between pixels. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. The transform process may be applied to square pixel blocks of the same size or non-square variable-size blocks.
[0064] The quantizer 233 quantizes the transform coefficients and transmits the quantized signal to the entropy encoder 240. The entropy encoder 240 encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 233 may rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoder 240 may perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). In addition to the quantized transform coefficients, the entropy encoder 240 may also encode information required for video / image restoration (e.g., values of syntax elements) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of network abstraction layer (NAL) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. In this document, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the above-described encoding procedure and included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcasting network and / or a communication network, and the digital storage medium can include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting the signal output from the entropy encoding unit 240 and / or a storage unit (not shown) for storing the signal can be configured as an internal / external element of the encoding device 200, or the transmitter can be included in the entropy encoding unit 240.
[0065] The quantized transform coefficients output from the quantizer 233 may be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) may be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantizer 234 and the inverse transformer 235. The adder 155 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter predictor 221 or the intra predictor 222. When there is no residual for the current block, such as when skip mode is applied, a predicted block may be used as the reconstructed block. The adder 250 may be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next block to be processed in the current picture, or may be used for inter prediction of the next picture after filtering, as described below.
[0066] Meanwhile, luma mapping with chroma scaling (LMCS) can be applied during picture encoding and / or reconstruction.
[0067] The filtering unit 260 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 260 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filtering unit 260 may generate various information related to filtering and transmit it to the entropy encoding unit 240, as will be described later in connection with each filtering method. The filtering information may be encoded by the entropy encoding unit 240 and output in the form of a bitstream.
[0068] The modified reconstructed picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. When inter prediction is applied through this, the encoding device can avoid prediction mismatch between the encoding device 100 and the decoding device, and can also improve coding efficiency.
[0069] The DPB of the memory 270 may store the modified reconstructed picture to be used as a reference picture in the inter predictor 221. The memory 270 may store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter predictor 221 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 270 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 222.
[0070] 3 is a diagram for explaining the configuration of a video / image decoding device to which the embodiments of this document can be applied. Hereinafter, the decoding device may include an image decoding device and / or a video decoding device.
[0071] As shown in FIG. 3, the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an intra predictor 331 and an inter predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. Depending on the embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be configured as a single hardware component (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) or may be configured as a digital storage medium. The hardware components may further include a memory 360 as an internal / external component.
[0072] When a bitstream including video / image information is input, the decoding device 300 can reconstruct an image corresponding to the process in which the video / image information was processed by the encoding device of FIG. 3. For example, the decoding device 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing unit applied by the encoding device. Therefore, the processing unit for decoding can be, for example, a coding unit, and the coding unit can be divided from a coding tree unit or a maximal coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. The reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a playback device.
[0073] The decoding device 300 may receive a signal output from the encoding device of FIG. 3 in the form of a bitstream, and the received signal may be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 may parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The decoding device may further decode pictures based on the information on the parameter sets and / or the general constraint information. Signaling / received information and / or syntax elements, which will be described later in this document, may be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 may decode information in a bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transform coefficients related to residuals. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using information on the syntax element to be decoded, decode information on adjacent and current blocks, or information on symbols / bins decoded in previous steps, predicts the occurrence probability of the bins based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method may update the context model using information on the decoded symbols / bins for the context model of the next symbol / bin.Among the information decoded by the entropy decoding unit 310, information related to prediction is provided to a prediction unit (inter prediction unit 332 and intra prediction unit 331), and residual values entropy decoded by the entropy decoding unit 310, i.e., quantized transform coefficients and related parameter information, may be input to the residual processing unit 320. The residual processing unit 320 may derive a residual signal (residual block, residual sample, residual sample array). In addition, among the information decoded by the entropy decoding unit 310, information related to filtering may be provided to the filtering unit 350. Meanwhile, a receiving unit (not shown) that receives a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310. Meanwhile, the decoding device according to this document may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, the inverse transform unit 322, the addition unit 340, the filtering unit 350, the memory 360, the inter prediction unit 332, and the intra prediction unit 331.
[0074] The inverse quantization unit 321 may inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 321 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit 321 may inverse quantize the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.
[0075] The inverse transform unit 322 performs inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0076] The prediction unit may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on the prediction information output from the entropy decoding unit 310, and may determine a specific intra / inter prediction mode.
[0077] The prediction unit 320 may generate a prediction signal based on various prediction methods, which will be described later. For example, the prediction unit may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as combined inter and intra prediction (CIIP). The prediction unit may also use an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content image / video coding, such as games, such as screen content coding (SCC). IBC basically performs prediction within a current picture, but may be similar to inter prediction in deriving a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, information regarding a palette table and a palette index may be included in the video / picture information and signaled.
[0078] The intra prediction unit 331 can predict a current block by referring to samples in a current picture. The referenced samples can be located adjacent to or distant from the current block depending on the prediction mode. In intra prediction, prediction modes can include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 can also determine a prediction mode to be applied to the current block using prediction modes applied to neighboring blocks.
[0079] The inter prediction unit 332 may derive a predicted block for the current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter prediction unit 332 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes, and the prediction information may include information indicating the inter prediction mode for the current block.
[0080] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to a predicted signal (predicted block, predicted sample array) output from a prediction unit (including an inter prediction unit 332 and / or an intra prediction unit 331). When there is no residual for the current block, such as when a skip mode is applied, the predicted block can be used as a reconstructed block.
[0081] The adder 340 may be referred to as a reconstruction unit or a reconstruction block generator. The generated reconstruction signal may be used for intra prediction of a next block to be processed in the current picture, may be output after filtering as described below, or may be used for inter prediction of a next picture.
[0082] Meanwhile, LMCS (luma mapping with chroma scaling) can be applied during the picture decoding process.
[0083] The filtering unit 350 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 350 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may transmit the modified reconstructed picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.
[0084] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter predictor 332. The memory 360 can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information can be transmitted to the inter predictor 260 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 360 can store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 331.
[0085] In this specification, the embodiments described for the filtering unit 260, inter prediction unit 221, and intra prediction unit 222 of the encoding device 200 can also be applied identically or correspondingly to the filtering unit 350, inter prediction unit 332, and intra prediction unit 331 of the decoding device 300, respectively.
[0086] As described above, prediction is performed to improve compression efficiency when performing video coding. Through this, a predicted block including predicted samples for a current block, which is a block to be coded, can be generated. Here, the predicted block includes predicted samples in the spatial domain (or pixel domain). The predicted block is derived in the same way by an encoding device and a decoding device. The encoding device signals information (residual information) regarding the residual between the original block and the predicted block, rather than the original sample values of the original block, to the decoding device, thereby improving image coding efficiency. The decoding device derives a residual block including residual samples based on the residual information, combines the residual block with the predicted block to generate a reconstructed block including reconstructed samples, and generates a reconstructed picture including the reconstructed block.
[0087] The residual information may be generated through a transform and quantization procedure. For example, an encoding device may derive a residual block between the original block and the predicted block, perform a transform procedure on residual samples (residual sample array) included in the residual block to derive transform coefficients, and perform a quantization procedure on the transform coefficients to derive quantized transform coefficients, thereby signaling the related residual information to a decoding device (via a bitstream). Here, the residual information may include information such as value information, position information, transform technique, transform kernel, and quantization parameter of the quantized transform coefficients. The decoding device may derive residual samples (or residual blocks) by performing an inverse quantization / inverse transform procedure based on the residual information. The decoding device may generate a reconstructed picture based on the predicted block and the residual block. The encoding device may also derive a residual block by inverse quantizing / inverse transforming quantized transform coefficients for reference for inter-prediction of a future picture, and generate a reconstructed picture based on the residual block.
[0088] In this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When the quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When the transform / inverse transform is omitted, the transform coefficients may also be referred to as coefficients or residual coefficients, or may still be referred to as transform coefficients for the sake of uniformity of expression.
[0089] In this document, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficient(s), and the information about the transform coefficient(s) may be signaled via residual coding syntax. Transform coefficients may be derived based on the residual information (or information about the transform coefficient(s), and scaled transform coefficients may be derived through an inverse transform (scaling) of the transform coefficients. Residual samples may be derived based on an inverse transform (transform) of the scaled transform coefficients. This may be similarly applied / expressed in other parts of this document.
[0090] Intra prediction may refer to a prediction that generates prediction samples for a current block based on reference samples in a picture to which the current block belongs (hereinafter referred to as the current picture). When intra prediction is applied to the current block, neighboring reference samples used for intra prediction of the current block may be derived. The neighboring reference samples of the current block may include samples adjacent to the left boundary and bottom-left neighboring samples of the current block having a size of nW×nH, a total of 2×nH samples adjacent to the top boundary and top-right neighboring samples of the current block, and one sample adjacent to the top-left neighboring sample of the current block. Alternatively, the neighboring reference samples of the current block may include upper neighboring samples of multiple columns and left neighboring samples of multiple rows. In addition, the neighboring reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block having a size of nW×nH, a total of nW samples adjacent to the bottom boundary of the current block, and one sample adjacent to the bottom-right of the current block.
[0091] However, some of the neighboring reference samples of the current block may not yet be decoded or may not be available. In this case, the decoder may substitute the unavailable samples as available samples to construct neighboring reference samples to be used for prediction, or may construct neighboring reference samples to be used for prediction through interpolation of available samples.
[0092] When neighboring reference samples are derived, (i) a predicted sample can be derived based on an average or interpolation of neighboring reference samples of the current block, or (ii) the predicted sample can be derived based on a reference sample that exists in a specific (prediction) direction with respect to the predicted sample among the neighboring reference samples of the current block. Case (i) can be called a non-directional mode or a non-angular mode, and case (ii) can be called a directional mode or an angular mode.
[0093] Furthermore, the prediction sample may be generated by interpolating a first neighboring sample located in the prediction direction of the intra prediction mode of the current block and a second neighboring sample located in the opposite direction to the prediction direction based on the prediction sample of the current block among the neighboring reference samples. This case may be called linear interpolation intra prediction (LIP). Alternatively, a chroma prediction sample may be generated based on a luma sample using a linear model. This case may be called LM mode.
[0094] Alternatively, a temporary predicted sample of the current block may be derived based on filtered neighboring reference samples, and the predicted sample of the current block may be derived by weighted summing the temporary predicted sample and at least one reference sample derived according to the intra prediction mode from the existing neighboring reference samples, i.e., non-filtered neighboring reference samples. The above case may be called Position Dependent Intra Prediction (PDPC).
[0095] In addition, intra-prediction coding can be performed by selecting a reference sample line with the highest prediction accuracy from among multiple adjacent reference sample lines of the current block, deriving a predicted sample using a reference sample located in the prediction direction of the corresponding line, and signaling the reference sample line used at this time to a decoding device. This case can be called multi-reference line intra-prediction or MRL-based intra-prediction.
[0096] In addition, the current block may be divided into vertical or horizontal sub-partitions, and intra prediction may be performed based on the same intra prediction mode, and neighboring reference samples may be derived and used in units of the sub-partitions. That is, in this case, the intra prediction mode for the current block is applied to the sub-partitions in the same way, and neighboring reference samples may be derived and used in units of the sub-partitions, thereby improving intra prediction performance in some cases. This prediction method may be called ISP (intra sub-partitions)-based intra prediction.
[0097] The above-described intra prediction methods may be referred to as intra prediction types, distinguished from intra prediction modes. The intra prediction types may be referred to by various terms, such as intra prediction techniques or additional intra prediction modes. For example, the intra prediction types (or additional intra prediction modes, etc.) may include at least one of the above-described LIP, PDPC, MRL, and ISP. A general intra prediction method excluding specific intra prediction types, such as LIP, PDPC, MRL, and ISP, may be referred to as a normal intra prediction type. The normal intra prediction type may be generally applied when the above-described specific intra prediction types are not applied, and prediction may be performed based on the above-described intra prediction modes. Meanwhile, post-processing filtering may be performed on the derived prediction samples, if necessary.
[0098] Specifically, the intra prediction procedure may include an intra prediction mode / type determination step, a neighboring reference sample derivation step, and an intra prediction mode / type-based prediction sample derivation step. If necessary, a post-processing filtering step may be performed on the derived prediction samples.
[0099] When intra prediction is applied, the intra prediction mode to be applied to the current block may be determined using the intra prediction mode of a neighboring block. For example, the decoding device may select one of most probable mode (MPM) candidates in an MPM (most probable mode) list derived based on the intra prediction mode of a neighboring block (e.g., a left and / or upper neighboring block) of the current block and additional candidate modes based on the received MPM index, or may select one of the remaining intra prediction modes not included in the MPM candidates (and planar mode) based on remaining intra prediction mode information. The MPM list may be configured with or without including a planar mode as a candidate. For example, if the MPM list includes a planar mode as a candidate, the MPM list may have six candidates, and if the MPM list does not include a planar mode as a candidate, the MPM list may have five candidates. If the MPM list does not include a planar mode as a candidate, a not planar flag (e.g., intra_luma_not_planar_flag) indicating that the intra prediction mode of the current block is not a planar mode may be signaled. For example, the MPM flag may be signaled first, and the MPM index and the not planar flag may be signaled if the MPM flag has a value of 1. Also, the MPM index may be signaled if the not planar flag has a value of 1. Here, the reason why the MPM list is configured not to include a planar mode as a candidate is that, rather than meaning that the planar mode is not an MPM, the planar mode is always considered as an MPM, and therefore the flag (not planar flag) is signaled first to first confirm whether the mode is a planar mode.
[0100] For example, whether the intra prediction mode applied to the current block is among the MPM candidates (and planar mode) or among the remaining modes may be indicated based on an MPM flag (e.g., intra_luma_mpm_flag). A value of 1 in the MPM flag may indicate that the intra prediction mode for the current block is among the MPM candidates (and planar mode), and a value of 0 in the MPM flag may indicate that the intra prediction mode for the current block is not among the MPM candidates (and planar mode). A value of 0 in the not planar flag (e.g., intra_luma_not_planar_flag) may indicate that the intra prediction mode for the current block is planar mode, and a value of 1 in the not planar flag may indicate that the intra prediction mode for the current block is not planar mode. The MPM index may be signaled in the form of an mpm_idx or intra_luma_mpm_idx syntax element, and the remaining intra prediction mode information may be signaled in the form of a rem_intra_luma_pred_mode or intra_luma_mpm_remainder syntax element. For example, the remaining intra prediction mode information may index the remaining intra prediction modes not included in the MPM candidates (and planar modes) among all intra prediction modes in order of prediction mode number, and point to one of them. The intra prediction mode is an intra prediction mode for a luma component (sample). Hereinafter, the intra prediction mode information may include at least one of the MPM flag (e.g., intra_luma_mpm_flag), the not planar flag (e.g., intra_luma_not_planar_flag), the MPM index (e.g., mpm_idx or intra_luma_mpm_idx), and the remaining intra prediction mode information (rem_intra_luma_pred_mode or intra_luma_mpm_remainder).In this document, the MPM list may be referred to by various terms such as an MPM candidate list, candModeList, etc. If MIP is applied to the current block, a separate mpm flag (e.g., intra_mip_mpm_flag), mpm index (e.g., intra_mip_mpm_idx), and remaining intra-prediction mode information (e.g., intra_mip_mpm_remainder) for MIP may be signaled, and the not planar flag is not signaled.
[0101] That is, in general, when an image is divided into blocks, a current block to be coded and neighboring blocks have similar image characteristics. Therefore, the current block and neighboring blocks are likely to have the same or similar intra prediction modes. Therefore, an encoder can use the intra prediction modes of neighboring blocks to encode the intra prediction mode of the current block.
[0102] For example, the encoder / decoder may construct an MPM (Most Probable Modes) list for the current block. The MPM list may also be referred to as an MPM candidate list. Here, MPM may refer to a mode used to improve coding efficiency by considering similarities between the current block and neighboring blocks during intra-prediction mode coding. As described above, the MPM list may be constructed to include or exclude the planar mode. For example, if the MPM list includes the planar mode, the number of candidates in the MPM list is six. If the MPM list does not include the planar mode, the number of candidates in the MPM list is five.
[0103] The encoder / decoder can construct an MPM list containing five or six MPMs.
[0104] To construct the MPM list, three types of modes can be considered: default intra modes, neighbor intra modes, and derived intra modes.
[0105] For the neighboring intra mode, two neighboring blocks can be considered: a left neighboring block and an upper neighboring block.
[0106] As mentioned above, if the MPM list is configured not to include a planar mode, the planar mode is excluded from the list, and the number of MPM list candidates can be set to five.
[0107] In addition, among the intra prediction modes, the non-directional mode (or non-angular mode) may include a DC mode based on the average of neighboring reference samples of the current block or a planar mode based on interpolation.
[0108] When inter prediction is applied, a prediction unit of an encoding / decoding device may perform inter prediction on a block-by-block basis to derive predicted samples. Inter prediction may refer to prediction derived in a manner dependent on data elements (e.g., sample values or motion information) of pictures other than the current picture. When inter prediction is applied to a current block, a predicted block (prediction sample array) for the current block may be derived based on a reference block (reference sample array) identified by a motion vector in a reference picture indicated by a reference picture index. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, motion information of the current block may be predicted on a block, sub-block, or sample-by-block basis based on correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction type (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). When inter prediction is applied, neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporally neighboring block may be the same or different. The temporally neighboring block may be called a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporally neighboring block may be called a collocated picture (colPic). For example, a candidate list of motion information may be constructed based on the neighboring blocks of the current block, and flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled.Inter prediction is performed based on various prediction modes. For example, in skip mode and merge mode, the motion information of the current block may be the same as the motion information of the selected neighboring block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected neighboring block is used as a motion vector predictor, and a motion vector difference may be signaled. In this case, the motion vector of the current block may be derived using the sum of the motion vector predictor and the motion vector difference.
[0109] The motion information may include L0 motion information and / or L1 motion information depending on the inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction may be referred to as an L0 motion vector or MVL0, and a motion vector in the L1 direction may be referred to as an L1 motion vector or MVL1. Prediction based on an L0 motion vector may be referred to as L0 prediction, prediction based on an L1 motion vector may be referred to as L1 prediction, and prediction based on both the L0 motion vector and the L1 motion vector may be referred to as bi-prediction (Bi) prediction. Here, an L0 motion vector may indicate a motion vector associated with a reference picture list L0 (L0), and an L1 motion vector may indicate a motion vector associated with a reference picture list L1 (L1). The reference picture list L0 may include pictures that are earlier in output order than the current picture as reference pictures, and the reference picture list L1 may include pictures that are later in output order than the current picture. The previous picture may be called a forward (reference) picture, and the subsequent picture may be called a backward (reference) picture. The reference picture list L0 may further include, as reference pictures, pictures that are subsequent to the current picture in output order. In this case, the previous picture may be indexed first in the reference picture list L0, and the subsequent picture may be indexed thereafter. The reference picture list L1 may further include, as reference pictures, pictures that are subsequent to the current picture in output order. In this case, the subsequent picture may be indexed first in the reference picture list L1, and the previous picture may be indexed thereafter. Here, the output order may correspond to a picture order count (POC) order.
[0110] FIG. 4 shows an example of a video / image encoding method based on inter prediction.
[0111] The encoding apparatus performs inter prediction on a current block (S400). The encoding apparatus may derive an inter prediction mode and motion information of the current block and generate a predicted sample for the current block. Here, the inter prediction mode determination, motion information derivation, and predicted sample generation procedures may be performed simultaneously, or one procedure may be performed before the other procedures. For example, the inter prediction unit of the encoding apparatus may include a prediction mode determination unit, a motion information derivation unit, and a predicted sample derivation unit. The prediction mode determination unit may determine a prediction mode for the current block, the motion information derivation unit may derive motion information for the current block, and the predicted sample derivation unit may derive a predicted sample for the current block. For example, the inter prediction unit of the encoding apparatus may search for a block similar to the current block within a certain region (search region) of a reference picture through motion estimation and derive a reference block whose difference from the current block is minimum or equal to or less than a certain criterion. Based on this, a reference picture index indicating a reference picture in which the reference block is located can be derived, and a motion vector can be derived based on a position difference between the reference block and the current block. The encoding device can determine a mode to be applied to the current block from various prediction modes. The encoding device can compare RD costs for the various prediction modes to determine an optimal prediction mode for the current block.
[0112] For example, when a skip mode or a merge mode is applied to the current block, the encoding device may construct a merge candidate list (described below) and derive a reference block whose difference from the current block is minimum or equal to or less than a certain criterion among reference blocks indicated by merge candidates included in the merge candidate list. In this case, a merge candidate associated with the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and signaled to the decoding device. Motion information of the current block may be derived using motion information of the selected merge candidate.
[0113] As another example, when the (A)MVP mode is applied to the current block, the encoding apparatus may construct an (A)MVP candidate list (described below) and use a motion vector of an MVP (motion vector predictor) candidate selected from the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. In this case, for example, a motion vector pointing to a reference block derived by the above-described motion estimation may be used as the motion vector of the current block, and the MVP candidate having the smallest difference from the motion vector of the current block may be the selected MVP candidate. A motion vector difference (MVD), which is a difference obtained by subtracting the MVP from the motion vector of the current block, may be derived. In this case, information regarding the MVD may be signaled to the decoding apparatus. Furthermore, when the (A)MVP mode is applied, the value of the reference picture index may be configured as reference picture index information and separately signaled to the decoding apparatus.
[0114] The encoding apparatus may derive residual samples based on the predicted samples (S410) by comparing the original samples of the current block with the predicted samples.
[0115] The encoding apparatus encodes image information including prediction information and residual information (S420). The encoding apparatus may output the encoded image information in the form of a bitstream. The prediction information is information related to the prediction procedure and may include prediction mode information (e.g., skip flag, merge flag, or mode index) and information about motion information. The information about the motion information may include candidate selection information (e.g., merge index, MVP flag, or MVP index) for deriving a motion vector. The information about the motion information may also include information about the above-mentioned MVD and / or reference picture index information. The information about the motion information may also include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information about the residual sample. The residual information may also include information about quantized transform coefficients for the residual sample.
[0116] The output bitstream can be stored in a (digital) storage medium and transmitted to the decoding device, or can be transmitted to the decoding device via a network.
[0117] Meanwhile, as described above, the encoding apparatus can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. This is because the encoding apparatus derives the same prediction result as that performed by the decoding apparatus, thereby improving coding efficiency. Therefore, the encoding apparatus can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in memory and use it as a reference picture for inter prediction. As described above, an in-loop filtering procedure can be further applied to the reconstructed picture.
[0118] A video / image decoding procedure based on inter prediction may generally include, for example:
[0119] FIG. 5 shows an example of a video / image decoding method based on inter prediction.
[0120] As shown in Figure 5, the decoding device may perform operations corresponding to those performed by the encoding device. The decoding device may perform prediction on the current block based on received prediction information and derive predicted samples.
[0121] Specifically, the decoding device may determine a prediction mode for the current block based on received prediction information (S500). The decoding device may determine which inter-prediction mode is applied to the current block based on prediction mode information in the prediction information.
[0122] For example, it may determine whether the merge mode or (A)MVP mode is applied to the current block based on the merge flag. Alternatively, it may select one of various inter prediction mode candidates based on the mode index. The inter prediction mode candidates may include skip mode, merge mode, and / or (A)MVP mode, or may include various inter prediction modes described below.
[0123] The decoding device derives motion information of the current block based on the determined inter prediction mode (S510). For example, when a skip mode or a merge mode is applied to the current block, the decoding device may construct a merge candidate list (described below) and select one merge candidate from among the merge candidates included in the merge candidate list. The selection may be made based on the selection information (merge index) described above. The motion information of the selected merge candidate may be used to derive motion information of the current block. The motion information of the selected merge candidate may be used as the motion information of the current block.
[0124] As another example, when the (A)MVP mode is applied to the current block, the decoding apparatus may construct an (A)MVP candidate list (described below) and use a motion vector predictor (MVP) selected from among the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. The selection may be performed based on the selection information (MVP flag or MVP index) described above. In this case, the MVD of the current block may be derived based on information related to the MVD, and the motion vector of the current block may be derived based on the MVP of the current block and the MVD. Furthermore, the decoding apparatus may derive a reference picture index of the current block based on the reference picture index information. A picture pointed to by the reference picture index in the reference picture list for the current block may be derived as a reference picture referenced for inter-prediction of the current block.
[0125] On the other hand, as will be described later, the motion information of the current block may be derived without constructing a candidate list, and in this case, the motion information of the current block may be derived according to a procedure disclosed in the prediction mode section, which will be described later. In this case, the candidate list construction as described above may be omitted.
[0126] The decoding device may generate predictive samples for the current block based on the motion information of the current block (S520). In this case, the reference picture may be derived based on a reference picture index of the current block, and the predictive samples of the current block may be derived using samples of a reference block pointed to in the reference picture by the motion vector of the current block. In this case, as described below, a predictive sample filtering procedure may further be performed on all or some of the predictive samples of the current block, depending on the circumstances.
[0127] For example, the inter-prediction unit of the decoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit, and may determine a prediction mode for the current block based on prediction mode information received by the prediction mode determination unit, derive motion information (such as a motion vector and / or a reference picture index) of the current block based on information regarding the motion information received by the motion information derivation unit, and derive a prediction sample of the current block by the prediction sample derivation unit.
[0128] The decoding device generates residual samples for the current block based on the received residual information (S530). The decoding device generates reconstructed samples for the current block based on the predicted samples and the residual samples, and can generate a reconstructed picture based on the reconstructed samples (S540). Thereafter, an in-loop filtering procedure, etc., can be further applied to the reconstructed picture, as described above.
[0129] FIG. 6 exemplarily illustrates an inter prediction procedure.
[0130] 6, as described above, the inter prediction procedure may include an inter prediction mode determination step, a motion information deriving step according to the determined prediction mode, and a prediction (prediction sample generation) step based on the derived motion information. The inter prediction procedure may be performed by an encoding device and a decoding device, as described above. In this document, a coding device may include an encoding device and / or a decoding device.
[0131] As shown in FIG. 6, the coding apparatus determines an inter prediction mode for a current block (S600). Various inter prediction modes may be used for predicting a current block in a picture. For example, various modes may be used, such as merge mode, skip mode, motion vector prediction (MVP) mode, affine mode, sub-block merge mode, and merge with MVD (MMVD) mode. Decoder side motion vector refinement (DMVR) mode, adaptive motion vector resolution (AMVR) mode, bi-prediction with CU-level weight (BCW), bi-directional optical flow (BDOF), etc. may be used additionally or alternatively as additional modes. Affine mode may also be referred to as affine motion prediction mode. MVP mode may also be referred to as advanced motion vector prediction mode. In this document, some modes and / or motion information candidates derived by some modes may be included as one of the motion information-related candidates of other modes. For example, an HMVP candidate can be added as a merge candidate in the merge / skip mode, or as an MVP candidate in the MVP mode. When the HMVP candidate is used as a motion information candidate in the merge mode or skip mode, the HMVP candidate can be referred to as an HMVP merge candidate.
[0132] Prediction mode information indicating the inter prediction mode of a current block may be signaled from an encoding apparatus to a decoding apparatus. The prediction mode information may be included in a bitstream and received by the decoding apparatus. The prediction mode information may include index information indicating one of a plurality of candidate modes. Alternatively, the inter prediction mode may be indicated through hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, a skip flag may be signaled to indicate whether a skip mode is applied, and if the skip mode is not applied, a merge flag may be signaled to indicate whether a merge mode is applied, and if the merge mode is not applied, an MVP mode may be applied, or a flag for additional classification may be further signaled. The affine mode may be signaled as an independent mode or as a mode dependent on the merge mode or MVP mode. For example, the affine mode may include affine merge mode and affine MVP mode.
[0133] Meanwhile, information indicating whether the above-mentioned list0 (L0) prediction, list1 (L1) prediction, or bi-prediction is used for the current block (current coding unit) may be signaled. This information may be referred to as motion prediction direction information, inter-prediction direction information, or inter-prediction indication information, and may be configured / encoded / signaled in the form of, for example, an inter_pred_idc syntax element. That is, the inter_pred_idc syntax element may indicate whether the above-mentioned list0 (L0) prediction, list1 (L1) prediction, or bi-prediction is used for the current block (current coding unit). For convenience of explanation, in this document, the inter-prediction type (L0 prediction, L1 prediction, or BI prediction) indicated by the inter_pred_idc syntax element may be referred to as a motion prediction direction. L0 prediction may be represented as pred_L0, L1 prediction as pred_L1, and bi-prediction as pred_BI. For example, the following prediction types can be expressed depending on the value of the inter_pred_idc syntax element:
[0134] [Table 1]
[0135] As described above, one picture may include one or more slices. A slice may have one of slice types including an I (intra) slice, a P (predictive) slice, and a B (bi-predictive) slice. The slice type may be indicated based on slice type information. For blocks in an I slice, inter prediction may not be used for prediction, and only intra prediction may be used. Of course, even in this case, original sample values may be coded and signaled without prediction. For blocks in a P slice, intra prediction or inter prediction may be used, and if inter prediction is used, only uni prediction may be used. On the other hand, for blocks in a B slice, intra prediction or inter prediction may be used, and if inter prediction is used, up to bi prediction may be used.
[0136] L0 and L1 may include reference pictures encoded / decoded before the current picture. For example, L0 may include reference pictures before and / or after the current picture in POC order, and L1 may include reference pictures after and / or before the current picture in POC order. In this case, L0 may be assigned a reference picture index that is lower relative to a reference picture before the current picture in POC order, and L1 may be assigned a reference picture index that is lower relative to a reference picture after the current picture in POC order. In the case of a B slice, bi-prediction may be applied, and in this case, unidirectional bi-prediction may be applied, or bi-directional bi-prediction may be applied. Bi-directional bi-prediction may be referred to as true bi-prediction.
[0137] The coding apparatus derives motion information for the current block (S610). The motion information may be derived based on the inter prediction mode.
[0138] A coding apparatus may perform inter-prediction using motion information of a current block. An encoding apparatus may derive optimal motion information for a current block through a motion estimation procedure. For example, the encoding apparatus may search for a similar reference block with high correlation using an original block in an original picture for the current block in fractional pixel units within a predetermined search range in the reference picture, thereby deriving motion information. Block similarity may be derived based on a difference in sample values based on phase. For example, block similarity may be calculated based on the SAD between the current block (or a template of the current block) and a reference block (or a template of the reference block). In this case, motion information may be derived based on the reference block with the smallest SAD within the search range. The derived motion information may be signaled to a decoding apparatus in various ways based on the inter-prediction mode.
[0139] The coding apparatus performs inter prediction based on motion information for the current block (S620). The coding apparatus may derive predictive samples (and the like) for the current block based on the motion information. The current block including the predictive samples may be referred to as a predicted block.
[0140] When a merge mode is applied, the motion information of the current prediction block is not directly transmitted, but is derived using the motion information of neighboring prediction blocks. Therefore, the motion information of the current prediction block can be indicated by transmitting flag information indicating that the merge mode is used and a merge index indicating which neighboring prediction block is used. The merge mode can be called a regular merge mode.
[0141] To perform the merge mode, the encoder must search for merge candidate blocks to be used to derive motion information for the current prediction block. For example, up to five merge candidate blocks may be used, but the embodiment of this document is not limited thereto. The maximum number of merge candidate blocks may be transmitted in a slice header or a tile group header. After searching for the merge candidate blocks, the encoder may generate a merge candidate list and select the merge candidate block with the smallest cost as the final merge candidate block.
[0142] The merge candidate list may include, for example, five merge candidate blocks. For example, four spatial merge candidates and one temporal merge candidate may be used. Hereinafter, the spatial merge candidates or spatial MVP candidates described below may be referred to as SMVPs, and the temporal merge candidates or temporal MVP candidates described below may be referred to as TMVPs.
[0143] FIG. 7 shows an exemplary hierarchical structure for coded images / videos.
[0144] Referring to Figure 7, the coded image / video is divided into a VCL (video coding layer) that handles the image / video decoding process and itself, a lower system that transmits and stores the coded information, and a NAL (network abstraction layer) that exists between the VCL and the lower system and is responsible for network adaptation functions.
[0145] The VCL can generate VCL data containing compressed image data (slice data), or parameter sets containing information such as a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), and a Video Parameter Set (VPS), or SEI (Supplemental Enhancement Information) messages that are additionally required for the image decoding process.
[0146] In NAL, NAL units can be generated by adding header information (NAL unit header) to RBSP (Raw Byte Sequence Payload) generated by VCL. In this case, RBSP refers to slice data, parameter sets, SEI messages, etc. generated by VCL. The NAL unit header can contain NAL unit type information identified by the RBSP data included in the corresponding NAL unit.
[0147] As shown in the drawing, NAL units can be classified into VCL NAL units and non-VCL NAL units according to the RBSP generated by the VCL. A VCL NAL unit can refer to a NAL unit containing information about an image (slice data), and a non-VCL NAL unit can refer to a NAL unit containing information necessary for decoding an image (parameter set or SEI message).
[0148] The VCL NAL unit and non-VCL NAL unit can be transmitted over a network with header information according to the data standard of the lower system. For example, the NAL unit can be transformed into a data format of a predetermined standard such as H.266 / VVC file format, RTP (Real-time Transport Protocol), TS (Transport Stream), etc. and transmitted over various networks.
[0149] As mentioned above, the NAL unit type of an NAL unit can be identified by the RBSP data structure included in the corresponding NAL unit, and information about such NAL unit type can be stored and signaled in the NAL unit header.
[0150] For example, NAL units can be broadly classified into VCL NAL unit types and non-VCL NAL unit types depending on whether the NAL unit contains information about an image (slice data). The VCL NAL unit types can be classified according to the nature and type of pictures included in the VCL NAL unit, and the non-VCL NAL unit types can be classified according to the type of parameter set.
[0151] The following is an example of a NAL unit type identified by the type of parameter set that the non-VCL NAL unit type includes:
[0152] -APS (Adaptation Parameter Set) NAL unit: Type for NAL units including APS
[0153] -DPS (Decoding Parameter Set) NAL unit: Type for NAL units containing DPS
[0154] -VPS (Video Parameter Set) NAL unit: Type for NAL unit containing VPS
[0155] -SPS (Sequence Parameter Set) NAL unit: Type for NAL unit including SPS
[0156] -PPS (Picture Parameter Set) NAL unit: Type for NAL unit including PPS
[0157] -PH (Picture header) NAL unit: Type for NAL unit including PH
[0158] The above-mentioned NAL unit type has syntax information for the NAL unit type, and the syntax information can be stored in a NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified by the nal_unit_type value.
[0159] Meanwhile, as described above, one picture may include multiple slices, and one slice may include a slice header and slice data. In this case, one picture header may be added to multiple slices (slice header and slice data set) in one picture. The picture header (picture header syntax) may include information / parameters commonly applicable to the picture. In this document, slices may be mixed with or replaced by tile groups. Also, in this document, slice headers may be mixed with or replaced by type group headers.
[0160] The slice header (slice header syntax) can include information / parameters commonly applicable to the slices. The APS (APS syntax) or PPS (PPS syntax) can include information / parameters commonly applicable to one or more slices or pictures. The SPS (SPS syntax) can include information / parameters commonly applicable to one or more sequences. The VPS (VPS syntax) can include information / parameters commonly applicable to multiple layers. The DPS (DPS syntax) can include information / parameters commonly applicable to video in general. The DPS can include information / parameters related to the concatenation of coded video sequences (CVSs). In this document, a high level syntax (HLS) can include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, and slice header syntax.
[0161] In this document, image / video information encoded from an encoding device to a decoding device and signaled in the form of a bitstream may include not only intra-picture partitioning-related information, intra / inter prediction information, residual information, in-loop filtering information, etc., but also information included in the slice header, information included in the picture header, information included in the APS, information included in the PPS, information included in the SPS, information included in the VPS, and / or information included in the DPS. In addition, the image / video information may further include information of a NAL unit header.
[0162] Meanwhile, to compensate for differences between an original image and a restored image due to errors that occur during the compression encoding process, such as quantization, an in-loop filtering procedure may be performed on the restored samples or pictures, as described above. As described above, the in-loop filtering may be performed in a filter unit of an encoding device and a filter unit of a decoding device, and a deblocking filter, SAO, and / or adaptive loop filter (ALF) may be applied. For example, the ALF procedure may be performed after the deblocking filtering procedure and / or SAO procedure is completed. However, in this case, the deblocking filtering procedure and / or SAO procedure may be omitted.
[0163] Meanwhile, to improve coding efficiency, luma mapping with chroma scaling (LMCS) can be applied as described above. LMCS can be called a loop reshaper (reshaping). To improve coding efficiency, LMCS control and / or LMCS-related information signaling can be performed hierarchically.
[0164] 8 illustrates an exemplary hierarchical structure of a CVS according to one embodiment of this document. A coded video sequence (CVS) may include a sequence parameter set (SPS), a picture parameter set (PPS), a tile group header, tile data, and / or a CTU(s). Here, the tile group header and the tile data may also be referred to as a slice header and slice data, respectively.
[0165] The SPS may primitively include flags for enabling tools used in the CVS. The SPS may also be referenced by a PPS, which contains information on parameters that vary from picture to picture. Each coded picture may include one or more coded rectangular tiles. The tiles may be grouped by raster scanning to form a tile group. Each tile group is encapsulated with header information called a tile group header. Each tile consists of a CTU containing coded data. Here, the data may include original sample values, predicted sample values, and their luma and chroma components (luma predicted sample values and chroma predicted sample values).
[0166] Figure 9 illustrates an exemplary LMCS structure according to one embodiment of this document. The LMCS structure 900 of Figure 9 may include an in-loop mapping portion 910 for luma components based on an adaptive piecewise linear (PWL) model, and a luma-dependent chroma residual scaling portion 920 for chroma components. The inverse quantization and inverse transform 911, reconstruction 912, and intra-prediction 913 blocks of the in-loop mapping portion 910 represent processes applied in the mapped (reshaped) domain. The loop filter 915, motion compensation or inter-prediction 917 blocks of the in-loop mapping portion 910, and the reconstruction 922, intra-prediction 923, motion compensation or inter-prediction 924, and loop filter 925 blocks of the chroma residual scaling portion 920 represent processes applied in the original (non-mapped, non-reshaped) domain.
[0167] As described in FIG. 9, when LMCS is enabled, at least one of an inverse reshaping (mapping) process 914, a forward reshaping (mapping) process 918, and a chroma scaling process 921 may be applied. For example, the inverse reshaping process may be applied to the (reconstructed) luma samples (or luma samples or luma sample arrays) of the reconstructed picture. The inverse reshaping process may be performed based on the piecewise function (inverse) index of the luma samples. The piecewise function (inverse) index may identify the piece (or portion) to which the luma sample belongs. The output of the inverse reshaping process is the modified (reconstructed) luma samples (or modified luma samples or modified luma sample arrays). The LMCS may be enabled or disabled at the tile group (or slice), picture, or higher level.
[0168] A forward reshaping process and / or a chroma scaling process may be applied to generate a reconstructed picture. A picture may include luma samples and chroma samples. A reconstructed picture with luma samples may be referred to as a reconstructed luma picture, and a reconstructed picture with chroma samples may be referred to as a reconstructed chroma picture. A combination of a reconstructed luma picture and a reconstructed chroma picture may be referred to as a reconstructed picture. The reconstructed luma picture may be generated based on a forward reshaping process. For example, if inter prediction is applied to the current block, forward reshaping is applied to luma prediction samples derived based on the (reconstructed) luma samples of a reference picture. Since the (reconstructed) luma samples of the reference picture are generated based on an inverse reshaping process, forward reshaping may be applied to the luma prediction samples to derive reshaped (mapped) luma prediction samples. The forward reshaping process may be performed based on piecewise function indexes of the luma prediction samples. The piecewise function index can be derived based on the values of the luma prediction samples or the values of the luma samples of the reference picture used for inter prediction. If intra prediction (or IBC (intra block copy)) is applied to the current block, forward mapping is not necessary because the inverse reshaping process has not yet been applied to the reconstructed samples of the current picture. The (reconstructed) luma samples in the reconstructed luma picture are generated based on the reshaped luma prediction samples and the corresponding luma residual samples.
[0169] The reconstructed chroma picture can be generated based on a chroma scaling process. For example, the (reconstructed) chroma samples in the reconstructed chroma picture are the chroma predicted samples and the chroma residual samples (c res ) can be derived based on the chroma residual sample (cres ) is the (scaled) chroma residual sample (c resScale ) and a chroma residual scaling factor (cScaleInv, which may be referred to as varScale). The chroma residual scaling factor may be calculated based on the reshaped luma predicted sample values of the current block. For example, the scaling factor may be calculated based on the reshaped luma predicted sample values (Y′ pred Average luma value (ave(Y′) pred )) can be calculated based on the inverse transform / inverse quantization. For reference, in FIG. 9, the (scaled) chroma residual samples derived based on the inverse transform / inverse quantization are calculated based on the resScale The chroma residual sample obtained by performing an (inverse) scaling procedure on the (scaled) chroma residual sample is called c res It can be called.
[0170] Figure 10 shows an LMCS structure according to another embodiment of the present document. Figure 10 will be described with reference to Figure 9. Here, the differences between the LMCS structure of Figure 10 and the LMCS structure 900 of Figure 9 will be mainly described. The in-loop mapping portion and the luma-dependent chroma residual scaling portion of Figure 10 can operate in the same manner as / similar to the in-loop mapping portion 910 and the luma-dependent chroma residual scaling portion 920 of Figure 9.
[0171] Referring to FIG. 10, a chroma residual scaling factor can be derived based on luma reconstruction samples. In this case, the average luma value (avgY r ), and the average luma value (avgY r) to derive a chroma residual scaling factor. Here, the neighboring luma reconstructed samples are neighboring luma reconstructed samples of a current block or neighboring luma reconstructed samples of a VPDU (virtual pipeline data unit) including the current block. For example, when intra prediction is applied to the current block, reconstructed samples may be derived based on prediction samples derived based on the intra prediction. Furthermore, for example, when inter prediction is applied to the current block, reconstructed samples may be generated based on reshaped (or forward-mapped) luma prediction samples by applying forward mapping to the prediction samples derived based on the inter prediction.
[0172] Video / image information signaled via a bitstream may include LMCS parameters (information regarding LMCS). The LMCS parameters may be configured using HLS (high level syntax, including slice header syntax), etc. A detailed description of LMCS parameters and configuration will be provided below. As described above, the syntax tables described in this document (and in the following embodiments) may be configured / encoded at the encoder end and signaled to the decoder end via the bitstream. The decoder may parse / decode the information regarding LMCS (in the form of syntax components) in the syntax tables. One or more of the embodiments described below may be combined. The encoder may encode a current picture based on information regarding the LMCS, and the decoder may decode the current picture based on information regarding the LMCS.
[0173] In-loop mapping of the luma component can adjust the dynamic range of the input signal by redistributing codewords across the dynamic range to improve compression efficiency. For luma mapping, a forward mapping (reshaping) function (FwdMap) and an inverse mapping (reshaping) function (InvMap) corresponding to the forward mapping function (FwdMap) can be used. The forward mapping function (FwdMap) can be signaled using a piecewise linear model. For example, the piecewise linear model can have 16 pieces or bins. The pieces can have the same length. In one example, the inverse mapping function (InvMap) is not separately signaled, but can instead be derived from the forward mapping function (FwdMap). That is, the inverse mapping is a function of the forward mapping. For example, the inverse mapping function is a function that is symmetric to the forward mapping function with respect to y=x.
[0174] In-loop (luma) reshaping can be used to map input luma values (samples) with modified values in a reshaped domain. The reshaped values can be encoded and then remapped to the original (unmapped, unreshaped) domain after decompression. Chroma residual scaling can be applied to compensate for differences between the luma and chroma signals. In-loop reshaping can be performed by specifying a high-level syntax for the reshaper model. The reshaper model syntax can signal a partially linear model (PWL model). A forward lookup table (FwdLUT) and / or an inverse lookup table (InvLUT) can be derived based on the partially linear model. For example, if a forward lookup table (FwdLUT) is derived, an inverse lookup table (InvLUT) can be derived based on the forward lookup table (FwdLUT). The forward lookup table (FwdLUT) is used to map the input luma values Y i The changed value Y r The inverse lookup table (InvLUT) maps the restored value Y based on the modified value. r The restored value Y′ i The restored value Y′ can be mapped as i is the input luma value Y i can be derived based on
[0175] In one example, the SPS may include the syntax in Table 2 below. The syntax in Table 2 may include sps_reshaper_enabled_flag as a tool enabling flag. Here, sps_reshaper_enabled_flag may be used to specify whether a reshaper is used in a coded video sequence (CVS). That is, sps_reshaper_enabled_flag is a flag that enables reshaping in the SPS. In one example, the syntax in Table 2 is part of the SPS.
[0176] [Table 2]
[0177] In one example, the semantics that sps_seq_parameter_set_id and sps_reshaper_enabled_flag can indicate are as shown in Table 3 below.
[0178] [Table 3]
[0179] In one example, a tile group header or slice header may include the syntax of Table 4 or Table 5 below.
[0180] [Table 4]
[0181] [Table 5]
[0182] The semantics of the syntax elements included in the syntax of Table 4 or Table 5 may include, for example, the items disclosed in the following table.
[0183] [Table 6]
[0184] [Table 7]
[0185] As an example, when sps_reshaper_enabled_flag is parsed, the tile group header can parse additional data (e.g., information contained in Table 6 or Table 7) used to configure the lookup tables (FwdLUT and / or InvLUT). To this end, the state of the SPS reshaper flag can be checked in the slice header or tile group header. If sps_reshaper_enabled_flag is true (or 1), an additional flag, tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag), can be parsed. The purpose of tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag) is to indicate the presence of a reshaper model. For example, if tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag) is true (or 1), it can be indicated that a reshaper exists for the current tile group (or current slice). If tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag) is false (or 0), it can be indicated that a reshaper does not exist for the current tile group (or current slice).
[0186] If a reshaper exists and is enabled for the current tile group (or current slice), the reshaper model (e.g., tile_group_reshaper_model() or slice_reshaper_model()) can be processed, and an additional flag, tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag) can also be parsed. tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag) can indicate whether a reshaper model is used for the current tile group (or slice). For example, if tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag) is 0 (or false), it can be indicated that a reshaper model is not used for the current tile group (or current slice). If tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag) is 1 (or true), the reshaper model can be designated to be used for the current tile group (or slice).
[0187] As an example, tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag) is true (or 1) and tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag) is false (or 0). This means that a reshaper model exists but is not currently being used in the tile group (or slice). In this case, the reshaper model can be used in the next tile group (or slice). Another example is when tile_group_reshaper_enable_flag is true (or 1) and tile_group_reshaper_model_present_flag is false (or 0).
[0188] When the reshaper model (e.g., tile_group_reshaper_model() or slice_reshaper_model()) and tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag) are parsed, it can be determined (evaluated) whether the conditions necessary for chroma scaling exist. The conditions can include condition 1 (the current tile group / slice is not intra-coded) and / or condition 2 (the current tile group / slice is not split into two separate coding quad-tree structures for luma and chroma, i.e., the current tile group / slice is not a dual-tree structure). If condition 1 and / or condition 2 are true and / or tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag) is true (or 1), tile_group_reshaper_chroma_residual_scale_flag (or slice_reshaper_chroma_residual_scale_flag) can be parsed. If tile_group_reshaper_chroma_residual_scale_flag (or slice_reshaper_chroma_residual_scale_flag) is enabled (1 or true), it can indicate that chroma residual scaling is enabled for the current tile group (or slice). If tile_group_reshaper_chroma_residual_scale_flag (or slice_reshaper_chroma_residual_scale_flag) is disabled (0 or false), it can be indicated that chroma residual scaling is disabled for the current tile group (or slice).
[0189] The purpose of the reshaping described above is to parse the data necessary to construct a lookup table (FwdLUT and / or InvLUT). In one example, the lookup table constructed based on the parsed data can divide the distribution of allowable luma value ranges into a number of bins (e.g., 16 bins). Thus, luma values within a given bin can be mapped to modified luma values.
[0190] Figure 11 shows a graph illustrating an exemplary forward mapping. Only five bins are shown in Figure 8 for illustrative purposes.
[0191] Referring to Figure 11, the x-axis represents input luma values and the y-axis represents modified output luma values. The x-axis is divided into five bins or pieces, each having a length L. That is, the five bins mapped to modified luma values have the same length as each other. A forward lookup table (FwdLUT) can be constructed using data available in the tile group header (e.g., reshaper data), facilitating the mapping.
[0192] In one embodiment, output pivot points associated with the bin index may be calculated. The output pivot points may mark the minimum and maximum boundaries of the output range of luma codeword reshaping. The process of calculating the output pivot points may be performed based on a piecewise cumulative distribution function of the number of codewords. The output pivot range may be divided based on the maximum number of bins used and the size of the lookup table (FwdLUT or InvLUT). For example, the output pivot range may be divided based on the product of the maximum number of bins and the size of the lookup table. For example, if the product of the maximum number of bins and the size of the lookup table is 1024, the output pivot range may be divided into 1024 entries. The division of the output pivot range may be performed (applied or achieved) based on (using) a scaling factor. In one example, the scaling factor may be derived based on Equation 1 below.
[0193]
number
[0194] In Equation 1, SF represents a scaling factor, y1 and y2 represent output pivot points corresponding to each bin, and FP_PREC and c are predetermined constants. The scaling factor determined based on Equation 1 can be called a scaling factor for forward reshaping.
[0195] In another embodiment, in connection with inverse reshaping (inverse mapping), for a defined range of bins (e.g., from reshaper_model_min_bin_idx to reshape_model_max_bin_idx), the input reshaped pivot points and the mapped inverse output pivot points (given as bin index * number of initial codewords) corresponding to the mapped pivot points of the forward lookup table (FwdLUT) are patched. In another example, the scaling factor (SF) can be derived based on the following Equation 2:
[0196]
number
[0197] In Equation 2, SF represents a scaling factor, x1 and x2 represent input pivot points, and y1 and y2 represent output pivot points corresponding to each piece (bin). Here, the input pivot points are pivot points mapped based on a forward lookup table (FwdLUT), and the output pivot points are pivot points inversely mapped based on an inverse lookup table (InvLUT). FP_PREC is a predetermined constant. FP_PREC in Equation 2 may be the same as or different from FP_PREC in Equation 1. The scaling factor determined based on Equation 2 may be referred to as a scaling factor for inverse reshaping. During inverse reshaping, division of the input pivot points may be performed based on the scaling factor of Equation 2. Based on the split input pivot points, pivot values corresponding to the minimum and maximum bin values are specified for bin indices ranging from 0 to the minimum bin index (reshaper_model_min_bin_idx) and / or from the minimum bin index (reshaper_model_min_bin_idx) to the maximum bin index (reshaper_model_max_bin_idx).
[0198] Table 8 below shows the syntax of a reshaper model according to one embodiment. The reshaper model may be referred to as an LMCS model. Here, the reshaper model is exemplarily described as a tile group reshaper, but this description is not necessarily limited to this embodiment. For example, the reshaper model may be included in an APS, or the tile group reshaper model may be referred to as a slice reshaper model or LMCS data. Also, the prefix reshaper_model or Rsp may be used interchangeably with lmcs. For example, in the following table and description, reshaper_model_min_bin_idx, reshaper_model_delta_max_bin_idx, reshaper_model_max_bin_idx, RspCW, and RsepDeltaCW can be used interchangeably with lmcs_min_bin_idx, lmcs_delta_max_bin_idx, lmcs_max_bin_idx, lmcsCW, and lmcsDeltaCW, respectively.
[0199] [Table 8]
[0200] The semantics of the syntax elements included in the syntax of Table 8 may include, for example, those disclosed in the following table.
[0201] [Table 9-1]
[0202] [Table 9-2]
[0203] The inverse mapping procedure for luma samples according to this document can be written in a standard document format as in the table below.
[0204] [Table 10]
[0205] The identification of the piecewise function index procedure for luma samples according to this document can be written in a standard document format such as the following table: In Table 11, idxYInv can be referred to as an inverse mapping index, and the inverse mapping index can be derived based on the restored luma sample (lumaSample).
[0206] [Table 11]
[0207] Luma mapping can be performed based on the above-described embodiments and examples, and the above-described syntax and components included therein are merely exemplary representations, and the embodiments are not limited by the detailed tables and formulas. A method for performing chroma residual scaling (scaling of chroma components of residual samples) based on luma mapping will be described below.
[0208] Luma-dependent chroma residual scaling compensates for differences between luma samples and corresponding chroma samples. For example, whether chroma residual scaling is enabled can be signaled at the tile group level or slice group level. In one example, if luma mapping is enabled and dual tree partitioning is not currently applied to the tile group, an additional flag can be signaled to indicate whether luma-dependent chroma residual scaling is enabled. In another example, if luma mapping is not used or dual tree partitioning is not currently used for the tile group, luma-dependent chroma residual scaling can be disabled. In another example, chroma residual scaling can always be disabled for chroma blocks having a size less than or equal to 4.
[0209] Chroma residual scaling may be performed based on the average value of a corresponding luma prediction block (a luma component of a prediction block to which an intra prediction mode and / or an inter prediction mode is applied). The scaling operation at the encoder end and / or decoder end may be realized using fixed-point constant arithmetic based on the following Equation 3:
[0210]
number
[0211] In the above-mentioned Equation 3, c' represents a scaled chroma residual sample (a scaled chroma component of a residual sample), c represents a chroma residual sample (a chroma component of a residual sample), s represents a chroma residual scaling factor, and CSCALE_FP_PREC can represent a predetermined constant, for example, CSCALE_FP_PREC is 11.
[0212] 12 is a flow diagram illustrating a method for deriving a chroma-residual scaling index according to one embodiment of the present document. The method described in conjunction with FIG. 12 may be performed based on the tables, formulas, variables, arrays, and functions contained in FIG. 9 and the associated description.
[0213] In step S1210, it may be determined whether the prediction mode is an intra prediction mode or an inter prediction mode based on the prediction mode information. If the prediction mode is an intra prediction mode, the current block or its predicted samples are considered to be in an already reshaped (mapped) region. If the prediction mode is an inter prediction mode, the current block or its predicted samples are considered to be in an original (unmapped, unreshaped) region.
[0214] In step S1220, if the prediction mode is an intra prediction mode, the average of the current block (or the luma prediction samples of the current block) may be calculated (derived). That is, the average of the current block in the already reshaped region is directly calculated. The average is also called a mean value.
[0215] In step S1221, if the prediction mode is an inter prediction mode, forward reshaping (forward mapping) may be performed (applied) to the luma prediction samples of the current block. Through forward reshaping, the luma prediction samples based on the inter prediction mode may be mapped from their original regions to reshaped regions. In one example, the forward reshaping of the luma prediction samples may be performed based on the reshaper model described above in conjunction with Table 4.
[0216] In step S1222, an average of the forward-reshaped (forward-mapped) luma prediction samples may be calculated (derived), i.e., an averaging process may be performed on the forward-reshaped result.
[0217] In step S1230, a chroma residual scaling index may be calculated. If the prediction mode is an intra prediction mode, the chroma residual scaling index may be calculated based on an average of the luma prediction samples. If the prediction mode is an inter prediction mode, the chroma residual scaling index may be calculated based on an average of forward-reshaped luma prediction samples.
[0218] In one embodiment, the chroma residual scaling index can be calculated based on a for loop syntax: The following table shows an exemplary for loop syntax for deriving (calculating) the chroma residual scaling index:
[0219] [Table 12]
[0220] In Table 12, idxS indicates the chroma-residual scaling index, idxS indicates an index that identifies whether a chroma-residual scaling index that satisfies the condition of the if statement has been found, S indicates a predetermined constant, and MaxBinIdx indicates the maximum allowable bin index. ReshapPivot[idxS+1] can be derived based on Table 8 and / or Table 9.
[0221] In one embodiment, a chroma resistive scaling factor can be derived based on the chroma resistive scaling index. Equation 4 below is an example for deriving the chroma resistive scaling factor.
[0222]
number
[0223] In Equation 4, s represents the chroma residual scaling factor, and ChromaScaleCoef is a variable (or array) derived based on Table 8 and / or Table 9 above.
[0224] As described above, an average luma value of the reference samples may be obtained, a chroma residual scaling factor may be derived based on the average luma value, scaling may be performed on the chroma component residual samples based on the chroma residual scaling factor, and chroma component restored samples may be generated based on the scaled chroma component residual samples.
[0225] In one embodiment of this document, a signaling structure for efficiently applying the above-mentioned LMCS is proposed. According to this embodiment, for example, LMCS data may be included in HLS (e.g., APS), and an LMCS model (reshaper model) may be adaptively derived by signaling a referenced APS ID through header information (e.g., a picture header, a slice header) that is a lower level of the APS. The LMCS model may be derived based on LMCS parameters. Also, for example, multiple APS IDs may be signaled through the header information, thereby allowing different LMCS models to be applied to blocks within the same picture / slice.
[0226] In one embodiment of this document, a method for efficiently performing the operations required for LMCS is proposed. According to the semantics detailed in Table 9, a division operation by the piece length lmcsCW[i] is required to derive InvScaleCoeff[i]. However, if the piece length is not a power of 2, the division operation cannot be performed by bit shifting.
[0227] For example, the calculation of InvScaleCoeff requires a maximum of 16 division operations per slice. According to Table 8, in the case of 10-bit coding, the range of lmcsCW[i] is from 8 to 511, so in order to implement the division operation by lmcsCW[i] using an LUT, the size of the LUT must be 504. In addition, in the case of 12-bit coding, the range of lmcsCW[i] is from 32 to 2047, so in order to implement the division operation by lmcsCW[i] using an LUT, the size of the LUT must be 2016. In other words, division operations can be quite expensive in hardware implementation, and therefore, it is better to omit division operations if possible.
[0228] In one aspect of this embodiment, lmcsCW[i] can be limited to a multiple of a fixed number (or a pre-defined or pre-determined number). This can reduce the size or capacity of a lookup table (LUT) for a division operation. For example, if lmcsCW[i] is a multiple of 2, the size of the LUT for substituting a division operation can be reduced by half.
[0229] In another aspect of this embodiment, high internal bit depth coding is proposed. High internal bit depth coding is an upper condition for the range restriction of lmcsCW[i]. For example, if the coding bit depth is higher than 10, lmcsCW[i] may be restricted to a multiple of 1<<(BitDepthY-10), where BitDepthY is the luma bit depth. As a result, the possible values of lmcsCW[i] do not change depending on the coding bit depth, and therefore the size of the LUT for calculating InvScaleCoeff does not increase even if the coding bit depth is high. In one example, for a 12-bit internal coding bit depth, the value of lmcsCW[i] may be restricted to a multiple of 4, and therefore the size of the LUT for replacing the division operation is the same as the size of the LUT used for 10-bit coding. This aspect can be implemented independently or in combination with the previously described aspect.
[0230] In another aspect of this embodiment, lmcsCW[i] can be restricted to a narrower range. For example, lmcsCW[i] can be restricted to the range from (OrgCW>>1) to (OrgCW<<1)-1. In the case of 10-bit coding, the range of lmcsCW[i] is [32, 127], and InvScaleCoeff can be calculated using only a LUT with a size of 96.
[0231] In another aspect of this embodiment, lmcsCW[i] can be approximated to a value close to a power of 2 and used in the design of the reshaper, so that the division operation in the inverse mapping procedure can be performed (replaced) by bit shifting.
[0232] In one embodiment of this document, a restriction on the LMCS codeword range is proposed. According to Table 9 above, the value of the LMCS codeword is in the range from (OrgCW>>3) to (OrgCW<<3)-1. If a large range of LMCS piece lengths occurs, there is a risk of visual degradation if a large difference occurs between RspCW[i] and OrgCW.
[0233] According to one embodiment of this document, it is proposed to restrict the codewords of the LMCS PWL mapping to a narrow range, for example, lmcsCW[i] ranges from (OrgCW>>1) to (OrgCW<<1)-1.
[0234] In one embodiment of this document, the use of a single chroma residual scaling factor is proposed for chroma residual scaling in LMCS. Existing methods for deriving a chroma residual scaling factor derive the slope of each piece in inverse luma mapping as the corresponding scaling factor using the average value of the corresponding luma block. Furthermore, a latency problem occurs due to the procedure for identifying a piecewise index that requires the availability of the corresponding luma block. This is undesirable in hardware implementation. In one embodiment of this document, scaling in a chroma block becomes independent of the luma block value, and piecewise index identification is not required. Therefore, the chroma residual scaling procedure in LMCS can be performed without latency problems.
[0235] In one embodiment according to this document, a single chroma scaling factor can be derived at both the encoder and the decoder based on the luma LMCS information. When the LMCS luma model is received, the chroma residual scaling factor can be updated. For example, if the LMCS model is updated, the single chroma residual scaling factor can be updated.
[0236] The table below shows an example for obtaining a single chroma scaling factor according to this embodiment.
[0237] [Table 13]
[0238] Referring to Table 13, a single chroma scaling factor (eg, ChromaScaleCoeff or ChromaScaleCoeffSingle) can be obtained by averaging the inverse lumah mapping slopes of all pieces in the range between lmcs_min_bin_idx and lmcs_max_bin_idx.
[0239] 13 illustrates a linear fitting of pivot points according to an embodiment of the present document. Pivot points P1, Ps, and P2 are shown in FIG. 13. The following embodiments or examples thereof will be described with reference to FIG. 13.
[0240] In one example of this embodiment, a single chroma scaling factor can be obtained based on a linear approximation of the luma PWL mapping between pivot points lmcs_min_bin_idx and lmcs_max_bin_idx+1. That is, the inverse slope of the linear mapping can be used as the chroma residual scaling factor. For example, linear line 1 in FIG. 13 is a line connecting pivot points P1 and P2. Referring to FIG. 13, at P1, the input value is x1 and the mapped value is 0, and at P2, the input value is x2 and the mapped value is y2. The inverse slope (inverse scale) of linear line 1 is (x2-x1) / y2, and the single chroma scaling factor ChromaScaleCoeffSingle can be calculated based on the input values and mapped values of pivot points P1 and P2 and the following formula:
[0241]
number
[0242] In Equation 5, CSCALE_FP_PREC represents a shift factor, for example, CSCALE_FP_PREC is a predetermined constant. In one example, CSCALE_FP_PREC is 11.
[0243] 13, at pivot point Ps, the input value is min_bin_idx+1 and the mapped value is ys, so that the inverse slope (inverse scale) of linear line 1 can be calculated as (xs-x1) / ys, and the single chroma scaling factor ChromaScaleCoeffSingle can be calculated based on pivot point P1, the input value and mapped value of Ps, and the following formula:
[0244]
number
[0245] In Equation 6, CSCALE_FP_PREC indicates a shift factor (factor for bit shifting), for example, CSCALE_FP_PREC is a predetermined constant. In one example, CSCALE_FP_PREC is 11, and bit shifting for the inverse scale can be performed based on CSCALE_FP_PREC.
[0246] In another example according to this embodiment, a single chroma residual scaling factor can be derived based on a linear approximation line. One example for deriving a linear approximation line can include a linear concatenation of pivot points (e.g., lmcs_min_bin_idx, lmcs_max_bin_idx+1). For example, the linear trend result can be represented by a codeword of PWL mapping. The mapped value y2 at P2 is the sum of the codewords of all bins (pieces), and the difference (x2-x1) between the input value at P2 and the input value at P1 is OrgCW*(lmcs_max_bin_idx-lmcs_min_bin_idx+1) (OrgCW is shown in Table 9 above). The following table shows an example of obtaining a single chroma scaling factor according to the above-described embodiment.
[0247] [Table 14]
[0248] Referring to Table 14, a single chroma scaling factor (e.g., ChromaScaleCoeffSingle) can be obtained from two pivot points (i.e., lmcs_min_bin_idx, lmcs_max_bin_idx). For example, the inverse slope of linear mapping can be used as the chroma scaling factor.
[0249] In another example of this embodiment, a single chroma scaling factor can be obtained by linear fitting of the pivot points to minimize the error (or mean square error) between the linear fitting and the existing PWL mapping. This example is more accurate than simply concatenating two pivot points, lmcs_min_bin_idx and lmcs_max_bin_idx. There can be various ways to find the optimal linear mapping, and one example is described below.
[0250] In one example, the parameters b1 and b0 of the linear fitting equation y=b1*x+b0 for minimizing the sum of the least squares error can be calculated based on the following Equation 7 and / or Table 8.
[0251]
number
[0252]
number
[0253] Here, x denotes the original luma value and y denotes the reshaped luma value. JPEG0007755039000024.jpg1391 indicates the average of x and y, respectively, and xi and yi indicate the value of the i-th pivot point.
[0254] Referring to FIG. 13, another approximation for identifying a linear mapping can be given as follows:
[0255] - Calculate lmcs_pivots_linear[i] for linear lines with input values that are multiples of OrgCW, obtaining linear line 1 by concatenating the pivot points of the PWL mapping at lmcs_min_bin_idx and lmcs_max_bin_idx+1.
[0256] - Add up the difference between the mapped values of the pivot points using Linear Line 1 and PWL mapping
[0257] -Get the average difference (avgDiff)
[0258] -Adjust the last pivot point of the linear line by the average difference (e.g., 2*avgDiff)
[0259] -Use the inverse slope of the adjusted linear line as the chroma residual scale
[0260] By the linear fitting described above, the chroma scaling factor (i.e., the inverse slope of forward mapping) can be derived (obtained) based on the following Equation 9 or 10.
[0261]
number
[0262]
number
[0263] In the above formula, lmcs_pivots_linear[i] is the mapped value of the linear mapping. Through the linear mapping, all pieces of the PWL mapping between the minimum and maximum bin indexes can have the same LMCS codeword (lmcsCW). That is, lmcs_pivots_linear[lmcs_min_bin_idx+1] is the same as lmcsCW[lmcs_min_bin_idx].
[0264] In addition, in Equation 9 and Equation 10, CSCALE_FP_PREC indicates a shift factor (factor for bit shifting), and for example, CSCALE_FP_PREC is a predetermined constant. In one example, CSCALE_FP_PREC is 11.
[0265] There is no need to calculate the average of the corresponding luma block through the single chroma residual scaling factor (ChromaScaleCoeffSingle), and there is no need to search for an index in the PWL linear mapping, which improves the coding efficiency using chroma residual scaling.
[0266] In another embodiment of this document, the encoder can determine parameters for a single chroma scaling factor and signal said parameters to the decoder. Through signaling, the encoder can utilize other information available to the encoder to derive the chroma residual scaling factor. This embodiment aims to eliminate the chroma residual scaling latency problem.
[0267] For example, a procedure for identifying a linear mapping used to determine the chroma-residual scaling factor can be given as follows:
[0268] - Calculate lmcs_pivots_linear[i] for linear lines with input values that are multiples of OrgCW, obtaining linear line 1 by concatenating the pivot points of the PWL mapping at lmcs_min_bin_idx and lmcs_max_bin_idx+1.
[0269] - Uses the pivot points of Linear Line 1 and Luma PWL mapping to obtain the weighted sum of the differences between the mapped values of the pivot points
[0270] -Get the weighted average difference (avgDiff)
[0271] -Adjust the last pivot point of Linear Line 1 by the weighted average difference (e.g., 2*avgDiff)
[0272] -Use the inverse slope of the adjusted linear line as the chroma residual scale
[0273] The following table shows an example syntax for signaling y values for chroma scaling factor derivation.
[0274] [Table 15]
[0275] In Table 15, the syntax element lmcs_chroma_scale can specify the single chroma (residual) scaling factor used for LMCS chroma residual scaling (ChromaScaleCoeffSingle=lmcs_chroma_scale). That is, information about the chroma residual scaling factor can be directly signaled, and the signaled information can be derived as a chroma residual scaling factor. In other words, the value of the signaled information about the chroma residual scaling factor can be (directly) derived as a single chroma residual scaling factor. Here, the syntax element lmcs_chroma_scale can be signaled together with other LMCS data (e.g., absolute value of a codeword, syntax elements related to a code, etc.).
[0276] Alternatively, the encoder can signal only the parameters necessary to derive the chroma-resistance scaling factor to the decoder. To derive the chroma-resistance scaling factor at the decoder, the input value x and the mapped value y are required. The x value indicates the bin length and is already known to the decoder, so it does not need to be signaled. Consequently, only the y value needs to be signaled to derive the chroma-resistance scaling factor. Here, the y value is the mapped value of an arbitrary pivot point in the linear mapping (e.g., the mapped value of P2 or Ps in FIG. 13).
[0277] The table below shows an example of signaling mapped values for chroma-residual scaling factor derivation.
[0278] [Table 16]
[0279] [Table 17]
[0280] One of the syntaxes in Tables 16 and 17 above can be used to signal the y value at any linear pivot point specified by the encoder and decoder, i.e., the encoder and decoder can derive the y value using the same syntax.
[0281] First, an embodiment according to Table 16 will be described. In Table 16, lmcs_cw_linear can indicate a value mapped to Ps or P2. That is, in the embodiment according to Table 16, a fixed number can be signaled via lmcs_cw_linear.
[0282] In one example according to this embodiment, when lmcs_cw_linear indicates a value mapped to one bin (i.e., lmcs_pivots_linear[lmcs_min_bin_idx+1] at Ps in FIG. 13), the chroma scaling factor can be derived based on the following formula:
[0283]
number
[0284] In another example according to this embodiment, when lmcs_cw_linear indicates lmcs_max_bin_idx+1 (ie, lmcs_pivots_linear[lmcs_max_bin_idx+1] at P2 in FIG. 13), the chroma scaling factor can be derived based on the following formula:
[0285]
number
[0286] In the above formula, CSCALE_FP_PREC indicates a shift factor (factor for bit shifting), for example, CSCALE_FP_PREC is a predetermined constant. In one example, CSCALE_FP_PREC is 11.
[0287] Next, an embodiment according to Table 17 will be described. In this embodiment, lmcs_cw_linear can also be signaled as a delta value associated with a fixed number (i.e., lmcs_delta_abs_cw_linear, lmcs_delta_sign_cw_linear_flag). In one example of this embodiment, if lmcs_cw_linear indicates a mapped value at lmcs_pivots_linear[lmcs_min_bin_idx+1] (i.e., Ps in FIG. 13), lmcs_cw_linear_delta and lmcs_cw_linear can be derived based on the following formulas:
[0288]
number
[0289]
number
[0290] In another example of this embodiment, when lmcs_cw_linear indicates the mapped value at lmcs_pivots_linear[lmcs_max_bin_idx+1] (i.e., P2 in FIG. 13), lmcs_cw_linear_delta and lmcs_cw_linear can be derived based on the following formula:
[0291]
number
[0292]
number
[0293] In the above formula, OrgCW is a value derived based on Table 9 above.
[0294] Figure 14 shows an example of linear reshaping (or linear reshaper, linear mapping) according to one embodiment of the present document. That is, in one embodiment of the present document, the use of a linear reshaper in LMCS is proposed. For example, the illustration in Figure 14 is associated with forward linear reshaping (mapping).
[0295] In existing examples, LMCS can use piecewise linear mapping with 16 fixed pieces. This can increase the complexity of reshaper design because abrupt transitions between pivot points can cause unavoidable degradation. Furthermore, for inverse luma mapping of the reshaper, piecewise function indices must be identified. The piecewise function index identification procedure is an iterative process involving many comparison steps. Furthermore, a luma piecewise index identification procedure is required for chroma residual scaling using the corresponding luma block average. This not only increases complexity but can also cause latency in chroma residual scaling, which depends on the restoration of the entire luma block. To solve this problem, the use of a linear reshaper in LMCS has been proposed, and a detailed description of the linear reshaper will be provided below.
[0296] Referring to FIG. 14, the linear reshaper may include two pivot points (i.e., P1 and P2). P1 and P2 may represent input and mapped values, e.g., P1 is (minInput, 0) and P2 is (maxInput, maxMapped). Here, minInput represents the minimum input value and maxInput represents the maximum input value. If the input value is less than or equal to minInput, it is mapped to 0, and if the input value is greater than maxInput, it is mapped to maxMapped. Input (luma) values between minInput and maxInput may be linearly mapped to other values. FIG. 14 shows an example of mapping. The pivot points P1 and P2 may be determined by the encoder, and linear fitting may be used to approximate a piecewise linear mapping for this purpose.
[0297] There are various ways to signal a linear reshaper. In one example of a method for signaling a linear reshaper, each luma range can be divided into an equal number of bins. That is, the luma mapping between the minimum and maximum bins can be achieved evenly. For example, all bins can have the same LMCS codeword (lmcsCW). For this purpose, the minimum and maximum bin indexes can be signaled. Also, only one set of reshape_model_bin_delta_abs_CW (or reshaper_model_delta_abs_CW, lmcs_delta_abs_CW) and reshaper_model_bin_delta_sign_CW_flag (or reshaper_model_delta_sign_CW_flag, lmcs_delta_sign_CW_flag) needs to be signaled.
[0298] The following table exemplarily illustrates the syntax and semantics for signaling a linear reshaper according to this example.
[0299] [Table 18]
[0300] [Table 19]
[0301] Referring to Tables 18 and 19, the syntax element log2_lmcs_num_bins_minus4 is information regarding the number of bins. Based on this information, the number of bins can be signaled, thereby enabling better control of the minimum and maximum pivot points. In other existing examples, the encoder and / or decoder can derive (explicitly) the (fixed) number of bins without signaling; for example, the number of bins is derived to be 16 or 32. However, according to the examples in Tables 18 and 19, log2_lmcs_num_bins_minus4 plus 4 can indicate the log2 (binary logarithm) of the number of bins. The number of bins derived based on this syntax element ranges from 4 to the value of the luma bit depth (BitDepthY).
[0302] In Table 19, ScaleCoeffSingle may be referred to as a single luma forward scaling factor, and InvScaleCoeffSingle may be referred to as a single luma inverse scaling factor. Forward mapping for predicted luma samples may be performed based on the single luma forward scaling factor, and inverse mapping for restored luma samples may be performed based on the single luma inverse scaling factor. ChromaSclaeCoeffSingle may be referred to as a single chroma residual scaling factor, as described above. ScaleCoeffSingle, InvScaleCoeffSingle, and ChromaSclaeCoeffSingle may be used for forward luma mapping, inverse luma mapping, and chroma residual scaling, respectively. ScaleCoeffSingle, InvScaleCoeffSingle, and ChromaSclaeCoeffSingle may be uniformly applied to all bins (16 bins PWL mappings) as a single factor.
[0303] Referring to Table 19, FP_PREC and CSCALE_FP_PREC are constants for bit shifting. FP_PREC and CSCALE_FP_PREC may or may not be the same as each other. For example, FP_PREC is greater than or equal to CSCALE_FP_PREC. In one example, FP_PREC and CSCALE_FP_PREC are both 11. In another example, FP_PREC is 15 and CSCALE_FP_PREC is 11.
[0304] In another example of a method for signaling a linear reshaper, an LMCS codeword (lmcsCWlinearALL) can be derived based on the following formula: In this example, the linear reshaping syntax element signaled by the syntax in Table 18 above can also be used. The following table shows an example of the semantics described by this example.
[0305] [Table 20]
[0306] Referring to Table 20, FP_PREC and CSCALE_FP_PREC are constants for bit shifting. FP_PREC and CSCALE_FP_PREC may or may not be the same as each other. For example, FP_PREC is greater than or equal to CSCALE_FP_PREC. In one example, FP_PREC and CSCALE_FP_PREC are both 11. In another example, FP_PREC is 15 and CSCALE_FP_PREC is 11.
[0307] In Table 20, lmcs_max_bin_idx can be used interchangeably with LmcsMaxBinIdx. lmcs_max_bin_idx and LmcsMaxBinIdx can be derived with reference to Table 19 based on lmcs_delta_max_bin_idx in the syntax of Table 18. In this way, Table 20 can be parsed with reference to Table 19.
[0308] In other embodiments according to this document, other examples of how to signal a linear reshaper can be proposed. The pivot points P1, P2 of the linear reshaper model can be explicitly signaled. The following table shows an example of the syntax and semantics of explicitly signaling a linear reshaper model according to this example:
[0309] [Table 21]
[0310] [Table 22]
[0311] Referring to Table 21 and Table 22, the input value of the first pivot point can be derived based on the syntax element lmcs_min_input, and the input value of the second pivot point can be derived based on the syntax element lmcs_max_input. The mapped value of the first pivot point is a predetermined value (a value known to both the encoder and the decoder), for example, 0. The mapped value of the second pivot point can be derived based on the syntax element lmcs_max_mapped. That is, the linear reshaper model can be explicitly (directly) signaled based on the information signaled based on the syntax of Table 21.
[0312] Alternatively, lmcs_max_input and lmcs_max_mapped can be signaled as delta values. The table below shows an example of the syntax and semantics of signaling a linear reshaper model as a delta value.
[0313] [Table 23]
[0314] [Table 24]
[0315] Referring to Table 24 above, the input value of the first pivot point can be derived based on the syntax element lmcs_min_input. For example, lmcs_min_input can have a mapped value of 0. lmcs_max_input_delta can indicate the difference between the input value of the second pivot point and the maximum luma value (i.e., (1<<bitdepthY)-1). lmcs_max_mapped_delta can indicate the difference between the mapped value of the second pivot point and the maximum luma value (i.e., (1<<bitdepthY)-1).
[0316] According to one embodiment of this document, based on the examples described above regarding the linear reshaper, forward mapping for luma prediction samples, inverse mapping for luma restoration samples, and chroma residual scaling can be performed. In one example, only one inverse scaling factor is required for inverse scaling of luma (restored) samples (pixels) in the linear reshaper-based inverse mapping. This is also the case for forward mapping and chroma residual scaling. That is, the step of determining ScaleCoeff[i], InvScaleCoeff[i], and ChromaScaleCoeff[i] for bin index i can be replaced by using only one factor. Here, one factor can mean that the (forward) slope or inverse slope of the linear mapping is represented in fixed-point notation. In one example, the inverse luma mapping scaling factor (the inverse scaling factor in the inverse mapping for luma restoration samples) can be derived based on at least one of the following equations.
[0317]
Equation
[0318]
Equation
[0319]
number
[0320] lmcsCWLinear in Equation 17 can be derived from the above-described Tables 18 and 19. lmcsCWLinearALL in Equation 18 and 19 can be derived from at least one of the above-described Tables 20 to 24. In Equation 17 or 18, OrgCW can be derived from Table 9 or 19.
[0321] The following table describes the formulas and syntax (conditional statements) that indicate the forward mapping procedure for luma samples (i.e., luma prediction samples) in picture reconstruction. In the following table and formulas, FP_PREC is a constant for bit shifting and is a predetermined value. For example, FP_PREC is 11 or 15.
[0322] [Table 25]
[0323] [Table 26]
[0324] Table 25 is for deriving the luma samples forward-mapped by the luma mapping procedure based on Table 8 and Table 9 described above. That is, Table 25 can be described together with Table 8 and Table 9. In Table 25, the forward-mapped luma (prediction) samples PredmAPSamples[i][j] as output can be derived from the luma (prediction) samples predSamples[i][j] as input. idxY in Table 25 can be called the (forward) mapping index, and the mapping index can be derived based on the predicted luma samples.
[0325] Table 26 is for deriving the luma samples forward-mapped from the luma mapping by applying the linear reshaper. For example, lmcs_min_input, lmcs_max_input, lmcs_max_mapped, and ScaleCoeffSingle in Table 26 can be derived by at least one of Tables 21 to 24. In Table 26, when lmcs_min_input < predSamples[i][j] < lmcs_max_input, the forward-mapped luma (prediction) samples PredmAPSamples[i][j] as output can be derived from the luma (prediction) samples predSamples[i][j] as input. Through the comparison between Table 25 and Table 26, the change from the existing LMCS by applying the linear reshaper can be seen from the perspective of forward mapping.
[0326] The following equations explain the inverse mapping procedure for luma samples (i.e., luma restoration samples). In the following equations, the lumaSample as input is the (pre-modification) luma restoration sample before inverse mapping. The invSample as output is the inversely mapped (modified) luma restoration sample. In other cases, the clipped invSample is also called the modified luma restoration sample.
[0327]
Equation
[0328]
number
[0329] Referring to Equation 20, the index idxInv can be derived based on Table 11. That is, Equation 20 is for deriving an inverse-mapped luma sample in the luma mapping procedure based on Tables 8 and 9. Equation 20 can be explained together with Table 10.
[0330] Equation 21 is for deriving inverse mapped luma samples in luma mapping by applying a linear reshaper. For example, lmcs_min_input in Equation 21 can be derived from at least one of Tables 21 to 24. By comparing Equation 20 with Equation 21, changes from the existing LMCS by applying a linear reshaper can be seen from the perspective of forward mapping.
[0331] Based on the example of the linear reshaper described above, the piecewise index identification procedure can be omitted. That is, in this example, since there is only one piece having reshaped luma pixels, the piecewise index identification procedure used in inverse luma mapping and chroma residual scaling can be eliminated. This can reduce the complexity of inverse luma mapping. In addition, the latency issue caused by relying on luma piecewise index identification can be eliminated during chroma residual scaling.
[0332] The above-described embodiment involving the use of a linear reshaper can provide the following advantages for LMCS: i) It simplifies the reshaper design in the encoder to prevent degradation due to abrupt changes between piecewise linear pieces; ii) It simplifies the decoder inverse mapping procedure by eliminating the piecewise index identification procedure; iii) It eliminates latency issues in chroma residual scaling caused by dependency on the corresponding luma block by eliminating the piecewise index identification procedure; iv) It reduces signaling overhead, making frequent reshaper updates more feasible; v) It eliminates loops where a 16-piece loop (e.g., a for construct) is required in many cases. For example, it reduces the number of division operations by lmcsCW[i] to derive InvScaleCoeff[i] to 1.
[0333] In another embodiment of this document, an LMCS based on flexible bins is proposed. Here, flexible bins may mean that the number of bins is not fixed to a predetermined number. In an existing embodiment, the number of bins in the LMCS is fixed to 16, and the 16 bins are evenly distributed among the input sample values. In this embodiment, a flexible number of bins is proposed, and the pieces (bins) are not evenly distributed among the original pixel values.
[0334] The following table exemplarily illustrates the syntax for LMCS data (data fields) according to this embodiment and the semantics of the syntax elements contained therein.
[0335] [Table 27]
[0336] [Table 28]
[0337] Referring to Table 27, information lmcs_num_bins_minus1 regarding the number of bins can be signaled. Referring to Table 28, lmcs_num_bins_minus1 + 1 is the same as the number of bins, and thus the number of bins is within the range from 1 to (1 << BitDepthY)-1. For example, lmcs_num_bins_minus1 or lmcs_num_bins_minus1 + 1 is a multiple of 2.
[0338] In the embodiments described with Tables 27 and 28, the number of pivot points can be derived based on lmcs_num_bins_minus1 (information regarding the number of bins) regardless of whether the reshaper is linear (signaling of lmcs_num_bins_minus1), and the input values and mapped values of the pivot points (LmcsPivot_input[i], LmcsPivot_mapped[i]) can be derived based on the sum of the signaled codeword values (lmcs_delta_input_cw[i], lmcs_delta_mapped_cw[i]) (where the initial input value LmcsPivot_input[0] and the initial output value LmcsPivot_mapped[0] are 0).
[0339] FIG. 15 shows an example of linear forward mapping in an embodiment of this document. FIG. 16 shows an example of inverse forward mapping in an embodiment of this document.
[0340] In the embodiment according to Figures 15 and 16, a method is proposed that supports both regular LMCS and linear LMCS. In one example according to this embodiment, regular LMCS and / or linear LMCS can be indicated based on the syntax element lmcs_is_linear. In the encoder, after the linear LMCS line is determined, the mapped value (e.g., the mapped value at pL in Figures 15 and 16) is divided into equal pieces (e.g., LmcsMaxBinIdx-lmcs_min_bin_idx+1). The codeword at binLmcsMaxBinIdx can be signaled using the syntax for the lmcs data or reshaper model described above.
[0341] The following table exemplarily illustrates the syntax for LMCS data (data fields) and the semantics for the syntax elements contained therein according to an example of this embodiment.
[0342] [Table 29]
[0343] [Table 30-1]
[0344] [Table 30-2]
[0345] The following table exemplarily shows the syntax for LMCS data (data fields) according to another example of this embodiment and the semantics of the syntax elements included therein.
[0346] [Table 31]
[0347] [Table 32-1]
[0348] [Table 32-2]
[0349] Referring to Tables 29 to 32, if lmcs_is_linear_flag is true, all lmcsDeltaCW[i] between lmcs_min_bin_idx and LmcsMaxBinIdx can have the same value. That is, among all pieces between lmcs_min_bin_idx and LmcsMaxBinIdx, lmcsCW[i] can have the same value. The scale, inverse scale, and chroma scale of all pieces between lmcs_min_bin_idx and lmcsMaxBinIdx are the same. If linear reshaper is true, there is no need to derive a piece index, and the scale or inverse scale can be used in one of the pieces.
[0350] The following table exemplarily illustrates the piecewise index identification procedure according to this embodiment.
[0351] [Table 33]
[0352] According to another embodiment of this document, the application of regular 16-piece PWL LMCS and linear LMCS can rely on high-level syntax (e.g., sequence level).
[0353] The following table exemplarily illustrates the syntax for an SPS according to this embodiment and the semantics for the syntax elements contained therein.
[0354] [Table 34]
[0355] [Table 35]
[0356] Referring to Tables 34 and 35, the availability of regular LMCS and / or linear LMCS can be determined (signaled) by the syntax elements included in the SPS. Referring to Table 35, either regular LMCS or linear LMCS can be used per sequence based on the syntax element sps_linear_lmcs_enabled_flag.
[0357] Additionally, whether one or both of linear and regular LMCSs are available is profile level dependent. In one example, for certain profile files (e.g., SDR profile files), only linear LMCSs may be allowed, and for other profile files (e.g., HDR profile files), only regular LMCSs may be allowed, and for other profile files, both regular and / or linear LMCSs may be allowed.
[0358] According to another embodiment of this document, an LMCS piecewise index identification procedure can be used in inverse luma mapping and chroma residual scaling. In this embodiment, the piecewise index identification procedure can be used for blocks for which chroma residual scaling is available, and can also be used for all luma samples in the reshaped (mapped) region. This embodiment aims to reduce the complexity for deriving the index.
[0359] The following table shows the identification (derivation) procedure for existing piecewise function indexes.
[0360] [Table 36]
[0361] In one example, the piecewise index identification procedure may classify input samples into at least two or more categories. For example, the input samples may be classified into three categories: first, second, and third categories. For example, the first category may indicate samples (values) smaller than LmcsPivot[lmcs_min_bin_idx+1], the second category may indicate samples (values) greater than or equal to LmcsPivot[LmcsMaxBinIdx], and the third category may indicate samples (values) between LmcsPivot[lmcs_min_bin_idx+1] and LmcsPivot[LmcsMaxBinIdx].
[0362] In this embodiment, a method for optimizing the identification procedure by removing category classification is proposed. Because the input to the piecewise index identification procedure is the reshaped (mapped) luma value, there should be no values that exceed the mapped values at the pivot points lmcs_min_bin_idx and LmcsMaxBinIdx+1. Therefore, the conditional procedure for classifying samples by category can be omitted from the existing piecewise index identification procedure, and a specific example will be described below with reference to a table.
[0363] In one example according to this embodiment, the identification procedure included in Table 36 can be replaced with one of Tables 37 or 38 below. Referring to Tables 37 and 38, the first two categories in Table 36 can be deleted, and the boundary value (second boundary value or ending point) in the iterative loop (for syntax) for the last category can be modified from LmcsMaxBinIdx to LmcsMaxBinIdx+1. In other words, the identification procedure can be simplified, and the complexity for deriving piecewise indexes can be reduced. Therefore, according to this embodiment, LMCS-related coding can be performed efficiently.
[0364] [Table 37]
[0365] [Table 38]
[0366] Referring to Table 37, a comparison procedure corresponding to the condition of the if clause (the mathematical expression corresponding to the condition of the if clause) may be repeatedly performed for all bin indexes from the minimum bin index to the maximum bin index. If the mathematical expression corresponding to the condition of the if clause is true, the bin index may be derived as an inverse mapping index for inverse luma mapping (or an inverse scaling index for chroma residual scaling). Based on the inverse mapping index, a modified restored luma sample (or a scaled chroma residual sample) may be derived.
[0367] The following drawings are created to explain a specific example of the present specification. The names of specific devices and names of specific signals / messages / fields shown in the drawings are provided for illustrative purposes only, and the technical features of the present specification are not limited to the specific names used in the following drawings.
[0368] 17 and 18 schematically illustrate an example of a video / image encoding method and related components according to an embodiment (or the like) of the present document. The method disclosed in FIG. 17 may be performed by the encoding device disclosed in FIG. 2. Specifically, for example, steps S1700 and S1710 of FIG. 17 may be performed by the prediction unit 220 of the encoding device, step S1720 may be performed by the residual processing unit 230 of the encoding device, step S1730 may be performed by the prediction unit 220 or the residual processing unit 230 of the encoding device, step S1740 may be performed by the residual processing unit 230 or the addition unit 250 of the encoding device, step S1750 may be performed by the residual processing unit 230 of the encoding device, and step S1760 may be performed by the entropy encoding unit 240 of the encoding device. The method disclosed in FIG. 17 may include the embodiments described above in this document.
[0369] 17, the encoding device may derive an inter prediction mode for a current block in a current picture (S1700). The encoding device may derive at least one of the various inter prediction modes disclosed in this document.
[0370] The encoding apparatus may generate predicted luma samples based on the inter prediction mode (S1710) by predicting original samples included in the current block.
[0371] The encoding apparatus may derive a predicted chroma sample. The encoding apparatus may derive the residual chroma sample based on the original chroma sample and the predicted chroma sample of the current block. For example, the encoding apparatus may derive the residual chroma sample based on a difference between the predicted chroma sample and the original chroma sample.
[0372] The encoding device can derive bins and LMCS codewords for luma mapping. The encoding device can derive the bins and / or LMCS codewords based on SDR or HDR.
[0373] The encoding device may derive LMCS-related information (S1720). The LMCS-related information may include bins and LMCS codewords for luma mapping.
[0374] The encoding device may generate a mapped predicted luma sample based on a mapping procedure for the luma sample (S1730). The encoding device may generate a mapped predicted luma sample based on a bin and / or an LMCS codeword for luma mapping. For example, the encoding device may derive an input value and a mapping value (output value) of a pivot point for luma mapping, and generate a mapped predicted luma sample based on the input value and the mapping value. As an example, the encoding device may derive a mapping index (idxY) based on a first predicted luma sample, and generate a first mapped predicted luma sample based on the input value and the mapping value of the pivot point corresponding to the mapping index. As another example, linear mapping (linear reshaping, linear LMCS) may be used, and a mapped predicted luma sample may be generated based on a forward mapping scaling factor derived from two pivot points in the linear mapping. Therefore, the index derivation procedure may be omitted due to the linear mapping.
[0375] The encoding device may generate scaled residual chroma samples. Specifically, the encoding device may derive a chroma residual scaling factor and generate scaled residual chroma samples based on the chroma residual scaling factor. Here, the chroma residual scaling at the encoding end may also be referred to as forward chroma residual scaling. Therefore, the chroma residual scaling factor derived by the encoding device may be referred to as a forward chroma residual scaling factor, and forward-scaled residual chroma samples may be generated.
[0376] The encoding device may generate reconstructed luma samples based on the mapped predicted luma samples. Specifically, the encoding device may sum the residual luma samples with the mapped predicted luma samples and generate reconstructed luma samples based on the summation result.
[0377] The encoding device may generate modified reconstructed luma samples based on an inverse mapping procedure for the luma samples (S1740). The encoding device may generate the modified reconstructed luma samples based on the bins for the luma mapping, the LMCS codeword, and the reconstructed luma samples. The encoding device may generate the modified reconstructed luma samples through an inverse mapping procedure for the reconstructed luma samples. For example, the encoding device may derive an inverse mapping index (e.g., invYIdx) based on a mapping value (e.g., LmcsPivot[i], i = lmcs_min_bin_idx...LmcsMaxBinIdx+1) assigned to each reconstructed luma sample and / or bin index in the inverse mapping procedure. The encoding device may generate the modified reconstructed luma samples based on the mapping value (LmcsPivot[invYIdx]) assigned to the inverse mapping index.
[0378] The encoding device may generate residual luma samples based on the mapped predicted luma samples. For example, the encoding device may derive residual luma samples based on a difference between the mapped predicted luma samples and original luma samples.
[0379] The encoding apparatus may derive residual information (S1750). For example, the encoding apparatus may generate residual information based on the mapped predicted luma sample and the modified restored luma sample. For example, the encoding apparatus may derive residual information based on the scaled residual chroma sample and / or the residual luma sample. The encoding apparatus may derive transform coefficients based on a transform procedure for the scaled residual chroma sample and / or the residual luma sample. For example, the transform procedure may include at least one of DCT, DST, GBT, or CNT. The encoding apparatus may derive quantized transform coefficients based on a quantization procedure for the transform coefficients. The quantized transform coefficients may have a one-dimensional vector form based on a coefficient scan order. The encoding apparatus may generate residual information representing the quantized transform coefficients. The residual information may be generated using various encoding methods, such as Exponential-Golomb, CAVLC, CABAC, etc.
[0380] The encoding device may encode image / video information (S1760). The image information may include information about LMCS data and / or residual information. For example, the LMCS-related information may include information about linear LMCS. In one example, at least one LMCS codeword may be derived based on the information about linear LMCS. The encoded video / image information may be output in the form of a bitstream. The bitstream may be transmitted to a decoding device via a network or a storage medium.
[0381] The image / video information may include various information according to embodiments of this document, for example, the image / video information may include at least one of the information disclosed in Tables 1 to 38 above.
[0382] In one embodiment, a minimum bin index and a maximum bin index may be derived based on information about the bins. For example, a mapping value based on bin indices from the minimum bin index to the maximum bin index may be derived based on information about the LMCS codeword. The mapping value may include a first mapping value (e.g., LmcsPivot[idxY]) and a second mapping value (e.g., LmcsPivot[idxYInv]). In the mapping procedure for the luma sample, a mapping index may be derived based on the predicted luma sample, and the mapped predicted luma sample may be generated using the first mapping value based on the mapping index. In the inverse mapping procedure for the luma sample, an inverse mapping index may be derived based on the mapping value based on bin indices from the minimum bin index to the maximum bin index, and the modified reconstructed luma sample may be generated using the second mapping value based on the inverse mapping index.
[0383] In one embodiment, a mapping value based on a bin index from a minimum bin index to a maximum bin index may be derived based on an LMCS codeword, For example, the mapped prediction luma samples and the modified reconstructed luma samples may be generated based on the mapping value based on the bin index from the minimum bin index to the maximum bin index.
[0384] In one embodiment, the inverse mapping index can be derived based on a comparison between a mapping value based on the bin index from the minimum bin index to the maximum bin index and the value of the restored luma sample.
[0385] In one embodiment, the bin index includes a first bin index that refers to a first mapping value, the first bin index is derived based on a comparison between a mapping value based on the bin index from the minimum bin index to the maximum bin index and the value of the restored luma sample, and at least one of the corrected restored luma samples can be derived based on the restored luma sample and the first mapping value.
[0386] In one embodiment, the bin index includes a first bin index that refers to a first mapping value, and the first bin index can be derived based on the comparison formula (lumaSample < LmcsPivot[idxYInv + 1]) included in Table 37. For example, in the formula, lumaSample represents the value of the target luma sample among the restored luma samples, idxYInv represents one of the bin indices, and LmcsPivot[idxYinv + 1] can represent one of the mapping values based on the bin index from the minimum bin index to the maximum bin index. At least one of the corrected restored luma samples can be derived based on the restored luma sample and the first mapping value. As an example, the bin index when the comparison formula (lumaSample < LmcsPivot[idxYInv + 1]) is true can be derived as the inverse mapping index.
[0387] In one embodiment, the comparison procedure (lumaSample < LmcsPivot[idxYInv + 1]) included in Table 37 can be performed for all of the bin indices from the minimum bin index to the maximum bin index.
[0388] In one embodiment, the image information includes residual information, a residual chroma sample is generated based on the residual information, an inverse scaling index is identified based on the LMCS-related information, a chroma residual scaling factor is derived based on the inverse scaling index, and a scaled residual chroma sample is generated based on the residual chroma sample and the chroma residual scaling factor.
[0389] In one embodiment, a residual chroma sample is generated based on the residual information, the bin index includes a first bin index indicating a first mapping value, the first bin index is derived based on a comparison between a mapping value based on the bin indexes from the minimum bin index to the maximum bin index and a value of the restored luma sample, a chroma residual scaling factor is derived based on the first bin index, a scaled residual chroma sample is generated based on the residual chroma sample and the chroma residual scaling factor, and a restored chroma sample may be generated based on the scaled residual chroma sample.
[0390] In one embodiment, the image information includes information about a linear LMCS, and the information about the LMCS data includes a linear LMCS flag indicating whether a linear LMCS is applied, and if the mapped predicted luma sample is generated based on the information about the linear LMCS, the value of the linear LMCS flag may be 1.
[0391] In one embodiment, the information about the linear LMCS may include information about a first pivot point (e.g., P1 in FIG. 12) and information about a second pivot point (e.g., P2 in FIG. 12). For example, the input value and mapping value of the first pivot point are a minimum input value and a minimum mapping value, respectively. The input value and mapping value of the second pivot point are a maximum input value and a maximum mapping value, respectively. Input values between the minimum input value and the maximum input value may be linearly mapped.
[0392] In one embodiment, the image information may include a sequence parameter set (SPS), which may include a linear LMCS availability flag indicating whether a linear LMCS is available.
[0393] In one embodiment, a minimum bin index (e.g., lmcs_min_bin_idx) and / or a maximum bin index (e.g., LmcsMaxBinIdx) may be derived based on the LMCS-related information. A first mapping value (LmcsPivot[lmcs_min_bin_idx]) may be derived based on the minimum bin index. A second mapping value (LmcsPivot[LmcsMaxBinIdx] or LmcsPivot[LmcsMaxBinIdx+1]) may be derived based on the maximum bin index. The value of the restored luma sample (e.g., lumaSample in Table 37 or Table 38) may range from the first mapping value to the second mapping value. For example, the values of all restored luma samples may range from the first mapping value to the second mapping value. For another example, the values of some of the restored luma samples may range from the first mapping value to the second mapping value.
[0394] In one embodiment, an encoding device can generate piecewise indices for chroma residual scaling, derive chroma residual scaling factors based on the piecewise indices, and generate scaled residual chroma samples based on the residual chroma samples and the chroma residual scaling factors.
[0395] In one embodiment, the chroma resistive scaling factor is a single chroma resistive scaling factor.
[0396] In one embodiment, the LMCS-related information may include an LMCS data field and information about a linear LMCS. The information about the linear LMCS may also be referred to as information about linear mapping. The LMCS data field may include a linear LMCS flag indicating whether the linear LMCS is applied. If the value of the linear LMCS flag is 1, the mapped predicted luma samples may be generated based on the information about the linear LMCS.
[0397] In one embodiment, the image information may include information about the maximum input value and information about the maximum mapping value. The maximum input value is the same as the value of the information about the maximum input value (e.g., lmcs_max_input in Table 21). The maximum mapping value is the same as the value of the information about the maximum mapping value (e.g., lmcs_max_mapped in Table 21).
[0398] In one embodiment, the information about the linear mapping may include information about an input delta value of the second pivot point (e.g., lmcs_max_input_delta in Table 23) and information about a mapped delta value of the second pivot point (e.g., lmcs_max_mapped_delta in Table 23). The maximum input value may be derived based on the input delta value of the second pivot point, and the maximum mapped value may be derived based on the mapped delta value of the second pivot point.
[0399] In one embodiment, the maximum input value and the maximum mapping value may be derived based on at least one of the formulas included in Table 24 above.
[0400] In one embodiment, generating the mapped prediction luma samples may include deriving a forward mapping scaling factor for the prediction luma samples (e.g., ScaleCoeffSingle as described above), and generating the mapped prediction luma samples based on the forward mapping scaling factor, where the forward mapping scaling factor is a single factor for the prediction luma samples.
[0401] In one embodiment, the forward mapping scaling factor may be derived based on at least one of the formulas included in Table 22 and / or Table 24 above.
[0402] In one embodiment, the mapped predicted luma samples may be derived based on at least one of the formulas included in Table 26 above.
[0403] In one embodiment, the encoding device can derive an inverse mapping scaling factor (e.g., the aforementioned InvScaleCoeffSingle) for the restored luma sample (e.g., the aforementioned lumaSample). Also, the encoding device can generate a corrected restored luma sample (e.g., invSample) based on the restored luma sample and the inverse mapping scaling factor. The inverse mapping scaling factor is a single factor for the restored luma sample.
[0404] In one embodiment, the inverse mapping scaling factor can be derived using a piecewise index derived based on the restored luma sample.
[0405] In one embodiment, the piecewise index can be derived based on Table 36 described above. That is, the comparison procedure (lumaSample < LmcsPivot[idxYInv+1]) included in Table 37 can be repeatedly executed from the case where the piecewise index is the minimum bin index to the case where the piecewise index is the maximum bin index.
[0406] In one embodiment, the inverse mapping scaling factor can be derived based on at least one mathematical formula included in Table 19, Table 20, Table 22, Table 23 described above, or Mathematical Formula 11 or Mathematical Formula 12.
[0407] In one embodiment, the corrected restored luma sample can be derived based on Mathematical Formula 21 described above.
[0408] In one embodiment, the LMCS-related information may include information regarding the number of bins for deriving the mapped predicted luma samples (e.g., lmcs_num_bins_minus1 in Table 27). For example, the number of pivot points for luma mapping can be set the same as the number of bins. In one example, the encoding device can generate the delta input values and delta mapping values of the pivot points for each of the number of bins. In one example, the input values and mapping values of the pivot points are derived based on the delta input values (e.g., lmcs_delta_input_cw[i] in Table 27) and the delta mapping values (e.g., lmcs_delta_mapped_cw[i] in Table 28), and the mapped predicted luma samples can be generated based on the input values (e.g., LmcsPivot_input[i] in Table 28, or InputPivot[i] in Table 10) and the mapping values (e.g., LmcsPivot_mapped[i] in Table 28, or LmcsPivot[i] in Table 10).
[0409] In one embodiment, the encoding device can derive the LMCS delta codeword based on at least one LMCS codeword included in the LMCS-related information and the original codeword (OrgCW), and can also derive the luma prediction sample mapped based on at least one LMCS codeword and the original codeword. In one example, the information regarding the linear mapping can include the information regarding the LMCS delta codeword.
[0410] In one embodiment, at least one LMCS codeword can be derived based on the addition of the LMCS delta codeword and OrgCW. For example, OrgCW is (1<<BitDepthY) / 16, where BitDepthY can indicate the luma bit depth. This embodiment can be performed based on Equation 12.
[0411] In one embodiment, the at least one LMCS codeword can be derived based on the addition of the LMCS delta codeword and OrgCW*(lmcs_max_bin_idx - lmcs_min_bin_idx + 1), where lmcs_max_bin_idx and lmcs_min_bin_idx are the maximum bin index and the minimum bin index, respectively, and OrgCW is (1<<BitDepthY) / 16. This embodiment can be performed based on Equations 15 and 16.
[0412] In one embodiment, the at least one LMCS codeword is a multiple of 2.
[0413] In one embodiment, when the luma bit depth (BitDepthY) of the restored luma sample is higher than 10, the at least one LMCS codeword is a multiple of 1<<(BitDepthY - 10).
[0414] In one embodiment, the at least one LMCS codeword is within the range from (OrgCW>>1) to (OrgCW<<1)-1.
[0415] FIGs. 19 and 20 schematically show an example of an image / video decoding method and related components according to an embodiment of this document. The method disclosed in FIG. 19 can be performed by the decoding device disclosed in FIG. 3. Specifically, for example, S1900 in FIG. 19 can be performed by the entropy decoding unit 310 of the decoding device, S1910 and S1920 can be performed by the prediction unit 330 of the decoding device, S1930 can be performed by the residual processing unit 320 or the prediction unit 330 of the decoding device, and S1940 can be performed by the residual processing unit 320, the prediction unit 330, and / or the addition unit 340 of the decoding device. The method disclosed in FIG. 19 can include the embodiments described above in this document.
[0416] As shown in FIG. 19, a decoding device may receive / acquire video / image information (S1900). The video / image information may include prediction mode information, LMCS-related information, and / or residual information. For example, the LMCS-related information may include information about the luma mapping (e.g., forward mapping, inverse mapping, linear mapping), information about chroma residual scaling, and / or an index (e.g., maximum bin index, minimum bin index, mapping index) related to the LMCS (or reshaping, reshaper). The decoding device may receive / acquire the image / video information via a bitstream.
[0417] The image / video information may include various information according to embodiments of this document, for example, the image / video information may include at least one of the information disclosed in Tables 1 to 38 above.
[0418] The decoding device may derive a prediction mode for the current block in the current picture based on the prediction mode information (S1910). The decoding device may derive at least one of various inter prediction modes disclosed in this document.
[0419] The decoding device may generate predicted luma samples (S1920). The decoding device may derive predicted luma samples of the current block based on a prediction mode. The decoding device may generate predicted luma samples by predicting original samples included in the current block.
[0420] The decoding device may generate a mapped predicted luma sample (S1930). The decoding device may generate the mapped predicted luma sample based on a mapping procedure for the luma sample. For example, the decoding device may derive input values and mapping values (output values) of pivot points for luma mapping, and generate a mapped predicted luma sample based on the input values and mapping values. As an example, the decoding device may derive a (forward) mapping index (idxY) based on a first predicted luma sample, and generate a first mapped predicted luma sample based on the input value and mapping value of the pivot point corresponding to the mapping index. As another example, linear mapping (linear reshaping, linear LMCS) may be used, and a mapped predicted luma sample may be generated based on a forward mapping scaling factor derived from two pivot points in the linear mapping. Therefore, the index derivation procedure may be omitted due to the linear mapping.
[0421] The decoding device may generate residual luma samples based on the residual information. For example, the decoding device may derive quantized transform coefficients based on the residual information. The quantized transform coefficients may have a one-dimensional vector form based on a coefficient scan order. The decoding device may derive transform coefficients based on an inverse quantization procedure for the quantized transform coefficients. The decoding device may derive residual samples based on an inverse transform procedure for the transform coefficients. The residual samples may include residual luma samples and / or residual chroma samples.
[0422] The decoding device may generate reconstructed luma samples. The decoding device may generate reconstructed luma samples based on the mapped predicted luma samples. Specifically, the decoding device may sum the residual luma samples with the mapped predicted luma samples and generate reconstructed luma samples based on the summation result.
[0423] The decoding device may generate modified reconstructed luma samples. The decoding device may generate modified reconstructed luma samples based on an inverse mapping procedure for the luma samples. The decoding device may generate modified reconstructed luma samples based on information about the LMCS data and the reconstructed luma samples (S1940). The decoding device may generate the modified reconstructed luma samples through an inverse mapping procedure for the reconstructed luma samples.
[0424] The decoding device may generate scaled residual chroma samples. Specifically, the decoding device may derive a chroma residual scaling factor and generate scaled residual chroma samples based on the chroma residual scaling factor. Here, the chroma residual scaling at the decoding end may also be referred to as inverse chroma residual scaling, as opposed to the encoding end. Therefore, the chroma residual scaling factor derived by the decoding device may be referred to as an inverse chroma residual scaling factor, and inverse-scaled residual chroma samples may be generated.
[0425] The decoding device may generate reconstructed chroma samples. The decoding device may generate reconstructed chroma samples based on the scaled residual chroma samples. Specifically, the decoding device may perform a prediction procedure on the chroma components and generate predicted chroma samples. The decoding device may generate reconstructed chroma samples based on a sum of the predicted chroma samples and the scaled residual chroma samples.
[0426] In one embodiment, a mapping value based on a bin index from a minimum bin index to a maximum bin index may be derived based on an LMCS codeword, For example, the mapped prediction luma samples and the modified reconstructed luma samples may be generated based on the mapping value based on the bin index from the minimum bin index to the maximum bin index.
[0427] In one embodiment, the LMCS-related information may include information about bins for mapping and inverse mapping and information about the LMCS codeword. For example, a minimum bin index and a maximum bin index may be derived based on the information about the bins. A mapping value based on bin indices from the minimum bin index to the maximum bin index may be derived based on the information about the LMCS codeword. The mapping value may include a first mapping value (e.g., LmcsPivot[idxY]) and a second mapping value (e.g., LmcsPivot[idxYInv]). In the mapping procedure for the luma sample, a mapping index (e.g., idxY) may be derived based on the predicted luma sample, and the mapped predicted luma sample may be generated using the first mapping value based on the mapping index. In the inverse mapping procedure for the luma sample, an inverse mapping index may be derived based on the mapping value based on bin indices from the minimum bin index to the maximum bin index, and the modified reconstructed luma sample may be generated using the second mapping value based on the inverse mapping index (e.g., idxYInv).
[0428] In one embodiment, the image information may include a sequence parameter set (SPS), which may include a linear LMCS availability flag indicating whether a linear LMCS is available.
[0429] In one embodiment, the bin index includes a first bin index pointing to a first mapping value, and the first bin index is derived based on a comparison between the mapping values based on the bin indexes from the minimum bin index to the maximum bin index and the value of the restored luma sample, and at least one of the modified restored luma sample can be derived based on the restored luma sample and the first mapping value.
[0430] In one embodiment, the bin index includes a first bin index that points to a first mapping value, and the first bin index can be derived based on the comparison formula (lumaSample < LmcsPivot[idxYInv+1]) included in Table 37. For example, in the formula, lumaSample represents the value of the target luma sample among the restored luma samples, idxYInv represents one of the bin indexes, and LmcsPivot[idxYinv+1] can represent one of the mapping values based on the bin indexes from the minimum bin index to the maximum bin index. At least one of the corrected restored luma samples can be derived based on the restored luma sample and the first mapping value. As an example, the bin index when the comparison formula (lumaSample < LmcsPivot[idxYInv+1]) is true can be derived as the inverse mapping index.
[0431] In one embodiment, the comparison procedure (lumaSample < LmcsPivot[idxYInv+1]) included in Table 37 can be performed for all of the bin indexes from the minimum bin index to the maximum bin index.
[0432] In one embodiment, the image information includes residual information, residual chroma samples are generated based on the residual information, an inverse scaling index is identified based on the LMCS related information, a chroma residual scaling factor is derived based on the inverse scaling index, and scaled residual chroma samples can be generated based on the residual chroma samples and the chroma residual scaling factor.
[0433] In one embodiment, a residual chroma sample is generated based on the residual information, the bin index includes a first bin index indicating a first mapping value, the first bin index is derived based on a comparison between a mapping value based on the bin indexes from the minimum bin index to the maximum bin index and a value of the restored luma sample, a chroma residual scaling factor is derived based on the first bin index, a scaled residual chroma sample is generated based on the residual chroma sample and the chroma residual scaling factor, and a restored chroma sample may be generated based on the scaled residual chroma sample.
[0434] In one embodiment, the information about the LMCS data may include information about a linear LMCS. The information about the linear LMCS may also be referred to as information about linear mapping. The information about the LMCS data may include a linear LMCS flag indicating whether a linear LMCS is applied. If the value of the linear LMCS flag is 1, the mapped predicted luma samples may be generated based on the information about the linear LMCS.
[0435] In one embodiment, the information about the linear LMCS may include information about a first pivot point (e.g., P1 in FIG. 12) and information about a second pivot point (e.g., P2 in FIG. 12). For example, the input value and mapping value of the first pivot point are a minimum input value and a minimum mapping value, respectively. The input value and mapping value of the second pivot point are a maximum input value and a maximum mapping value, respectively. Input values between the minimum input value and the maximum input value may be linearly mapped.
[0436] In one embodiment, a minimum bin index (e.g., lmcs_min_bin_idx) and / or a maximum bin index (e.g., LmcsMaxBinIdx) may be derived based on the LMCS-related information. A first mapping value (LmcsPivot[lmcs_min_bin_idx]) may be derived based on the minimum bin index. A second mapping value (LmcsPivot[LmcsMaxBinIdx] or LmcsPivot[LmcsMaxBinIdx+1]) may be derived based on the maximum bin index. The values of the restored luma samples (e.g., lumaSample in Table 37 or Table 38) range from the first mapping value to the second mapping value. In one example, the values of all restored luma samples range from the first mapping value to the second mapping value. In another example, the values of some of the restored luma samples range from the first mapping value to the second mapping value.
[0437] In one embodiment, a piecewise index (e.g., idxYInv in Table 36, Table 37, or Table 38) may be identified based on the LMCS-related information. A decoding device may derive a chroma residual scaling factor based on the piecewise index. A decoding device may generate scaled residual chroma samples based on the residual chroma samples and the chroma residual scaling factor.
[0438] In one embodiment, the chroma resistive scaling factor can be a single chroma resistive scaling factor.
[0439] In one embodiment, the image information may include information about the maximum input value and information about the maximum mapping value. The maximum input value is the same as the value of the information about the maximum input value (e.g., lmcs_max_input in Table 21). The maximum mapping value is the same as the value of the information about the maximum mapping value (e.g., lmcs_max_mapped in Table 21).
[0440] In one embodiment, the information about the linear mapping may include information about an input delta value of the second pivot point (e.g., lmcs_max_input_delta in Table 23) and information about a mapped delta value of the second pivot point (e.g., lmcs_max_mapped_delta in Table 23). The maximum input value may be derived based on the input delta value of the second pivot point, and the maximum mapped value may be derived based on the mapped delta value of the second pivot point.
[0441] In one embodiment, the maximum input value and the maximum mapping value may be derived based on at least one of the formulas included in Table 24 above.
[0442] In one embodiment, generating the mapped prediction luma samples may include deriving a forward mapping scaling factor for the prediction luma samples (e.g., ScaleCoeffSingle as described above), and generating the mapped prediction luma samples based on the forward mapping scaling factor, where the forward mapping scaling factor is a single factor for the prediction luma samples.
[0443] In one embodiment, the inverse mapping scaling factor may be derived using piecewise indices derived based on the reconstructed luma samples.
[0444] In one embodiment, the piecewise index can be derived based on Table 36 described above. That is, the comparison procedure (lumaSample < LmcsPivot[idxYInv+1]) included in Table 37 can be repeatedly executed from the case where the piecewise index is the minimum bin index to the case where the piecewise index is the maximum bin index.
[0445] In one embodiment, the forward mapping scaling factor can be derived based on at least one mathematical formula included in Table 22 and / or Table 24 described above.
[0446] In one embodiment, the mapped predicted luma sample can be derived based on at least one mathematical formula included in Table 26 described above.
[0447] In one embodiment, the decoding device can derive an inverse mapping scaling factor (e.g., InvScaleCoeffSingle described above) for the restored luma sample (e.g., lumaSample described above). Further, the decoding device can generate a corrected restored luma sample (e.g., invSample) based on the restored luma sample and the inverse mapping scaling factor. The inverse mapping scaling factor is a single factor for the restored luma sample.
[0448] In one embodiment, the inverse mapping scaling factor can be derived based on at least one mathematical formula included in Table 19, Table 20, Table 22, Table 24 described above, or Mathematical Formula 11 or Mathematical Formula 12.
[0449] In one embodiment, the corrected restored luma sample can be derived based on Mathematical Formula 21 described above.
[0450] In one embodiment, the LMCS-related information may include information regarding the number of bins for deriving the mapped predicted luma samples (e.g., lmcs_num_bins_minus1 in Table 27). For example, the number of pivot points for luma mapping can be set the same as the number of bins. In one example, the decoding device can generate the delta input values and delta mapping values of the pivot points for the number of bins respectively. In one example, the input values and mapping values of the pivot points are derived based on the delta input values (e.g., lmcs_delta_input_cw[i] in Table 27) and the delta mapping values (e.g., lmcs_delta_mapped_cw[i] in Table 28), and the mapped predicted luma samples can be generated based on the input values (e.g., LmcsPivot_input[i] in Table 28, or InputPivot[i] in Table 10) and the mapping values (e.g., LmcsPivot_mapped[i] in Table 28, or LmcsPivot[i] in Table 10).
[0451] In one embodiment, the decoding device can derive an LMCS delta codeword based on at least one LMCS codeword included in the LMCS-related information and the original codeword (OrgCW), and can also derive a luma prediction sample mapped based on at least one LMCS codeword and the original codeword. In one example, the information regarding the linear mapping can include information regarding the LMCS delta codeword.
[0452] In one embodiment, at least one LMCS codeword can be derived based on the addition of the LMCS delta codeword and OrgCW. For example, OrgCW is (1<<BitDepthY) / 16, where BitDepthY can indicate the luma bit depth. This embodiment can be performed based on Equation 12.
[0453] In one embodiment, the at least one LMCS codeword can be derived based on the sum of the LMCS delta codeword and OrgCW*(lmcs_max_bin_idx - lmcs_min_bin_idx + 1), where lmcs_max_bin_idx and lmcs_min_bin_idx are the maximum bin index and the minimum bin index respectively, and OrgCW is (1<<BitDepthY) / 16. This embodiment can be performed based on Formulas 15 and 16.
[0454] In one embodiment, the at least one LMCS codeword can be a multiple of 2.
[0455] In one embodiment, when the luma bit depth (BitDepthY) of the restored luma sample is higher than 10, the at least one LMCS codeword is a multiple of 1<<(BitDepthY - 10).
[0456] In one embodiment, the at least one LMCS codeword is within the range from (OrgCW>>1) to (OrgCW<<1)-1.
[0457] In the foregoing embodiments, the method is described based on a flowchart as a series of steps or blocks, but the corresponding embodiments are not limited to the order of the steps. A certain step can occur in a different order from the steps described above, or simultaneously. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, and different steps may be included, or one or more steps of the flowchart can be deleted without affecting the scope of the embodiments of this document.
[0458] The method according to the embodiments of this document described above can be realized in the form of software, and the encoding device and / or decoding device according to this document can be included in devices that perform image processing, such as TVs, computers, smartphones, set-top boxes, display devices, etc.
[0459] When an embodiment of this document is implemented in software, the method described above may be implemented with modules (processes, functions, etc.) that perform the functions described above. The modules may be stored in memory and executed by a processor. The memory may be internal or external to the processor and may be coupled to the processor in various well-known ways. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described herein may be implemented and performed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each figure may be implemented and performed on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information on instructions) or algorithms for implementation may be stored on a digital storage medium.
[0460] In addition, the decoding device and encoding device to which the embodiments of this document are applied may be included in a multimedia broadcast transmitting / receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video interaction device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a customized video (VoD) service providing device, an over-the-top (OTT) video (over-the-top) device, an internet streaming service providing device, a three-dimensional (3D) video device, a virtual reality (VR) device, an augmented reality (AR) device, an image telephone video device, a transportation terminal (e.g., a vehicle terminal (including an autonomous vehicle), an airplane terminal, a ship terminal, etc.), a medical video device, etc., and may be used to process video signals or data signals. For example, over-the-top (OTT) video (over-the-top) devices may include a game console, a Blu-ray player, an internet access TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.
[0461] In addition, a processing method to which an embodiment of this document is applied may be produced in the form of a computer-executable program and stored on a computer-readable recording medium. Multimedia data having a data structure according to an embodiment of this document may also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices on which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The computer-readable recording medium may also include media implemented in the form of a carrier wave (e.g., transmission via the Internet). The bitstream generated by the encoding method may be stored on a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0462] Furthermore, the embodiments of the present document may be implemented in a computer program product by program code, which may be executed by a computer in accordance with the embodiments of the present document. The program code may be stored on a computer-readable carrier.
[0463] FIG. 21 illustrates an example of a content streaming system in which the embodiments disclosed herein can be applied.
[0464] Referring to FIG. 21, a content streaming system to which the embodiments of this document are applied can broadly include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0465] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted.
[0466] The bitstream can be generated by an encoding method or a bitstream generation method to which an embodiment of this document is applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0467] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which controls commands and responses between devices in the content streaming system.
[0468] The streaming server can receive content from a media repository and / or an encoding server. For example, if content is received from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.
[0469] Examples of the user device include a mobile phone, a smartphone, a laptop computer, a digital broadcasting terminal, a PDA (personal digital assistant), a PMP (portable multimedia player), a navigation system, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, a smart glass, a head mounted display (HMD)), a digital TV, a desktop computer, a digital signage, etc.
[0470] Each server in the content streaming system can be operated as a distributed server, in which case data received by each server can be processed in a distributed manner.
[0471] The claims described in this specification may be combined in various ways. For example, the technical features of the method claims herein may be combined to be realized as an apparatus, and the technical features of the apparatus claims herein may be combined to be realized as a method. Furthermore, the technical features of the method claims herein and the technical features of the apparatus claims herein may be combined to be realized as an apparatus, and the technical features of the method claims herein and the technical features of the apparatus claims herein may be combined to be realized as a method.
Claims
1. An image decoding method performed by a decoding device, Obtaining image information including prediction mode information, residual information, and LMCS (luma mapping with chroma scaling) related information from a bitstream; deriving an inter prediction mode for a current block in a current picture based on the prediction mode information; deriving motion information for the current block based on the inter prediction mode; generating a predictive luma sample for the current block based on the motion information for the current block; generating mapped predicted luma samples based on a mapping procedure for the predicted luma samples; generating residual luma samples for the current block based on the residual information; generating reconstructed luma samples based on the mapped prediction luma samples and the residual luma samples; generating modified reconstructed luma samples based on an inverse mapping procedure for the reconstructed luma samples; Including, the LMCS-related information includes information about bins for mapping and inverse mapping, and information about LMCS codewords; deriving a minimum bin index and a maximum bin index based on information about the bins; an LMCS codeword is derived based on information about the LMCS codeword; a mapping value indexed by an index from the minimum bin index to the maximum bin index is derived based on the LMCS codeword; the mapping procedure for the predictive luma sample, a mapping index is derived based on the predictive luma sample, and the mapped predictive luma sample is generated using a mapping value indexed by the mapping index; In the inverse mapping procedure for the reconstructed luma samples, an inverse mapping index is derived based on the mapping values indexed by bin indexes from the minimum bin index to the maximum bin index, and the modified reconstructed luma samples are generated using the mapping values indexed by the inverse mapping indexes; To derive the inverse mapping index, the inverse mapping procedure for the reconstructed luma samples comprises: determining whether a mapping value is greater than a value of the reconstructed luma sample when the inverse mapping index is equal to the smallest bin index; determining whether a mapping value is greater than a value of the reconstructed luma sample when the inverse mapping index is equal to a maximum bin index minus one; determining whether a mapping value is greater than a value of the reconstructed luma sample when the inverse mapping index is equal to the maximum bin index based on the determination that the mapping value is not greater than a value of the reconstructed luma sample when the inverse mapping index is equal to the maximum bin index minus one; An image decoding method, comprising:
2. An image encoding method performed by an encoding device, deriving an inter prediction mode; generating motion information for a current block based on the inter prediction mode; generating a predictive luma sample for the current block based on the motion information for the current block; generating mapped predicted luma samples based on a mapping procedure for the predicted luma samples; generating luma mapping with chroma scaling (LMCS) related information based on the mapped predicted luma samples; generating residual luma samples based on the mapped predicted luma samples; generating residual information based on the residual luma samples; generating reconstructed luma samples based on the mapped prediction luma samples and the residual luma samples; generating modified reconstructed luma samples based on an inverse mapping procedure for the reconstructed luma samples; encoding image information including the LMCS-related information and the residual information; Including, the LMCS-related information includes information about bins for mapping and inverse mapping and information about LMCS codewords; deriving a minimum bin index and a maximum bin index based on information about the bins; an LMCS codeword is derived based on information about the LMCS codeword; mapping values indexed by bin indices from the minimum bin index to the maximum bin index are derived based on the LMCS codeword; the mapping procedure for the predictive luma sample, a mapping index is derived based on the predictive luma sample, and the mapped predictive luma sample is generated using a mapping value indexed by the mapping index; In the inverse mapping procedure for the reconstructed luma samples, an inverse mapping index is derived based on the mapping values indexed by bin indexes from the minimum bin index to the maximum bin index, and the modified reconstructed luma samples are generated using the mapping values indexed by the inverse mapping indexes; To derive the inverse mapping index, the inverse mapping procedure for the reconstructed luma samples comprises: determining whether a mapping value is greater than a value of the reconstructed luma sample when the inverse mapping index is equal to the smallest bin index; determining whether a mapping value is greater than a value of the reconstructed luma sample when the inverse mapping index is equal to a maximum bin index minus one; determining whether a mapping value is greater than a value of the reconstructed luma sample when the inverse mapping index is equal to the maximum bin index based on the determination that the mapping value is not greater than a value of the reconstructed luma sample when the inverse mapping index is equal to the maximum bin index minus one; An image encoding method, including:
3. 1. A method for transmitting data for image information, comprising: obtaining a bitstream of the image information including LMCS (luma mapping with chroma scaling)-related information and residual information, wherein the bitstream is generated based on: deriving an inter prediction mode; generating motion information for a current block based on the inter prediction mode; generating a predicted luma sample for the current block based on the motion information for the current block; generating a mapped predicted luma sample based on a mapping procedure for the predicted luma sample; generating LMCS-related information based on the mapped predicted luma sample; generating a residual luma sample based on the mapped predicted luma sample; generating the residual information based on the residual luma sample; generating a reconstructed luma sample based on the mapped predicted luma sample and the residual luma sample; generating a modified reconstructed luma sample based on an inverse mapping procedure for the reconstructed luma sample; and encoding the image information including the LMCS-related information and the residual information; transmitting the data including the bitstream of the image information including the LMCS-related information and the residual information; Including, the LMCS-related information includes information about bins for mapping and inverse mapping and information about LMCS codewords; deriving a minimum bin index and a maximum bin index based on information about the bins; an LMCS codeword is derived based on information about the LMCS codeword; mapping values indexed by bin indices from the minimum bin index to the maximum bin index are derived based on the LMCS codeword; the mapping procedure for the predictive luma sample, a mapping index is derived based on the predictive luma sample, and the mapped predictive luma sample is generated using a mapping value indexed by the mapping index; In the inverse mapping procedure for the reconstructed luma samples, an inverse mapping index is derived based on the mapping values indexed by bin indexes from the minimum bin index to the maximum bin index, and the modified reconstructed luma samples are generated using the mapping values indexed by the inverse mapping indexes; To derive the inverse mapping index, the inverse mapping procedure for the reconstructed luma samples comprises: determining whether a mapping value is greater than a value of the reconstructed luma sample when the inverse mapping index is equal to the smallest bin index; determining whether a mapping value is greater than a value of the reconstructed luma sample when the inverse mapping index is equal to a maximum bin index minus one; determining whether a mapping value is greater than a value of the reconstructed luma sample when the inverse mapping index is equal to the maximum bin index based on the determination that the mapping value is not greater than a value of the reconstructed luma sample when the inverse mapping index is equal to the maximum bin index minus one; a transmission method,
Citation Information
Patent Citations
Integrated image reshaping and video coding
WO2019006300A1