Image coding method, computer readable storage medium and data transmission method
By using LMCS processing, including constraining the LMCS codeword range and using a single chroma residual scaling factor, the problem of low compression efficiency for high-resolution images/videos is solved, achieving efficient image/video encoding and decoding, improving visual quality and reducing resource consumption.
Patent Information
- Application Number
- CN202310777675.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-24
- Filing Date
- 2020-06-24
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2040-06-24
AI Technical Summary
Existing technologies are inefficient and costly in the compression, transmission, and storage of high-resolution, high-quality images/videos, making it difficult to meet the needs of virtual reality and immersive media.
It employs Luminance Mapping and Chromaticity Scaling (LMCS) processing, constrains the LMCS codeword range, uses a single chroma residual scaling factor and linear mapping, flexibly utilizes binaries, simplifies index derivation processing, adapts to the encoding of dual-tree structure blocks, and limits the number and range of APS.
It improves image/video compression efficiency, enhances subjective/objective visual quality, reduces resource requirements and complexity, and supports efficient encoding of dual-tree structure blocks.
Smart Images

Figure CN116600144B_ABST
Abstract
Description
[0001] This application is a divisional application of the original application No. 202080058312.1 (International Application No. PCT / KR2020 / 008208, filed on June 24, 2020, entitled "Video or image encoding based on mapping of luma samples and scaling of chroma samples") for an invention. TECHNICAL FIELD
[0002] The technology of the present document relates to video or image encoding based on mapping of luma samples and scaling of chroma samples. BACKGROUND
[0003] Recently, there has been an increasing demand for high-resolution, high-quality image / video such as 4K or 8K or higher ultra-high-definition (UHD) image / video in various fields. As image / video data has high resolution and high quality, the amount of information or bits to be transmitted increases relative to existing image / video data, and thus, transmitting image data using a medium such as an existing wired / wireless broadband line or an existing storage medium or storing image / video data using the existing storage medium increases transmission and storage costs.
[0004] In addition, interest and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms have recently increased, and broadcasting of image / video having different characteristics from real images (for example, game images) has increased.
[0005] Therefore, a very efficient image / video compression technique is needed to effectively compress, transmit, store, and reproduce information of high-resolution, high-quality image / video having various characteristics as described above.
[0006] In addition, a luma mapping with chroma scaling (LMCS) process is performed to improve compression efficiency and increase subjective / objective visual quality, and discussions are made on how to efficiently apply the LMCS process. SUMMARY
[0007] TECHNICAL SOLUTION
[0008] According to embodiments of the present document, a method and apparatus for increasing image encoding efficiency are provided.
[0009] According to embodiments of the present document, an efficient filter application method and apparatus are provided.
[0010] According to embodiments of the present document, an efficient LMCS application method and apparatus are provided.
[0011] According to embodiments of the present document, an LMCS codeword (or a range thereof) can be constrained.
[0012] According to an embodiment of the present document, a single chroma residual scaling factor directly signaled in the chroma scaling of LMCS can be used.
[0013] According to an embodiment of the present document, a linear mapping (linear LMCS) can be used.
[0014] According to an embodiment of the present document, information about the pivot point required for the linear mapping can be explicitly signaled.
[0015] According to an embodiment of the present document, a flexible number of bins can be used for the luma mapping.
[0016] According to an embodiment of the present document, the index derivation process for inverse luma mapping and / or chroma residual scaling can be simplified.
[0017] According to an embodiment of the present document, the LMCS process can be applied even when the luma block and the chroma block in one coding tree unit (CTU) have separate block tree structures (dual tree structure).
[0018] According to an embodiment of the present document, the number of LMCS APSs can be limited.
[0019] According to an embodiment of the present document, the range occupied by the LMCS APSs among all APSs can be limited.
[0020] According to an embodiment of the present document, a video / image decoding method performed by a decoding device is provided.
[0021] According to an embodiment of the present document, a decoding device for performing video / image decoding is provided.
[0022] According to an embodiment of the present document, a video / image encoding method performed by an encoding device is provided.
[0023] According to an embodiment of the present document, an encoding device for performing video / image encoding is provided.
[0024] According to an embodiment of the present document, a computer-readable digital storage medium having stored therein encoded video / image information generated by the video / image encoding method disclosed in at least one embodiment of the present document is provided.
[0025] According to an embodiment of the present document, a computer-readable digital storage medium having stored therein encoded information or encoded video / image information that causes a decoding device to perform the video / image decoding method disclosed in at least one embodiment of the present document is provided.
[0026] Advantageous effects
[0027] According to embodiments of the present document, total image / video compression efficiency can be improved.
[0028] According to embodiments of the present document, subjective / objective visual quality can be improved by efficient filtering.
[0029] According to embodiments of the present document, LMCS processing for image / video coding can be efficiently performed.
[0030] According to embodiments of the present document, resources / cost (of software or hardware) required for LMCS processing can be minimized.
[0031] According to embodiments of the present document, hardware implementation of LMCS processing can be facilitated.
[0032] According to embodiments of the present document, division operations required for derivation of LMCS codewords in mapping (reshaping) can be removed or minimized by constraining LMCS codewords (or their range).
[0033] According to embodiments of the present document, a single chroma residual scaling factor can be used to remove the delay identified according to the segment index.
[0034] According to embodiments of the present document, chroma residual scaling processing can be performed using linear mapping in LMCS without relying on (reconstruction of) luma blocks, thus removing the delay in scaling.
[0035] According to embodiments of the present document, mapping efficiency in LMCS can be increased.
[0036] According to embodiments of the present document, complexity of LMCS can be reduced by simplifying index derivation processing for inverse luma mapping and / or chroma residual scaling, thus video / image coding efficiency can be increased.
[0037] According to embodiments of the present document, LMCS process can be performed even for blocks with dual tree structure, thus efficiency of LMCS can be increased. Furthermore, coding performance (e.g., objective / subjective image quality) of blocks with dual tree structure can be improved.
[0038] According to embodiments of the present document, complexity of LMCS can be reduced and LMCS can consume (use) less resources (e.g., memory) since the number of LMCS APS is limited.
[0039] According to embodiments of the present document, complexity of LMCS can be reduced and video / image coding efficiency can be increased by limiting the range occupied by LMCS APS among all APS. BRIEF DESCRIPTION OF DRAWINGS
[0040] Figure 1An example of a video / image encoding system to which embodiments of the present document can be applied is shown.
[0041] Figure 2 is a diagram schematically showing a configuration of a video / image encoding apparatus to which embodiments of the present document can be applied.
[0042] Figure 3 is a diagram schematically showing a configuration of a video / image decoding apparatus to which embodiments of the present document can be applied.
[0043] Figure 4 An exemplary block tree structure is shown.
[0044] Figure 5 An exemplary layered structure of an encoded image / video is shown.
[0045] Figure 6 An exemplary layered structure of a CVS according to embodiments of the present document is shown.
[0046] Figure 7 An exemplary LMCS structure according to embodiments of the present document is shown.
[0047] Figure 8 An LMCS structure according to another embodiment of the present document is shown.
[0048] Figure 9 A graph showing an exemplary forward mapping is shown.
[0049] Figure 10 is a flowchart showing a method of deriving a chroma residual scaling index according to embodiments of the present document.
[0050] Figure 11 A linear fit of a pivot point according to embodiments of the present document is shown.
[0051] Figure 12 An example of linear reshaping (or linear reshaping, linear mapping) according to embodiments of the present document is shown.
[0052] Figure 13 An example of a linear forward mapping in embodiments of the present document is shown.
[0053] Figure 14 An example of an inverse forward mapping in embodiments of the present document is shown.
[0054] Figure 15 and Figure 16 An example of a video / image encoding method and related components according to embodiments of the present document is shown schematically.
[0055] Figure 17 and Figure 18An example of an image / video decoding method and related components according to an embodiment of the present document is schematically illustrated.
[0056] Figure 19 An example of a content streaming system to which embodiments disclosed in the present document can be applied is illustrated. DETAILED DESCRIPTION
[0057] The present document can be modified in various forms, and specific embodiments thereof will be described and illustrated in the drawings. However, these embodiments are not intended to limit the present document. The terms used in the following description are used to describe specific embodiments only and are not intended to limit the present document. Singular expressions include plural expressions as long as they are clearly different from the context. Terms such as "include" and "have" are intended to indicate that there is a feature, number, step, operation, element, component, or a combination thereof described in the following description, and it should be understood that the possibility of existence or addition of one or more different features, numbers, steps, operations, elements, components, or a combination thereof is not excluded.
[0058] In addition, the various configurations in the drawings described in the present document are independently illustrated for the convenience of describing different characteristic functions, and do not mean that the various configurations are implemented as separate hardware or separate software. For example, two or more of the various components among them can be combined to form one component, or one component can be divided into multiple components. Embodiments in which the various components are integrated and / or separated are also included in the scope of disclosure of the present document.
[0059] Hereinafter, examples of the present embodiments will be described in detail with reference to the accompanying drawings. In addition, like reference numerals are used to refer to like elements throughout the drawings, and the same description will be omitted with respect to similar elements.
[0060] Figure 1 An example of a video / image encoding system to which embodiments of the present document can be applied is illustrated.
[0061] REFERENCE Figure 1 The video / image encoding system can include a first device (a source device) and a second device (a receiving device). The source device can transmit encoded video / image information or data in the form of a file or a stream to the receiving device through a digital storage medium or a network.
[0062] The source device can include a video source, an encoding apparatus, and a transmitter. The receiving device can include a receiver, a decoding apparatus, and a renderer. The encoding apparatus can be referred to as a video / image encoding apparatus, and the decoding apparatus can be referred to as a video / image decoding apparatus. The transmitter can be included in the encoding apparatus. The receiver can be included in the decoding apparatus. The renderer can include a display, and the display can be configured as a separate device or an external component.
[0063] The video source can acquire a video / image through a process of capturing, synthesizing, or generating a video / image. The video source can include a video / image capturing device and / or a video / image generating device. For example, the video / image capturing device can include one or more cameras, a video / image archive including previously captured videos / images, etc. For example, the video / image generating device can include a computer, a tablet, and a smartphone, and can generate a video / image (electronically). For example, a virtual video / image can be generated through a computer, etc. In this case, the video / image capturing process can be replaced by a process of generating related data.
[0064] The encoding device can encode an input video / image. For compression and encoding efficiency, the encoding device can perform a series of processes such as prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0065] The transmitter can transmit the encoded image / image information or data output in the form of a bitstream to a receiver of a receiving device in the form of a file or a stream through a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include an element for generating a media file through a predetermined file format, and can include an element for transmission through a broadcasting / communication network. The receiver can receive / extract a bitstream and transmit the received bitstream to a decoding device.
[0066] The decoding device can decode a video / image by performing a series of processes such as dequantization, inverse transformation, and prediction corresponding to the operations of the encoding device.
[0067] The renderer can render the decoded video / image. The rendered video / image can be displayed through a display.
[0068] This document relates to video / image encoding. For example, the methods / embodiments disclosed in this document can be applied to methods disclosed in the Versatile Video Coding (VVC) standard, the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the 2nd generation Audio Video Coding standard (AVS2), or the next generation video / image encoding standard (e.g., H.267, H.268, etc.).
[0069] This document proposes various embodiments of video / image encoding, and the above-described embodiments can also be executed in combination with each other unless otherwise specified.
[0070] In this document, a video can refer to a series of pictures over time. A picture generally refers to a unit of one picture representing a certain time range, and a slice / tile refers to a unit that constitutes a part of a picture in terms of coding. A slice / tile can include one or more coding tree units (CTUs). One picture can be composed of one or more slices / tiles. One picture can be composed of one or more tile groups. One tile group can include one or more tiles. A tile can represent a rectangular region of CTU rows within a tile in a picture. A tile can be partitioned into multiple tiles, each of which can be composed of one or more CTU rows within a tile. A tile that is not partitioned into multiple tiles can also be referred to as a tile. Tile scanning can refer to a particular order of ordering CTUs that partition a picture, in which CTUs can be ordered in a raster scan of CTUs within a tile, and tiles within a picture can be ordered consecutively in a raster scan of tiles of a picture, and pictures within a picture can be ordered consecutively in a raster scan of pictures of a picture. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. A tile column is a rectangular region of CTUs with a height equal to a height of a picture and a width specified by a syntax element in a picture parameter set. A tile row is a rectangular region of CTUs with a height specified by a syntax element in a picture parameter set and a width equal to a width of a picture. Tile scanning is a particular order of ordering CTUs that partition a picture, in which CTUs are ordered consecutively in a raster scan of CTUs of a tile, and pictures within a picture are ordered consecutively in a raster scan of tiles of a picture. A slice includes an integer number of tiles of a picture that can be contained exclusively in a single NAL unit. A slice can be composed of a consecutive sequence of multiple complete tiles or only complete tiles of a single tile. In this document, a tile group and a slice can be used instead of each other. For example, in this document, a tile group / tile group header can be referred to as a slice / slice header.
[0071] In addition, one picture can be divided into two or more sub-pictures. A sub-picture can be a rectangular region of one or more slices within a picture.
[0072] A pixel or pel can mean a minimum unit constituting one picture (or image). In addition, a "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component.
[0073] A unit can represent a basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to the region. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. In some cases, a unit can be used interchangeably with terms such as a block or a region. In general, an MxN block can include a set (or an array) of M columns and N rows of samples (or sample array) or transform coefficients. Alternatively, a sample can mean a pixel value in a spatial domain and can mean a transform coefficient in a frequency domain when such a pixel value is transformed into the frequency domain.
[0074] In this document, "A or B" can mean "A and / or B". In other words, "A or B" in this document can be interpreted as "A and / or B". For example, in this document, "A, B, or C (A, B, or C)" means "only A", "only B", "only C", or "any combination of A, B, and C".
[0075] A slash ( / ) or a virgule (,) used in this document can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B, or C".
[0076] In this document, "at least one of A and B" can mean "only A", "only B", or "both A and B". In addition, in this document, the expression "at least one of A or B" or "at least one of A and / or B" can be interpreted identically to "at least one of A and B".
[0077] In addition, in this document, "at least one of A, B, and C" means "only A", "only B", "only C", or "any combination of A, B, and C". In addition, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0078] In addition, a parenthesis used in this document can mean "for example". Specifically, when indicating "prediction (intra prediction)", "intra prediction" can be proposed as an example of "prediction". In other words, "prediction" in this document is not limited to "intra prediction", and "intra prediction" can be proposed as an example of "prediction". In addition, even when indicating "prediction (i.e., intra prediction)", "intra prediction" can be proposed as an example of "prediction".
[0079] Technological features described separately in this document can be implemented separately or simultaneously.
[0080] Figure 2FIG. 1 is a diagram schematically illustrating a configuration of a video / image encoding apparatus to which embodiments of the present document can be applied. Hereinafter, the so-called video encoding apparatus can include an image encoding apparatus.
[0081] Referring to Figure 2 The encoding apparatus 200 includes an image partitioner 210, a predictor 220, a residual processor 230, and an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor 231. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. According to embodiments, the image partitioner 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 can be configured by at least one hardware component (e.g., an encoder chipset or a processor). In addition, the memory 270 can include a decoded picture buffer (DPB), or can be configured by a digital storage medium. The hardware component can further include the memory 270 as an internal / external component.
[0082] The image partitioner 210 can partition an input image (or picture or frame) input to the encoding apparatus 200 into one or more processors. For example, the processor can be referred to as a coding unit (CU). In this case, the coding unit can be recursively partitioned from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quad tree binary tree ternary (QTBT TT) structure. For example, one coding unit can be partitioned into a plurality of coding units at a deeper depth based on a quad tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad tree structure can be applied first, and the binary tree structure and / or the ternary structure can be applied later. Alternatively, the binary tree structure can be applied first. The encoding process according to the present disclosure can be performed based on the final coding unit that is no longer partitioned. In this case, the largest coding unit can be used as the final coding unit based on coding efficiency according to the image characteristics, or if necessary, the coding unit can be recursively partitioned into coding units at a deeper depth and the coding unit having an optimal size can be used as the final coding unit. Here, the encoding process can include processes of prediction, transformation, and reconstruction (to be described later). As another example, the processor can further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit can be split or partitioned from the final coding unit described above. The prediction unit can be a unit of sample prediction, and the transform unit can be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0083] In some cases, the term of unit can be used interchangeably with the term of block or area. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. The samples can represent pixels or pixel values, can represent only pixels / pixel values of a luminance component, or can represent only pixels / pixel values of a chrominance component. The samples can be used as a term corresponding to one picture (or image) of pixels or picture elements.
[0084] In the encoding apparatus 200, a prediction signal (prediction block, prediction sample array) output from the inter-predictor 221 or the intra-predictor 222 is subtracted from an input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the transformer 232. In this case, as illustrated, a unit that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) in the encoding apparatus 200 can be referred to as a subtracter 231. The predictor can perform prediction on a block to be processed (hereinafter, referred to as a current block) and generate a prediction block including predicted samples of the current block. The predictor can determine whether to apply intra-prediction or inter-prediction based on the current block or CU. As described later in the description of each prediction mode, the predictor can generate various information (for example, prediction mode information) related to prediction and transmit the generated information to the entropy encoder 240. The information about prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0085] The intra-predictor 222 can predict the current block with reference to samples in the current picture. Depending on the prediction mode, the referred samples can be located in the vicinity of the current block or can be spaced apart. In intra-prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. For example, the non-directional modes can include a DC mode and a planar mode. For example, depending on the degree of detail of the prediction direction, the directional modes can include 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or less directional prediction modes can be used according to settings. The intra-predictor 222 can determine a prediction mode applied to the current block using a prediction mode applied to a neighboring block.
[0086] The inter predictor 221 can derive a prediction block of a current block based on a reference block (a reference sample array) designated by a motion vector on a reference picture. Here, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in units of a block, a sub-block, or a sample based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring block can be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, the inter predictor 221 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of a skip mode and a merge mode, the inter predictor 221 can use motion information of the neighboring blocks as motion information of the current block. In the skip mode, unlike the merge mode, a residual signal can not be transmitted. In the case of a motion vector prediction (MVP) mode, a motion vector of the neighboring block can be used as a motion vector predictor, and a motion vector of the current block can be indicated by signaling a motion vector difference.
[0087] The predictor 220 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra prediction or inter prediction to predict one block, but also simultaneously apply both intra prediction and inter prediction. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can predict a block based on an intra block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or the palette mode can be used for content image / video encoding of a game or the like, such as screen content coding (SCC). The IBC basically performs prediction in the current picture, but can be performed similarly to inter prediction, such that a reference block is derived in the current picture. That is, the IBC can use at least one inter prediction technique described in the disclosure. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, sample values within a picture can be signaled based on information about a palette table and a palette index.
[0088] The prediction signal generated by the predictor (including the inter-predictor 221 and / or the intra-predictor 222) can be used to generate a reconstructed signal or to generate a residual signal. The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a karhunen-loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, the GBT means a transform obtained from a graph when relationship information between pixels is represented by the graph. The CNT refers to a transform generated based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process can be applied to square blocks of pixels having the same size or can be applied to blocks having variable sizes other than squares.
[0089] The quantizer 233 can quantize the transform coefficients and transmit them to the entropy encoder 240, and the entropy encoder 240 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 233 can rearrange the quantized transform coefficients of a block type into a one-dimensional vector form based on a coefficient scan order, and generate information about the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients. Information about the transform coefficients can be generated. The entropy encoder 240 can perform various encoding methods such as exponential Golomb, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), or the like. The entropy encoder 240 can encode information (for example, values of syntax elements, or the like) required for video / image reconstruction, together with or separately from the quantized transform coefficients. The encoded information (for example, encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer). The video / image information can further include information about various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information can further include general constraint information. In the disclosure, information and / or syntax elements transmitted / signaled from the encoding apparatus to the decoding apparatus can be included in the video / picture information. The video / image information can be encoded through the encoding process described above and included in the bitstream. The bitstream can be transmitted via a network or can be stored in a digital storage medium. The network can include a broadcast network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, or the like. A transmitter (not shown) that transmits a signal output from the entropy encoder 240 and / or a storage unit (not shown) that stores the signal can be included as an internal / external element of the encoding apparatus 200, and alternatively, the transmitter can be included in the entropy encoder 240.
[0090] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, a residual signal (a residual block or a residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients via the dequantizer 234 and the inverse transformer 235. The adder 250 adds the reconstructed residual signal to the prediction signal output from the inter-predictor 221 or the intra-predictor 222 to generate a reconstructed signal (a reconstructed picture, a reconstructed block, a reconstructed sample array). If there is no residual for a block to be processed (e.g., in the case where a skip mode is applied), the prediction block can be used as the reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. As described below, the generated reconstructed signal can be used for intra prediction of a next block to be processed in a current picture and can be filtered for inter prediction of a next picture.
[0091] In addition, luma mapping with chroma scaling (LMCS) can be applied during picture encoding and / or reconstruction.
[0092] The filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 270 (particularly, a DPB of the memory 270). For example, the various filtering methods can include a deblocking filter, a sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filter 260 can generate various information related to filtering and transmit the generated information to the entropy encoder 240 as described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoder 240 and output in the form of a bitstream.
[0093] The modified reconstructed picture transmitted to the memory 270 can be used as a reference picture in the inter-predictor 221. When inter prediction is applied by the encoding apparatus, prediction mismatch between the encoding apparatus 200 and the decoding apparatus 300 can be avoided and encoding efficiency can be improved.
[0094] The DPB of the memory 270 can store the modified reconstructed picture used as a reference picture in the inter-predictor 221. The memory 270 can store motion information of a block for which motion information of a current picture is derived (or encoded) and / or motion information of a block that has been reconstructed in a picture. The stored motion information can be transmitted to the inter-predictor 221 and used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory 270 can store reconstructed samples of a reconstructed block in a current picture and can deliver the reconstructed samples to the intra-predictor 222.
[0095] Figure 3 FIG. 1 is a schematic diagram illustrating a configuration of a video / image encoding apparatus to which embodiments of the disclosure can be applied.
[0096] Referring to Figure 3 , the decoding apparatus 300 can include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, a memory 360. The predictor 330 can include an inter-predictor 332 and an intra-predictor 331. The residual processor 320 can include a dequantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 can be configured by hardware components (e.g., a decoder chipset or a processor). In addition, the memory 360 can include a decoded picture buffer (DPB) or can be configured by a digital storage medium. The hardware components can further include the memory 360 as an internal / external component.
[0097] When a bitstream including video / image information is input, the decoding apparatus 300 can reconstruct an image corresponding to the processing of the video / image information processed in the encoding apparatus. Figure 2 For example, the decoding apparatus 300 can derive a unit / block based on block partitioning related information obtained from the bitstream. The decoding apparatus 300 can perform decoding using a processor applied in the encoding apparatus. Thus, for example, the processor of the decoding can be an encoding unit, and the encoding unit can be partitioned from a coding tree unit or a largest coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the encoding unit. The reconstructed image signal decoded and output by the decoding apparatus 300 can be reproduced by a reproduction apparatus.
[0098] The decoding apparatus 300 can receive a bitstream from Figure 2The signal outputted in the form of a bitstream by the encoding apparatus, and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information can further include general constraint information. The decoding apparatus can also decode a picture based on the information on the parameter sets and / or the general constraint information. The information and / or the syntax elements signaled / received later described in the disclosure can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 310 decodes information in the bitstream based on an encoding method such as exponential Golomb encoding, CAVLC, or CABAC, and outputs syntax elements and quantized values of transform coefficients of a residual required for image reconstruction. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using information of a decoding target syntax element, decoding information of a decoding target block, or a symbol / bin decoded in a previous stage, and generate a symbol corresponding to a value of each syntax element by performing arithmetic decoding on a bin by predicting a probability of bin occurrence according to the determined context model. In this case, the CABAC entropy decoding method can update the context model by using information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Information related to prediction among the information decoded by the entropy decoder 310 can be provided to the predictor (inter-predictor 332 and intra-predictor 331), and the residual values (i.e., quantized transform coefficients and related parameter information) on which entropy decoding is performed in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive a residual signal (a residual block, a residual sample, a residual sample array). In addition, information on filtering among the information decoded by the entropy decoder 310 can be provided to the filter 350. Furthermore, a receiver (not shown) for receiving a signal outputted from the encoding apparatus can also be configured as an internal / external element of the decoding apparatus 300, or the receiver can be a component of the entropy decoder 310. Furthermore, the decoding apparatus according to the disclosure can be referred to as a video / image / picture decoding apparatus, and the decoding apparatus can be classified into an information decoder (a video / image / picture information decoder) and a sample decoder (a video / image / picture sample decoder). The information decoder can include the entropy decoder 310, and the sample decoder can include at least one of the dequantizer 321, the inverse transformer 322, the adder 340, the filter 350, the memory 360, the inter-predictor 332, and the intra-predictor 331.
[0099] The dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on a coefficient scan order performed in the encoding apparatus. The dequantizer 321 can perform dequantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step length information) and obtain the transform coefficients.
[0100] The inverse transformer 322 inverse-transforms the transform coefficients to obtain a residual signal (a residual block, a residual sample array).
[0101] The predictor can perform prediction on the current block and generate a prediction block including predicted samples of the current block. The predictor can determine whether to apply intra prediction or inter prediction to the current block based on information on prediction output from the entropy decoder 310 and can determine a specific intra / inter prediction mode.
[0102] The predictor 330 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra prediction or inter prediction to predict one block, but also simultaneously apply intra prediction and inter prediction. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can predict a block based on an intra block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or the palette mode can be used for content image / video encoding of games and the like, such as screen content coding (SCC). The IBC basically performs prediction in a current picture, but can be performed similarly to inter prediction, such that a reference block is derived in the current picture. That is, the IBC can use at least one inter prediction technique described in the present disclosure. The palette mode can be regarded as an example of intra encoding or intra prediction. When the palette mode is applied, sample values within a picture can be signaled based on information on a palette table and a palette index.
[0103] The intra predictor 331 can predict the current block with reference to samples in the current picture. Depending on the prediction mode, the referred samples can be located in the vicinity of the current block or can be spaced apart. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra predictor 331 can determine a prediction mode applied to the current block using a prediction mode applied to a neighboring block.
[0104] The inter predictor 332 can derive a prediction block of a current block based on a reference block (a reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in units of a block, a sub-block, or a sample based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter predictor 332 can configure a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index of the current block based on received candidate selection information. Inter prediction can be performed based on various prediction modes, and information about the prediction can include information indicating an inter prediction mode of the current block.
[0105] The adder 340 can generate a reconstructed signal (a reconstructed picture, a reconstructed block, a reconstructed sample array) by adding the obtained residual signal to a prediction signal (a prediction block, a prediction sample array) output from the predictor (including the inter predictor 332 and / or the intra predictor 331). If there is no residual for a block to be processed, for example, when a skip mode is applied, the prediction block can be used as the reconstructed block.
[0106] The adder 340 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra prediction of a next block to be processed in the current picture, can be output by filtering as described below, or can be used for inter prediction of a next picture.
[0107] In addition, luma mapping for chroma scaling (LMCS) can be applied in the picture decoding process.
[0108] The filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 360 (specifically, the DPB of the memory 360). For example, the various filtering methods can include deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.
[0109] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction 332. The memory 360 can store motion information of a block for which motion information in the current picture is derived (or decoded) and / or motion information of a block in the picture for which the reconstruction is already done. The stored motion information can be sent to the inter prediction 332 to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory 360 can store reconstructed samples of a reconstructed block in the current picture and transfer the reconstructed samples to the intra prediction 331.
[0110] In this document, the implementations described in the filter 260, the inter prediction 221 and the intra prediction 222 of the encoding device 200 can be applied identically or respectively corresponding to the filter 350, the inter prediction 332 and the intra prediction 331 of the decoding device 300. This can also apply to the unit 332 and the intra prediction 331.
[0111] As described above, in video coding, prediction is performed to increase the coding efficiency. Thereby, a prediction block including prediction samples of a current block (a block to be coded) can be generated. Here, the prediction block includes prediction samples in a spatial domain (or pixel domain). The prediction block is derived identically from the encoding apparatus and the decoding apparatus, and the encoding apparatus decodes information on a residual between the original block and the prediction block (residual information) rather than original sample values of the original block itself. By signaling the apparatus, the image coding efficiency can be increased. The decoding device can derive a residual block including residual samples based on the residual information, and generate a reconstructed block including reconstructed samples by summing the residual block and the prediction block, and generate a reconstructed picture including the reconstructed block.
[0112] The residual information can be generated by a transform process and a quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, and perform a transform process on residual samples (an array of residual samples) included in the residual block to derive transform coefficients, and then, by performing a quantization process on the transform coefficients, derive quantized transform coefficients to signal residual-related information (via a bitstream) to the decoding device. Here, the residual information can include position information, a transform technique, a transform kernel and a quantization parameter, value information of the quantized transform coefficients, etc. The decoding device can perform a dequantization / inverse transform process based on the residual information and derive residual samples (or a residual block). The decoding device can generate a reconstructed picture based on the prediction block and the residual block. The encoding device can also dequantize / inverse transform the quantized transform coefficients for inter prediction reference of a later picture to derive a residual block, and generate a reconstructed picture based thereon.
[0113] In this document, at least one of quantization / dequantization and / or transform / inverse transform can be omitted. When quantization / dequantization is omitted, a quantized transform coefficient can be referred to as a transform coefficient. When transform / inverse transform is omitted, a transform coefficient can be referred to as a coefficient or a residual coefficient, or for consistency of expression, can still be referred to as a transform coefficient.
[0114] In this document, a quantized transform coefficient and a transform coefficient can be referred to as a transform coefficient and a scaled transform coefficient, respectively. In this case, residual information can include information on a transform coefficient, and the information on the transform coefficient can be signaled through residual coding syntax. A transform coefficient can be derived based on the residual information (or the information on the transform coefficient), and a scaled transform coefficient can be derived through inverse transform (scaling) of the transform coefficient. A residual sample can be derived based on inverse transform (transform) of the scaled transform coefficient. This can also be applied / expressed in other parts of this document.
[0115] Intra prediction can refer to prediction that generates a prediction sample of a current block based on reference samples in a picture to which the current block belongs (hereinafter, referred to as a current picture). When intra prediction is applied to the current block, neighboring reference samples to be used for intra prediction of the current block can be derived. The neighboring reference samples of the current block can include a total of 2×nH samples neighboring a left boundary of the current block of size nW×nH and neighboring a lower left, samples neighboring an upper boundary of the current block and a total of 2×nW samples neighboring an upper right, and one sample neighboring a left upper of the current block. Alternatively, the neighboring reference samples of the current block can include a plurality of upper neighboring samples and a plurality of left neighboring samples. In addition, the neighboring reference samples of the current block can include a total of nH samples neighboring a right boundary of the current block of size nW×nH, a total of nW samples neighboring a lower boundary of the current block, and one sample neighboring a lower right of the current block.
[0116] However, some of the neighboring reference samples of the current block can not be decoded or available. In this case, the decoder can configure the neighboring reference samples to be used for prediction by replacing the unavailable samples with available samples. Alternatively, the neighboring reference samples to be used for prediction can be configured through interpolation of the available samples.
[0117] When the neighboring reference samples are derived, (i) a prediction sample can be derived based on an average or interpolation of the neighboring reference samples of the current block, and (ii) a prediction sample can be derived based on a reference sample existing in a certain (prediction) direction of the prediction sample among the peripheral reference samples of the current block. The case of (i) can be referred to as a non-directional mode or a non-angular mode, and the case of (ii) can be referred to as a directional mode or an angular mode.
[0118] In addition, the prediction sample can also be generated by interpolation between a second neighboring sample and a first neighboring sample among the neighboring reference samples based on the prediction sample of the current block being located in a direction opposite to a prediction direction of the intra prediction mode of the current block. The above case can be referred to as linear interpolation intra prediction (LIP). In addition, the chroma prediction sample can be generated based on the luma sample using a linear model. This case can be referred to as LM mode.
[0119] In addition, a temporary prediction sample of the current block can be derived based on the filtered neighboring reference samples, and at least one reference sample (i.e., unfiltered neighboring reference sample) derived according to the intra prediction mode among the existing neighboring reference samples and the temporary prediction sample can be weighted and summed to derive the prediction sample of the current block. The above case can be referred to as position dependent intra prediction (PDPC).
[0120] In addition, a reference sample row with the highest prediction accuracy among the neighboring multiple reference sample rows of the current block can be selected to derive the prediction sample using the reference sample located in the prediction direction on the corresponding row, and then the reference sample row used herein can be indicated (signaled) to the decoding device to perform the intra prediction encoding. The above case can be referred to as multiple reference row (MRL) intra prediction or MRL-based intra prediction.
[0121] In addition, the intra prediction can be performed by dividing the current block into vertical or horizontal sub-partitions based on the same intra prediction mode, and the neighboring reference samples can be derived and used in units of the sub-partitions. That is, in this case, the intra prediction mode of the current block is also applied to the sub-partitions, and in some cases, the intra prediction performance can be improved by deriving and using the neighboring reference samples in units of the sub-partitions. This prediction method can be referred to as intra sub-partition (ISP) or ISP-based intra prediction.
[0122] The above intra prediction methods can be referred to as intra prediction types separately from the intra prediction modes. The intra prediction types can be referred to by various terms such as intra prediction techniques or additional intra prediction modes. For example, the intra prediction types (or additional intra prediction modes) can include at least one of the above LIP, PDPC, MRL, and ISP. A general intra prediction method other than a specific intra prediction type such as LIP, PDPC, MRL, or ISP can be referred to as a normal intra prediction type. The normal intra prediction type can be generally applied when the specific intra prediction type is not applied, and can perform prediction based on the above intra prediction modes. In addition, post-filtering can be performed on the derived prediction sample as needed.
[0123] Specifically, the intra prediction process can include an intra prediction mode / type determination step, a neighboring reference sample derivation step, and a prediction sample derivation step based on the intra prediction mode / type. In addition, a post-filtering step can be performed on the derived prediction sample as needed.
[0124] When intra prediction is applied, the intra prediction mode applied to the current block can be determined using intra prediction modes of neighboring blocks. For example, a decoding device can select one of the intra prediction mode candidates of an mpm list derived based on intra prediction modes of neighboring blocks (e.g., left and / or top neighboring blocks) of the current block based on a received most probable mode (mpm) index, and select one of the other remaining intra prediction modes not included in the mpm candidates (and the planar mode) based on remaining intra prediction mode information. The mpm list can be configured to include or not include the planar mode as a candidate. For example, if the mpm list includes the planar mode as a candidate, the mpm list can have six candidates. If the mpm list does not include the planar mode as a candidate, the mpm list can have three candidates. When the mpm list does not include the planar mode as a candidate, a non-planar flag (e.g., intra luma not planar flag) indicating whether the intra prediction mode of the current block is not the planar mode can be signaled. For example, the mpm flag can be first signaled, when the value of the mpm flag is 1, the mpm index and the non-planar flag can be signaled. In addition, when the value of the non-planar flag is 1, the mpm index can be signaled. Here, the mpm list is configured to not include the planar mode as a candidate, the non-planar flag is not first signaled to first check whether it is the planar mode because the planar mode is always considered as an mpm.
[0125] For example, whether the intra prediction mode applied to the current block is in the mpm candidates (and the planar mode) or in the remaining modes can be indicated based on an mpm flag (e.g., Intra_luma_mpm_flag). A value of 1 of the mpm flag can indicate that the intra prediction mode of the current block is in the mpm candidates (and the planar mode), and a value of 0 of the mpm flag can indicate that the intra prediction mode of the current block is not in the mpm candidates (and the planar mode). A value of 0 of a non-planar flag (e.g., Intra_luma_not_planar_flag) can indicate that the intra prediction mode of the current block is the planar mode, and a value of 1 of the non-planar flag can indicate that the intra prediction mode of the current block is not the planar mode. An mpm index can be signaled in the form of mpm_idx or intra_luma_mpm_idx, and remaining intra prediction mode information can be signaled in the form of rem_intra_luma_pred_mode or intra_luma_mpm_remainder. For example, the remaining intra prediction mode information can index the remaining intra prediction modes among all the intra prediction modes that are not included in the mpm candidates (and the planar mode) in order to indicate one of them in the order of the prediction mode numbers. The intra prediction mode can be an intra prediction mode of a luma component (sample). Hereinafter, the intra prediction mode information can include at least one of the mpm flag (e.g., Intra_luma_mpm_flag), the non-planar flag (e.g., Intra_luma_not_planar_flag), the mpm index (e.g., mpm_idx or intra_luma_mpm_idx), and the remaining intra prediction mode information (rem_intra_luma_pred_mode or intra_luma_mpm_remainder). In this document, the MPM list can be referred to as various terms such as MPM candidate list and candModeList. When MIP is applied to the current block, a separate mpm flag (e.g., intra_mip_mpm_flag) for MIP, an mpm index (e.g., intra_mip_mpm_idx), and remaining intra prediction mode information (e.g., intra_mip_mpm_remainder) can be signaled, and the non-planar flag can not be signaled.
[0126] In other words, generally, when block splitting is performed on an image, a current block to be encoded and a neighboring block have similar image characteristics. Accordingly, there is a high probability that the current block and the neighboring block have the same or similar intra prediction modes. Accordingly, the encoder can use the intra prediction mode of the neighboring block to encode the intra prediction mode of the current block.
[0127] For example, the encoder / decoder can configure a list of most probable modes (MPMs) for the current block. The MPM list can also be referred to as an MPM candidate list. Herein, MPM can refer to a mode that considers the similarity between the current block and neighboring blocks to improve coding efficiency in intra prediction mode coding. As described above, the MPM list can be configured to include the planar mode, or can be configured not to include the planar mode. For example, when the MPM list includes the planar mode, the number of candidates in the MPM list can be 6. And, if the MPM list does not include the planar mode, the number of candidates in the MPM list can be 5.
[0128] The encoder / decoder can configure the MPM list to include 5 or 6 MPMs.
[0129] To configure the MPM list, three types of modes can be considered: a default intra mode, a neighbor intra mode, and a derived intra mode.
[0130] For the neighbor intra mode, two neighboring blocks, i.e., a left neighboring block and an above neighboring block, can be considered.
[0131] As described above, if the MPM list is configured not to include the planar mode, the planar mode is excluded from the list, and the number of MPM list candidates can be set to 5.
[0132] In addition, a non-directional mode (or non-angular mode) among the intra prediction modes can include a DC mode based on an average of neighboring reference samples of the current block or a planar mode based on interpolation.
[0133] When inter prediction is applied, a predictor of an encoding / decoding device can derive prediction samples by performing inter prediction in a unit of a block. Inter prediction can be prediction derived in a manner that depends on data elements (e.g., sample values or motion information) of a picture other than a current picture. When inter prediction is applied to a current block, a prediction block (prediction sample array) of the current block can be derived based on a reference block (reference sample array) specified by a motion vector on a reference picture indicated by a reference picture index. Here, to reduce the amount of motion information transmitted in an inter prediction mode, motion information of a current block can be predicted in a unit of a block, a sub-block, or a sample based on correlation of motion information between neighboring blocks and the current block. Motion information can include a motion vector and a reference picture index. Motion information can also include inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, neighboring blocks can include spatial neighboring blocks existing in a current picture and temporal neighboring blocks existing in a reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring block can be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, a motion information candidate list can be configured based on neighboring blocks of a current block, and flag or index information indicating which candidate is selected (used) to derive a motion vector and / or a reference picture index of the current block can be signaled. Inter prediction can be performed based on various prediction modes. For example, in the case of a skip mode and a merge mode, motion information of a current block can be the same as that of a neighboring block. In the skip mode, unlike the merge mode, a residual signal can not be transmitted. In the case of a motion vector prediction (MVP) mode, a motion vector of a selected neighboring block can be used as a motion vector predictor, and a motion vector of a current block can be signaled. In this case, the motion vector of the current block can be derived using a sum of the motion vector predictor and a motion vector difference.
[0134] Motion information can include L0 motion information and / or LI motion information depending on the inter prediction type (L0 prediction, LI prediction, Bi prediction, etc.). A motion vector in the L0 direction can be referred to as an L0 motion vector or MVL0, and a motion vector in the LI direction can be referred to as an LI motion vector or MVLl. Prediction based on an L0 motion vector can be referred to as L0 prediction, prediction based on an LI motion vector can be referred to as LI prediction, and prediction based on both an L0 motion vector and an LI motion vector can be referred to as bi-prediction. Here, an L0 motion vector can indicate a motion vector associated with a reference picture list L0 (L0), and an LI motion vector can indicate a motion vector associated with a reference picture list LI (LI). A reference picture list L0 can include pictures earlier than the current picture in output order as reference pictures, and a reference picture list LI can include pictures later than the current picture in output order. A previous picture can be referred to as a forward (reference) picture, and a subsequent picture can be referred to as a backward (reference) picture. A reference picture list L0 can also include pictures later than the current picture in output order as reference pictures. In this case, previous pictures can be indexed first in the reference picture list L0, and subsequent pictures can be indexed later. A reference picture list LI can also include previous pictures earlier than the current picture in output order as reference pictures. In this case, subsequent pictures can be indexed first in the reference picture list 1, and previous pictures can be indexed later. The output order can correspond to a picture order count (POC) order.
[0135] Figure 4 An example block tree structure is shown. Figure 4 An example CTU is shown as being partitioned into CUs based on a quadtree and a nested multi-type tree structure.
[0136] Bold block edges represent quadtree partitioning, and remaining edges represent multi-type tree partitioning. Quadtree partitioning with nested multi-type trees can provide a content adaptive coding tree structure. A CU can correspond to a coding block (CB). Alternatively, a CU can include a coding block of luma samples and two coding blocks of corresponding chroma samples. The size of a CU can be as large as a CTU or as small as 4x4 luma sample units, for example. In the case of a 4:2:0 color format (or chroma format), the maximum chroma CB size can be 64x64, and the minimum chroma CB size can be 2x2, for example.
[0137] In this document, the maximum allowed luma TB size can be 64x64, and the maximum allowed chroma TB size can be 32x32, for example. When the width or height of a CB partitioned according to a tree structure is larger than the maximum transform width or the maximum transform height, the CB can be automatically (or implicitly) partitioned until the TB size limit in the horizontal and vertical directions is satisfied.
[0138] In this document, a coding tree scheme can support that luma (component) blocks and chroma (component) blocks have separate block tree structures. Blocks having separate block tree structures can be blocks that have been coded as separate trees. A case where luma blocks and chroma blocks in one CTU have the same block tree structure can be denoted as SINGLE_TREE. A case where luma blocks and chroma blocks in one CTU have separate block tree structures can be denoted as DUAL_TREE. In this case, the block tree type of luma components can be referred to as DUAL_TREE_LUMA, and the block tree type of chroma components can be referred to as DUAL_TREE_CHROMA. Blocks having dual tree structures can be blocks that have been coded as dual trees. For P and B slices / tile groups, luma and chroma CTBs in one CTU can be restricted to have the same coding tree structure. However, for I slices / tile groups, luma blocks and chroma blocks can have separate block tree structures, respectively. When separate block tree mode is applied, luma CTBs can be partitioned into CUs based on a certain coding tree structure, and chroma CTBs can be partitioned into chroma CUs based on another coding tree structure. This can mean that CUs in an I slice / tile group are composed of coding blocks of luma components or coding blocks of two chroma components, and CUs of P or B slices / tile groups are composed of blocks of three color components. In this document, a slice can be referred to as a tile / tile group, and a tile / tile group can be referred to as a slice.
[0139] Although a quadtree coding tree structure having a nested multi-type tree has been described, the structure of partitioning a CU is not limited thereto. For example, the BT structure and the TT structure can be interpreted as including the concept of a multi-partition tree (MPT) structure, and the CU can be interpreted as being partitioned by the QT structure and the MPT structure. In an example of partitioning a CU by the QT structure and the MPT structure, the partitioning structure can be determined by signaling a syntax element including information about how many blocks a leaf node of the QT structure is partitioned into (e.g., MPT_split_type) and a syntax element including information about whether the leaf node of the QT structure is partitioned in a vertical direction or a horizontal direction (e.g., MPT_split_mode).
[0140] Further, in another example, a CU can be partitioned by any other method other than the QT structure, the BT structure, or the TT structure. That is, unlike according to the QT structure, partitioning a CU at a lower depth into 1 / 4 size CUs at a higher depth, or according to the BT structure, partitioning a CU at a lower depth into 1 / 2 size CUs at a higher depth, or according to the TT structure, partitioning a CU at a lower depth into 1 / 4 or 1 / 2 size CUs at a higher depth, a CU at a lower depth can be partitioned into 1 / 5, 1 / 3, 3 / 8, 3 / 5, 2 / 3, or 5 / 8 size CUs at a higher depth, as the case can be, and the CU partitioning method is not limited thereto.
[0141] Figure 5 An exemplary layered structure of an encoded image / video is illustrated.
[0142] Referring to Figure 5 , an encoded image / video is divided into a video coding layer (VCL) which processes an image / video and a decoding process thereof, a subsystem which transmits and stores encoded information, and a NAL (network abstraction layer) which is responsible for a function and exists between the VCL and the subsystem.
[0143] In the VCL, VCL data including compressed image data (slice data) is generated, or a parameter set including a picture parameter set (PSP), a sequence parameter set (SPS), and a video parameter set (VPS) or supplemental enhancement information (SEI) messages additionally required for an image decoding process can be generated.
[0144] In the NAL, a NAL unit can be generated by adding a header information (NAL unit header) to a raw byte sequence payload (RBSP) generated in the VCL. In this case, the RBSP refers to slice data, a parameter set, an SEI message, etc. generated in the VCL. The NAL unit header can include NAL unit type information designated according to RBSP data included in the corresponding NAL unit.
[0145] As illustrated in the drawing, the NAL unit can be classified into a VCL NAL unit and a non-VCL NAL unit according to the RBSP generated in the VCL. The VCL NAL unit can mean a NAL unit including information about an image (slice data), and the non-VCL NAL unit can mean a NAL unit including information required to decode an image (a parameter set or an SEI message).
[0146] The above-described VCL NAL unit and non-VCL NAL unit can be transmitted through a network by attaching header information according to a data standard of the subsystem. For example, the NAL unit can be transformed into a data format of a predetermined standard such as an H.266 / VVC file format, a real-time transport protocol (RTP), a transport stream (TS), etc. and transmitted through various networks.
[0147] As described above, the NAL unit can be designated in a NAL unit type according to an RBSP data structure included in the corresponding NAL unit, and information about the NAL unit type can be stored in the NAL unit header and signaled.
[0148] For example, NAL units can be classified into VCL NAL unit types and non-VCL NAL unit types according to whether the NAL units include information (slice data) about a picture. VCL NAL unit types can be classified according to the nature and type of a picture included in a VCL NAL unit, and non-VCL NAL unit types can be classified according to the type of parameter sets.
[0149] The following are examples of NAL unit types specified according to the type of parameter sets included in non-VCL NAL unit types.
[0150] - APS (Adaptive Parameter Set) NAL unit: type of NAL unit including an APS
[0151] - DPS (Decoding Parameter Set) NAL unit: type of NAL unit including a DPS
[0152] - VPS (Video Parameter Set) NAL unit: type of NAL unit including a VPS
[0153] - SPS (Sequence Parameter Set) NAL unit: type of NAL unit including an SPS
[0154] - PPS (Picture Parameter Set) NAL unit: type of NAL unit including a PPS
[0155] - PH (Picture Header) NAL unit: type of NAL unit including a PH
[0156] The above-described NAL unit types can have syntax information of the NAL unit type, and the syntax information can be stored in a NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified by a nal_unit_type value.
[0157] Further, as described above, one picture can include a plurality of slices, and one slice can include a slice header and slice data. In this case, one picture header can be further added to the plurality of slices (slice header and slice data sets) in one picture. The picture header (picture header syntax) can include information / parameters generally applicable to a picture. In this document, a slice can be mixed with or replaced by a tile group. Also, in this document, a slice header can be mixed with or replaced by a tile group header.
[0158] A slice header (slice header syntax) can include information / parameters that can be generally applicable to a slice. An APS (APS syntax) or a PPS (PPS syntax) can include information / parameters that can be generally applicable to one or more slices or pictures. An SPS (SPS syntax) can include information / parameters that can be generally applicable to one or more sequences. A VPS (VPS syntax) can include information / parameters that can be generally applicable to a plurality of layers. A DPS (DPS syntax) can include information / parameters that can be generally applicable to a total video. The DPS can include information / parameters related to a concatenation of coded video sequences (CVSs). High-level syntax (HLS) in this document can include at least one of the APS syntax, the PPS syntax, the SPS syntax, the VPS syntax, the DPS syntax, and the slice header syntax.
[0159] In this document, an image / image information encoded from an encoding apparatus and signaled to a decoding apparatus in the form of a bitstream includes not only partitioning-related information in a picture, intra / inter prediction information, residual information, loop filtering information, etc., but also information included in a slice header, information included in an APS, information included in a PPS, information included in an SPS, and / or information included in a VPS.
[0160] Further, in order to compensate for a difference between an original image and a reconstructed image due to an error occurring in a compression encoding process such as quantization, a loop filtering process can be performed on a reconstructed sample or a reconstructed picture as described above. As described above, the loop filtering can be performed by a filter of an encoding apparatus and a filter of a decoding apparatus, and a deblocking filter, SAO, and / or an adaptive loop filter (ALF) can be applied. For example, the ALF process can be performed after a deblocking filtering process and / or an SAO process are completed. However, even in this case, the deblocking filtering process and / or the SAO process can be omitted.
[0161] Further, in order to increase encoding efficiency, luminance mapping and chrominance scaling (LMCS) can be applied as described above. The LMCS can be referred to as a loop shaper (shaping). In order to increase encoding efficiency, LMCS control and / or signaling of LMCS-related information can be performed hierarchically.
[0162] Figure 6 A hierarchical structure of a CVS according to an embodiment of this document is exemplarily shown. A coded video sequence (CVS) can include an SPS, a PPS, a tile group header, tile data, and / or a CTU. Here, the tile group header and the tile data can be referred to as a slice header and slice data, respectively.
[0163] The SPS can include a local flag to enable a tool to be used in the CVS. Also, the SPS can be referred to by a PPS including information on parameters changed for each picture. Each coded picture can include one or more coded rectangular tile patches. The tile patches can be grouped into a raster scan forming a tile group. Each tile group is encapsulated with header information called a tile group header. Each tile is composed of a CTU including coded data. Here, the data can include original sample values, predicted sample values, and their luma and chroma components (luma predicted sample values and chroma predicted sample values).
[0164] According to the existing method, ALF data (ALF parameters) or LMCS data (LMCS parameters) are incorporated into the tile group header. Since one video is composed of a plurality of pictures and one picture includes a plurality of tiles, frequent signaling of the ALF data (ALF parameters) or the LMCS data (LMCS parameters) in units of tile groups causes a problem of a decrease in coding efficiency.
[0165] According to the embodiments proposed in the present document, the ALF parameters or the LMCS data (LMCS parameters) can be incorporated into the APS and signaled.
[0166] In the embodiments, the APS can be defined, and the APS can carry necessary ALF data (ALF parameters). Further, the APS can have a self-identification parameter and ALF data. The self-identification parameter of the APS can include an APS ID. That is, the APS can include information indicating the APS ID in addition to the ALF data field. The tile group header or the slice header can refer to the APS using APS index information. In other words, the tile group header or the slice header can include the APS index information, and can perform the ALF process of a target block based on the ALF data (ALF parameters) included in the APS having the APS ID indicated by the APS index information. Here, the APS index information can be referred to as APS ID information.
[0167] Also, the SPS can include a flag allowing the use of the ALF. For example, when the CVS starts, the SPS can be checked, and the flag in the SPS can be checked. For example, the SPS can include the syntax of Table 1 below. The syntax in Table 1 can be a part of the SPS.
[0168] [Table 1]
[0169]
[0170] For example, the semantics of the syntax elements included in the syntax of Table 1 can be expressed as shown in the following table.
[0171] [Table 2]
[0172]
[0173] That is, the sps_alf_enabled_flag syntax element can indicate whether ALF is enabled based on whether its value is 0 or 1. The sps_alf_enabled_flag syntax element can be referred to as an ALF enabled flag (this can be referred to as a first ALF enabled flag), and can be included in the SPS. That is, the ALF enabled flag can be signaled in the SPS (or SPS level). When the value of the ALF enabled flag signaled in the SPS is 1, ALF can be determined to be substantially enabled for pictures referring to the SPS in the CVS. Also, as described above, ALF can be turned on / off separately by signaling an additional enabled flag at a lower level than the SPS.
[0174] For example, if the ALF tool is enabled for the CVS, an additional enabled flag (this can be referred to as a second ALF enabled flag) can be signaled in a tile group header or a slice header. For example, when ALF is enabled at the SPS level, the second ALF enabled flag can be parsed / signaled. When the value of the second ALF enabled flag is 1, ALF data can be parsed through a tile group header or a slice header. For example, the second ALF enabled flag can specify ALF enabled conditions for a luma component and a chroma component. The ALF data can be accessed through APS ID information.
[0175] [Table 3]
[0176]
[0177] [Table 4]
[0178]
[0179] For example, the semantics of the syntax elements included in the above Table 3 or Table 4 can be represented as shown in the following table.
[0180] [Table 5]
[0181]
[0182] [Table 6]
[0183]
[0184] The second ALF enabled flag can include a tile_group_alf_enabled_flag syntax element or a slice_alf_enabled_flag syntax element.
[0185] An APS corresponding to a tile group or a slice reference can be identified based on APS ID information (e.g., tile_group_aps_id syntax element or slice_aps_id syntax element). The APS can include ALF data.
[0186] Further, for example, a structure of an APS including ALF data can be described based on the following syntax and semantics. The syntax of Table 7 can be a part of an APS.
[0187] [Table 7]
[0188]
[0189] [Table 8]
[0190]
[0191] As described above, the adaptation_parameter_set_id syntax element can represent an identifier of a corresponding APS. That is, the APS can be identified based on the adaptation_parameter_set_id syntax element. The adaptation_parameter_set_id syntax element can be referred to as APS ID information. In addition, the APS can include an ALF data field. The ALF data field can be parsed / signalized after the adaptation_parameter_set_id syntax element.
[0192] In addition, for example, an APS extension flag (e.g., aps_extension_flag syntax element) can be parsed / signalized in the APS. The APS extension flag can indicate whether an APS extension data flag (aps_extension_data_flag) syntax element is present. For example, the APS extension flag can be used to provide an extension point for later versions of the VVC standard.
[0193] Figure 7 An exemplary LMCS structure according to an embodiment of the present document is shown. Figure 7 The LMCS structure 700 comprises a loop mapping section 710 for the luminance component based on an adaptive piecewise linear (adaptive PWL) model and a luma-dependent chroma residual scaling section 720 for the chroma components. The dequantization and inverse transform 711, reconstruction 712 and intra prediction 713 blocks of the loop mapping section 710 represent processing applied in the mapped (reshaped) domain. The in-loop filter 715, the motion compensation or inter prediction 717 block of the loop mapping section 710 and the reconstruction 722, intra prediction 723, motion compensation or inter prediction 724, in-loop filter 725 blocks of the chroma residual scaling section 720 represent processing applied in the original (non-mapped, non-reshaped) domain.
[0194] As Figure 7 indicated, when LMCS is enabled, at least one of inverse mapping (reshaping) processing 714, forward mapping (reshaping) processing 718, and chroma scaling processing 721 can be applied. For example, inverse mapping processing can be applied to (reconstructed) luma samples (or luma samples or luma sample array) in the reconstructed picture. The inverse mapping processing can be performed based on a piecewise function (inverse) index of the luma samples. The piecewise function (inverse) index can identify a piece to which the luma sample belongs. The output of the inverse mapping processing is modified (reconstructed) luma samples (or modified luma samples or modified luma sample array). LMCS can be enabled or disabled at tile group (or slice), picture, or higher level.
[0195] Forward mapping processing and / or chroma scaling processing can be applied to generate the reconstructed picture. The picture can include luma samples and chroma samples. The reconstructed picture in luma samples can be referred to as a reconstructed luma picture, and the reconstructed picture in chroma samples can be referred to as a reconstructed chroma picture. The combination of the reconstructed luma picture and the reconstructed chroma picture can be referred to as the reconstructed picture. The reconstructed luma picture can be generated based on the forward mapping processing. For example, if inter prediction is applied to the current block, forward mapping is applied to luma prediction samples derived based on (reconstructed) luma samples in the reference picture. Since the (reconstructed) luma samples in the reference picture are generated based on the inverse mapping processing, forward mapping can be applied to the luma prediction samples, and thus mapped (reshaped) luma prediction samples can be derived. The forward mapping processing can be performed based on a piecewise function index of the luma prediction samples. The piecewise function index can be derived based on the value of the luma prediction samples or the value of luma samples in the reference picture used for inter prediction. If intra prediction (or intra block copy (IBC)) is applied to the current block, no forward mapping is needed because no inverse mapping processing has been applied to the reconstructed samples in the current picture yet. The (reconstructed) luma samples in the reconstructed luma picture are generated based on the mapped luma prediction samples and corresponding luma residual samples.
[0196] The reconstructed chroma picture can be generated based on the chroma scaling processing. For example, (reconstructed) chroma samples in the reconstructed chroma picture can be derived based on chroma residual samples (c res ) in the current block and chroma prediction samples. The chroma residual samples (c resScale ) are derived based on (scaled) chroma residual samples (c res ) of the current block and a chroma residual scaling factor (cScaleInv pred ). The chroma residual scaling factor can be calculated based on the reshaped luma prediction sample value of the current block. For example, the average luma value ave(Y pred) to calculate the scaling factor. For reference, the (scaled) chroma residual samples derived based on the inverse transform / dequantization can be referred to as c resScale , the chroma residual samples derived by performing the (inverse) scaling process on the (scaled) chroma residual samples can be referred to as c res .
[0197] Figure 8 An LMCS structure according to another embodiment of the present document is shown. Figure 8 is described with reference to Figure 7 . Here, the difference between the LMCS structure of Figure 8 and the LMCS structure 700 of Figure 7 is mainly described. Figure 8 The loop mapping section and the luma-dependent chroma residual scaling section of Figure 7 may operate identically (similarly) to the loop mapping section 810 and the luma-dependent chroma residual scaling section 720.
[0198] With reference to Figure 8 , the chroma residual scaling factor can be derived based on the luma reconstructed samples. In this case, the average luma value (avgYr) can be obtained (derived) based on the neighboring luma reconstructed samples outside the reconstructed block, rather than the internal luma reconstructed samples of the reconstructed block, and the chroma residual scaling factor can be derived based on the average luma value (avgYr). Here, the neighboring luma reconstructed samples can be the neighboring luma reconstructed samples of the current block, or can be the neighboring luma reconstructed samples of a virtual pipeline data unit (VPDU) including the current block. For example, when intra prediction is applied to the target block, the reconstructed samples can be derived from the prediction samples derived based on the intra prediction. In another example, when inter prediction is applied to the target block, forward mapping is applied to the prediction samples derived based on the inter prediction, and the reconstructed samples are generated (derived) based on the shaped (or forward mapped) luma prediction samples.
[0199] The video / image information signaled through the bitstream can include the LMCS parameters (information about the LMCS). The LMCS parameters can be configured as high-level syntax (HLS, including slice header syntax), etc. The detailed description and configuration of the LMCS parameters will be described later. As described above, the syntax table described in the present document (and the following embodiments) can be configured / encoded at the encoder end, and signaled to the decoder end through the bitstream. The decoder can parse / decode the information about the LMCS in the syntax table (in the form of a syntax component). One or more embodiments to be described below can be combined. The encoder can encode the current picture based on the information about the LMCS, and the decoder can decode the current picture based on the information about the LMCS.
[0200] Loop mapping of the luma component can adjust the dynamic range of the input signal by redistributing the codewords across the dynamic range to improve compression efficiency. For luma mapping, a forward mapping (reshaping) function (FwdMap) can be used along with an inverse mapping (reshaping) function (InvMap) corresponding to the forward mapping function (FwdMap). The FwdMap function can be signaled using a piecewise linear model, for example, the piecewise linear model can have 16 segments or bins, the segments can have equal length. In one example, the InvMap function does not need to be signaled, but is derived from the FwdMap function. That is, the inverse mapping can be the symmetric function of the forward mapping function. For example, the inverse mapping function can be mathematically constructed as the symmetric function of the forward mapping, as reflected by the line y = x.
[0201] Loop (luma) reshaping can be used to map input luma values (samples) to altered values in the reshaped domain. The reshaped values can be encoded and then mapped back to the original (unmapped, unreshaped) domain after reconstruction. To compensate for the interaction between the luma signal and the chroma signal, chroma residual scaling can be applied. Loop reshaping is done by specifying a high-level syntax of a reshaper model. The reshaper model syntax can signal a piecewise linear model (PWL model). For example, the reshaper model syntax can signal a PWL model with 16 bins or segments of equal length. A forward lookup table (FwdLUT) and / or an inverse lookup table (InvLUT) can be derived based on the piecewise linear model. For example, the PWL model precomputes 1024 forward (FwdLUT) and inverse (InvLUT) lookup tables (LUTs). As an example, when deriving the forward lookup table FwdLUT, the inverse lookup table InvLUT can be derived based on the forward lookup table FwdLUT. The forward lookup table FwdLUT can map an input luma value Yi to an altered value Yr, and the inverse lookup table InvLUT can map the altered value Yr to a reconstructed value Y'i. The reconstructed value Y'i can be derived based on the input luma value Yi.
[0202] In one example, the SPS can include the syntax of Table 9 below. The syntax of Table 9 can include sps_reshaper_enabled_flag as a tool enabling flag. Here, the sps_reshaper_enabled_flag can be used to specify whether a reshaper is used in the coded video sequence (CVS). That is, the sps_reshaper_enabled_flag can be a flag to enable reshaping in the SPS. In one example, the syntax of Table 9 can be part of the SPS.
[0203] [Table 9]
[0204]
[0205] In one example, semantics for the syntax elements sps_seq_parameter_set_id and sps_reshaper_enabled_flag can be as shown in Table 10 below.
[0206] [Table 10]
[0207]
[0208] In one example, a tile group header or slice header can include the syntax of Table 11 or Table 12 below.
[0209] [Table 11]
[0210]
[0211] [Table 12]
[0212]
[0213] Semantics for the syntax elements included in the syntax of Table 11 or Table 12 can include, for example, the items disclosed in the following tables.
[0214] [Table 13]
[0215]
[0216] [Table 14]
[0217]
[0218] As an example, once the flag enabling reshaping in SPS (i.e., sps_reshaper_enabled_flag) is parsed, the tile group header can parse additional data (i.e., the information included in Table 13 or Table 14) for constructing the lookup tables (FwdLUT and / or InvLUT). To this end, the status of the SPS reshaper flag (sps_reshaper_enabled_flag) can be first checked in the slice header or tile group header. When sps_reshaper_enabled_flag is true (or 1), additional flag, i.e., tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag), can be parsed. The purpose of tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag) can be to indicate the presence of the reshaping model. For example, when tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag) is true (or 1), it can indicate the presence of the reshaper for the current tile group (or current slice). When tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag) is false (or 0), it can indicate the absence of the reshaper for the current tile group (or current slice).
[0219] If a reshaper is present and enabled in the current tile group (or current slice), the reshaper model (i.e., tile group reshaper model() or slice reshaper model()) can be processed. Still further, an additional flag tile group reshaper enable flag (or slice reshaper enable flag) can also be parsed. The tile group reshaper enable flag (or slice reshaper enable flag) can indicate whether the reshaping model is used for the current tile group (or slice). For example, if the tile group reshaper enable flag (or slice reshaper enable flag) is 0 (or false), it can indicate that the reshaping model is not used for the current tile group (or current slice). If the tile group reshaper enable flag (or slice reshaper enable flag) is 1 (or true), it can indicate that the reshaping model is used for the current tile group (or slice).
[0220] As one example, the tile group reshaper model present flag (or slice reshaper model present flag) can be true (or 1) and the tile group reshaper enable flag (or slice reshaper enable flag) can be false (or 0). This means that the reshaping model is present but not used in the current tile group (or slice). In this case, the reshaping model can be used in future tile groups (or slices). As another example, the tile group reshaper enable flag can be true (or 1) and the tile group reshaper model present flag can be false (or 0). In this case, the decoder uses the reshaper from a previous initialization.
[0221] When parsing the reshaping model (i.e., tile group reshaper model() or slice reshaper model()) and tile group reshaper enable flag (or slice reshaper enable flag), it can be determined (evaluated) whether the conditions required for chroma scaling exist. The above conditions include condition 1 (the current tile group / slice has not been intra coded) and / or condition 2 (the current tile group / slice has not been split into two separate coding quad-tree structures for luma and chroma, i.e., the block structure of the current tile group / slice is not a dual tree structure). If condition 1 and / or condition 2 is true and / or tile group reshaper enable flag (or slice reshaper enable flag) is true (or 1), tile group reshaper chroma residual scale flag (or slice reshaper chroma residual scale flag) can be parsed. When tile group reshaper chroma residual scale flag (or slice reshaper chroma residual scale flag) is enabled (if 1 or true), it can indicate that chroma residual scaling is enabled for the current tile group (or slice). When tile group reshaper chroma residual scale flag (or slice reshaper chroma residual scale flag) is disabled (if 0 or false), it can indicate that chroma residual scaling is disabled for the current tile group (or slice).
[0222] The purpose of the tile group reshaping model is to parse the data required to construct the look-up tables (LUTs). The idea behind the construction of these LUTs is that the distribution of the allowed range of luma values can be divided into a number of bins (e.g., 16 bins), which can be represented using a set of 16 PWL equations. Thus, any luma value that falls within a given bin can be mapped to a modified luma value.
[0223] Figure 9 A graph representing an example forward mapping is shown. In Figure 9 Five bins are shown exemplarily.
[0224] Referring to Figure 9, the x-axis represents input luma values, and the y-axis represents altered output luma values. The x-axis is divided into 5 bins or slices, each bin having a length of L. That is, the five bins mapped to altered luma values have the same length. A forward lookup table (FwdLUT) can be constructed using data available from the tile group header (i.e., reshaper data), and thus can facilitate mapping.
[0225] In one embodiment, output pivot points associated with bin indices can be calculated. The output pivot points can set (mark) the minimum and maximum boundaries of the output range of luma codeword reshaping. The calculation process of the output pivot points can be performed by calculating the piecewise cumulative distribution function (CDF) of the number of codewords. The output pivot range can be sliced based on the maximum number of bins to be used and the size of the lookup table (FwdLUT or InvLUT). As one example, the output pivot range can be sliced based on the product between the maximum number of bins and the size of the lookup table (size of LUT * maximum number of bin indices). For example, if the product between the maximum number of bins and the size of the lookup table is 1024, the output pivot range can be sliced into 1024 entries. This sawtooth of the output pivot range can be performed (applied or implemented) based on (using) a scaling factor. In one example, the scaling factor can be derived based on the following Equation 1.
[0226] [Equation 1]
[0227] SF = (y2 - y1) * (1 « FP_PREC) + c
[0228] In Equation 1, SF denotes the scaling factor, y1 and y2 denote output pivot points corresponding to respective bins. In addition, FP_PREC and c can be predetermined constants. The scaling factor determined based on Equation 1 can be referred to as a scaling factor for forward reshaping.
[0229] In another embodiment, with respect to inverse reshaping (inverse mapping), for the defined range of bins to be used (i.e., from reshaper_model_min_bin_idx to reshaper_model_max_bin_idx), the input reshaping pivot points corresponding to the mapping pivot points of the forward LUT and the inverse output pivot points (given by the considered bin index * initial number of codewords) are obtained. In another example, the scaling factor SF can be derived based on the following Equation 2.
[0230] [Equation 2]
[0231] SF = (y2 - y1) * (1 « FP_PREC) / (x2 - x1)
[0232] In Equation 2, SF denotes a scaling factor, x1 and x2 denote input pivot points, and y1 and y2 denote output pivot points (inverse-mapped output pivot points) corresponding to respective bins. Here, the input pivot points can be pivot points mapped based on a forward lookup table (FwdLUT), and the output pivot points can be inverse-mapped pivot points based on an inverse lookup table (InvLUT). Also, FP_PREC can be a predetermined constant value. FP_PREC of Equation 2 can be the same as or different from FP_PREC of Equation 1. The scaling factor determined based on Equation 2 can be referred to as a scaling factor for inverse reshaping. During inverse reshaping, binning of the input pivot points can be performed based on the scaling factor of Equation 2. The scaling factor SF is used to slice the range of the input pivot points. Based on the binned input pivot points, bin indices in the range from 0 to a minimum bin index (reshaper_model_min_bin_idx) and / or from a minimum bin index (reshaper_model_min_bin_idx) to a maximum bin index (reshape_model_max_bin_idx) are assigned pivot values corresponding to a minimum bin value and a maximum bin value.
[0233] In one example, LMCS data (lmcs_data) can be included in the APS. For example, semantics of the APS can be to signal 32 APSs for encoding.
[0234] The following table shows syntax and semantics of an exemplary APS according to embodiments of this document.
[0235] [Table 15]
[0236]
[0237] [Table 16]
[0238]
[0239] Referring to Table 15, type information (e.g., aps_params_type) of APS parameters can be parsed / signaled in the APS. The type information of the APS parameters can be parsed / signaled after adaptation_parameter_set_id.
[0240] The APS_params_type, ALF_APS, and LMCS_APS included in the above Table 15 can be described according to Table 3.2 included in Table 16. That is, according to the APS_params_type included in the above Table 15, the type of the APS parameter applied to the APS can be set as shown in Table 3.2 included in Table 16. The syntax elements included in Table 15 can be described with reference to Table 8. The description related to the APS can be supported by the description provided above together with Tables 1 to 8.
[0241] Referring to Table 16, for example, aps_params_type can be a syntax element for classifying the type of a corresponding APS parameter. The type of the APS parameter can include an ALF parameter and an LMCS parameter. Referring to Table 16, when the value of the type information (aps_params_type) is 0, the name of aps_params_type can be determined as ALF_APS (or ALF APS), and the type of the APS parameter can be determined as the ALF parameter (the APS parameter can mean the ALF parameter). In this case, the ALF data field (i.e., alf_data()) can be parsed / signalized to the APS. When the value of the type information (aps_params_type) is 1, the name of aps_params_type can be determined as LMCS_APS (or LMCS APS), and the type of the APS parameter can be determined as the LMCS parameter (the APS parameter can mean the LMCS parameter). In this case, the LMCS data field (i.e., lmcs_data()) can be parsed / signalized to the APS.
[0242] The following Table 17 and / or Table 18 show syntax of a reshaper model according to an embodiment. The reshaper model can be referred to as an LMCS model. Although the reshaper model is exemplarily described as a tile group reshaper here, the present specification is not necessarily limited to this embodiment. For example, the reshaper model can be included in the APS, or the tile group reshaper model can be referred to as a slice reshaper model or LMCS data (LMCS data field). In addition, the prefix "reshaper_model" or "Rsp" can be used interchangeably with "lmcs". For example, in the following tables and the following description, reshaper_model_min_bin_idx, reshaper_model_delta_max_bin_idx, reshaper_model_max_bin_idx, RspCW, and RsepDeltaCW can be used interchangeably with lmcs_min_bin_idx, lmcs_delta_cs_bin_idx, lmcs_max_bin_idx, lmcs_delta_csDcs_bin_idx, and Wlm_idx, respectively.
[0243] The LMCS data (lmcs_data()) or the shaper model (tile group shaper or slice shaper) included in Table 15 above can be expressed as the syntax included in the following table.
[0244] [Table 17]
[0245]
[0246] [Table 18]
[0247]
[0248] The semantics of the syntax elements included in the syntax of Table 17 and / or Table 18 can include, for example, the items disclosed in the following table.
[0249] [Table 19]
[0250]
[0251]
[0252]
[0253] [Table 20]
[0254]
[0255]
[0256] The inverse mapping process of the luma sample according to the present document can be described in the form of a standard document as shown in the following table.
[0257] [Table 21]
[0258]
[0259] The identification of the piecewise function index process of the luma sample according to the present document can be described in the form of a standard document as shown in the following table. In Table 22, idxYInv can be referred to as an inverse mapping index, and the inverse mapping index can be derived based on the reconstructed luma sample (lumaSample).
[0260] [Table 22]
[0261]
[0262]
[0263] The luma mapping can be performed based on the above-described embodiments and examples, and the above-described syntax and components included therein can be merely exemplary representations, and the embodiments in this document are not limited to the above-described tables or formulas. Hereinafter, a method of performing chroma residual scaling (scaling of the chroma component of a residual sample) based on luma mapping is described.
[0264] The (luma-dependent) chroma residual scaling is designed to compensate for the interaction between the luma signal and its corresponding chroma signal. For example, it is also signaled at the tile group level whether chroma residual scaling is enabled or not. In one example, if luma mapping is enabled and if no dual tree partitioning (also referred to as separate chroma tree) is applied for the current tile group, an additional flag is signaled to indicate whether luma-dependent chroma residual scaling is enabled or not. In other examples, when no luma mapping is used, or when dual tree partitioning is used in the current tile group, luma-dependent chroma residual scaling is disabled. In another example, for chroma blocks with an area smaller than or equal to 4, luma-dependent chroma residual scaling is always disabled.
[0265] The chroma residual scaling can be based on the average value of the corresponding luma prediction block (luma component of the prediction block to which the intra prediction mode and / or the inter prediction mode is applied). The scaling operation at the encoder side and / or the decoder side can be implemented in fixed-point integer arithmetic based on the following formula 3.
[0266] [Formula 3]
[0267] c’ = sign(c) * ((abs(c) * s + 2CSCALE_FP_PREC - 1) » CSCALE_FP_PREC)
[0268] In formula 3, c’ denotes the scaled chroma residual sample (scaled chroma component of the residual sample), c denotes the chroma residual sample (chroma component of the residual sample), s denotes the chroma residual scaling factor, and CSCALE_FP_PREC denotes a (predefined) constant value to specify the precision. For example, CSCALE_FP_PREC can be 11.
[0269] Figure 10 is a flowchart illustrating a method for deriving a chroma residual scaling index according to an embodiment of the present document. Figure 10 The method in can be based on Figure 7 and the tables, formulas, variables, arrays, and functions included in the description related to Figure 7 may be performed.
[0270] In step S1010, it can be determined, based on the prediction mode information, whether the prediction mode of the current block is the intra prediction mode or the inter prediction mode. If the prediction mode is the intra prediction mode, the current block or the prediction samples of the current block are considered to have been in the reshaped (mapped) region. If the prediction mode is the inter prediction mode, the current block or the prediction samples of the current block are considered to be in the original (unmapped, unreshaped) region.
[0271] In step S1020, when the prediction mode is the intra prediction mode, the average of the current block (or the luma prediction samples of the current block) can be calculated (derived). That is, the average of the current block in the already reshaped region is directly calculated. The average can also be referred to as the mean or the average value.
[0272] In step S1021, when the prediction mode is the inter prediction mode, forward reshaping (forward mapping) can be performed (applied) on the luma prediction samples of the current block. With the forward reshaping, the luma prediction samples based on the inter prediction mode can be mapped from the original region to the reshaped region. In one example, the forward reshaping of the luma prediction samples can be performed based on the reshaping model described above in Table 17 and / or Table 18.
[0273] In step S1022, the average of the forward reshaped (forward mapped) luma prediction samples can be calculated (derived). That is, the average processing of the forward reshaping result can be performed.
[0274] In step S1030, the chroma residual scaling index can be calculated. When the prediction mode is the intra prediction mode, the chroma residual scaling index can be calculated based on the average of the luma prediction samples. When the prediction mode is the inter prediction mode, the chroma residual scaling index can be calculated based on the average of the forward reshaped luma prediction samples.
[0275] In an embodiment, the chroma residual scaling index can be calculated based on a for loop syntax. The following table shows an example for loop syntax for deriving (calculating) the chroma residual scaling index.
[0276] [Table 23]
[0277]
[0278] In Table 23, idxS denotes the chroma residual scaling index, idxFound denotes an index identifying whether the chroma residual scaling index satisfying the condition of the if statement is obtained, S denotes a predetermined constant value, and MaxBinIdx denotes the maximum allowed bin index. ReshapPivot[idxS+1] (in other words, LmcsPivot[idxS+1]) can be derived based on Table 19 and / or Table 20 described above.
[0279] In an embodiment, a chroma residual scaling factor can be derived based on a chroma residual scaling index. Equation 4 is an example for deriving the chroma residual scaling factor.
[0280] [Equation 4]
[0281] s = ChromaScaleCoef [ idxS ]
[0282] In Equation 4, s represents the chroma residual scaling factor, and ChromaScaleCoef can be a variable (or an array) derived based on Table 19 and / or Table 20 described above.
[0283] As described above, an average luma value of reference samples can be obtained, and a chroma residual scaling factor can be derived based on the average luma value. As described above, chroma component residual samples can be scaled based on the chroma residual scaling factor, and chroma component reconstructed samples can be generated based on the scaled chroma component residual samples.
[0284] In one embodiment of the present document, a signaling structure for efficiently applying the above-described LMCS is proposed. According to this embodiment of the present document, for example, LMCS data can be included in HLS (i.e., APS), and by lower level header information (i.e., picture header, slice header) of the APS, a LMCS model (reshaper model) can be adaptively derived by signaling an ID of the APS (referred to as header information). The LMCS model can be derived based on LMCS parameters. In addition, for example, a plurality of APS IDs can be signaled by the header information, and thereby, different LMCS models can be applied in a block unit within the same picture / slice.
[0285] In one embodiment according to the present document, a method of efficiently performing operations required for LMCS is proposed. According to the semantics described above in Table 19 and / or Table 20, division operation by a segment length lmcsCW[i] (also referred to as RspCW[i] in the present document) is required to derive InvScaleCoeff[i]. The inversely mapped segment length can not be a power of 2, which means that the division operation cannot be performed by bit shift.
[0286] For example, calculating the InvScaleCoeff can require up to 16 divisions per slice. According to Table 19 and / or Table 20 described above, for 10-bit coding, the range of lmcsCW[i] is from 8 to 511, and thus in order to implement the division operation by lmcsCW[i] using a LUT, the size of the LUT must be 504. Also, for 12-bit coding, the range of lmcsCW[i] is from 32 to 2047, and thus the LUT size needs to be 2016 to implement the division operation by lmcsCW[i] using a LUT. That is, division is expensive in terms of hardware implementation, and thus it is desirable to avoid division as much as possible.
[0287] In one aspect of the present embodiment, lmcsCW[i] can be constrained to be a multiple of a fixed number (or a predetermined number). Thus, the lookup table (LUT) for division (the capacity or size of the LUT) can be reduced. For example, if lmcsCW[i] becomes a multiple of 2, the size of the LUT for replacing the division process can be reduced by half.
[0288] In another aspect of the present embodiment, it is proposed that, for coding with higher internal bit depth coding, on top of the existing constraint that "the value of lmcsCW[i] should be in the range of (OrgCW » 3) to (OrgCW « 3 - 1)", if the coding bit depth is higher than 10, lmcsCW[i] is further constrained to be a multiple of 1 « (BitDepthY - 10). Here, BitDepthY can be the luma bit depth. Thus, the possible number of lmcsCW[i] does not change with the coding bit depth, and the size of the LUT required for calculating InvScaleCoeff does not increase due to higher coding bit depth. For example, for 12-bit internal coding bit depth, the limit value of lmcsCW[i] is a multiple of 4, and thus the LUT for replacing the division process will be the same as for 10-bit coding. This aspect can be implemented alone, but can also be implemented in combination with the above aspect.
[0289] In another aspect of the present embodiment, lmcsCW[i] can be constrained to a narrower range. For example, lmcsCW[i] can be constrained to be in the range of (OrgCW » 1) to (OrgCW « 1) - 1. Then for 10-bit coding, the range of lmcsCW[i] can be [32, 127], and thus only a LUT of size 96 is required to calculate InvScaleCoeff.
[0290] In another aspect of the present embodiment, lmcsCW[i] can be approximated to the number closest to a power of 2, and used in the shaper design. Thus, the division in the inverse mapping can be performed (and replaced) by bit shifting.
[0291] In one embodiment according to the present document, a constraint of the LMCS codeword range is proposed. According to Table 8 above, the value of the LMCS codeword is in the range from (OrgCW » 3) to (OrgCW « 3) - 1. This codeword range is too wide. It can cause visual artifact problems when there is a large difference between RspCW[i] and OrgCW.
[0292] According to one embodiment of the present document, it is proposed to constrain the codeword of the LMCS PWL mapping to a narrow range. For example, the range of lmcsCW[i] can be in the range (OrgCW » 1) to (OrgCW « 1) - 1.
[0293] In one embodiment according to the present document, it is proposed to use a single chroma residual scaling factor for chroma residual scaling in LMCS. The existing method of deriving the chroma residual scaling factor uses the average of the corresponding luma block and derives the slope of each segment of the inverse luma mapping as the corresponding scaling factor. In addition, the process of identifying the segment index requires the availability of the corresponding luma block, which causes a delay problem. This is not desirable for hardware implementation. According to this embodiment of the present document, the scaling in the chroma block can not depend on the luma block value and can not require the identification of the segment index. Therefore, the chroma residual scaling process in LMCS can be performed without the delay problem.
[0294] In one embodiment according to the present document, a single chroma scaling factor can be derived from both the encoder and the decoder based on the luma LMCS information. When the LMCS luma model is received, the chroma residual scaling factor can be updated. For example, the single chroma residual scaling factor can be updated when the LMCS model is updated.
[0295] The following table shows an example of obtaining a single chroma scaling factor according to the present embodiment.
[0296] [Table 24]
[0297]
[0298] Referring to Table 24, a single chroma scaling factor (e.g., ChromaScaleCoeff or ChromaScaleCoeffSingle) can be obtained by averaging the inverse luma mapping slopes of all segments within LMCS_min_bin_idx and lmcs_max_bin_idx.
[0299] Figure 11 A linear fitting of the pivot points according to an embodiment of the present document is shown. In Figure 11 , pivot points P1, Ps and P2 are shown. The following embodiments or examples thereof will be described in Figure 11 .
[0300] In an example of this embodiment, a single chroma scaling factor can be obtained based on a linear approximation of the luminance PWL mapping between the pivot points lmcs_min_bin_idx and lmcs_max_bin_idx + 1 (LmcsMaxBinIdx + 1). That is, the inverse slope of the linear mapping can be used as the chroma residual scaling factor. For example, Figure 11 The linear line 1 can be a straight line connecting the pivot points P1 and P2. Referring to Figure 11 In P1, the input value is x1 and the mapping value is 0, and in P2, the input value is x2 and the mapping value is y2. The inverse slope (inverse scale) of the linear line 1 is (x2 - x1) / y2, and a single chroma scaling factor ChromaScaleCoeffSingle can be calculated based on the input values and mapping values of the pivot points P1, P2 and the following equation.
[0301] [Equation 5]
[0302] ChromaScaleCoeffSingle = (x2 - x1) * (1 « CSCALE_FP_PREC) / y2
[0303] In Equation 5, CSCALE_FP_PREC represents a shift factor, for example, CSCALE_FP_PREC can be a predetermined constant value. In one example, CSCALE_FP_PREC can be 11.
[0304] In another example according to this embodiment, referring to Figure 11 , the input value at the pivot point Ps is min_bin_idx + 1, and the mapping value at the pivot point Ps is ys. Therefore, the inverse slope (inverse scale) of the linear line 1 can be calculated as (xs - x1) / ys, and a single chroma scaling factor ChromaScaleCoeffSingle can be calculated based on the input values and mapping values of the pivot points P1, Ps and the following equation.
[0305] [Equation 6]
[0306] ChromaScaleCoeffSingle = (xs - x1) * (1 « CSCALE_FP_PREC) / ys
[0307] In Equation 6, CSCALE_FP_PREC represents a shift factor (bit shift factor), for example, CSCALE_FP_PREC can be a predetermined constant value. In one example, CSCALE_FP_PREC can be 11, and a bit shift of the inverse scale can be performed based on CSCALE_FP_PREC.
[0308] In another example according to this embodiment, a single chroma residual scaling factor can be derived based on a linear approximation line. An example of deriving a linear approximation line can include a linear connection of the pivot points (i.e., lmcs_min_bin_idx, lmcs_max_bin_idx + 1). For example, the linear approximation result can be represented by the code word of a PWL mapping. The mapping value y2 at P2 can be the sum of the code words of all bins (slices), and the difference (x2 - xl) of the input value at P2 and the input value at P1 is OrgCW * (lmcs_max_bin_idx - lmcs_min_bin_idx + 1) (for OrgCW, see Table 19 and / or Table 20 above). The following equation shows an example of obtaining a single chroma scaling factor according to the above embodiment.
[0309] [Table 25]
[0310]
[0311] Referring to Table 25, a single chroma scaling factor (e.g., ChromaScaleCoeffSingle) can be obtained from two pivot points (i.e., lmcs_min_bin_idx, lmcs_max_bin_idx). For example, the inverse slope of a linear mapping can be used as a chroma scaling factor.
[0312] In another example of this embodiment, a single chroma scaling factor can be obtained by a linear fitting of the pivot points to minimize the error (or mean square error) between the linear fitting and the existing PWL mapping. This example can be more accurate than simply connecting the two pivot points at lmcs_min_bin_idx and lmcs_max_bin_idx. There are many ways to find the optimal linear mapping, and an example of this is described below.
[0313] In one example, the parameters b1 and b0 of the linear fitting formula y = b1 * x + b0 to minimize the sum of the least square errors can be calculated based on Equation 7 and / or Equation 8 below.
[0314] [Equation 7]
[0315]
[0316] [Equation 8]
[0317]
[0318] In Equation 7 and Equation 8, x is the original luma value, y is the reshaped luma value, and are the mean values of x and y, x i and y i represent the value of the i-th pivot point.
[0319] Referring to Figure 11 Another simple approximation to identify the linear mapping is given as:
[0320] - Get linear line 1 by connecting the pivot points of the PWL mapping at lmcs_min_bin_idx and lmcs_max_bin_idx + 1, calculate lmcs_pivots_linear[i] of this linear line in multiples of OrgW
[0321] - Sum the difference of the pivot point mapped values using the linear line 1 and using the PWL mapping.
[0322] - Get the average difference avgDiff.
[0323] - Adjust the last pivot point of the linear line according to the average difference, e.g. 2*avgDiff
[0324] - Use the inverse slope of the adjusted linear line as the chroma residual scale.
[0325] According to the above linear fitting, the chroma scale factor (i.e. the inverse slope of the forward mapping) can be derived (obtained) based on the following equation 9 or equation 10.
[0326] [Equation 9]
[0327] ChromaScaleCoeffSingle = OrgCW * (1 « CSCALE_FP_PREC) / lmcs_pivots_linear[lmcs_min_bin_idx + 1]
[0328] [Equation 10]
[0329] ChromaScaleCoeffSingle = OrgCW * (lmcs_max_bin_idx - lmcs_max_bin_idx + 1) * (1 « CSCALE_FP_PREC) / lmcs_pivots_linear[lmcs_max_bin_idx + 1]
[0330] In the above equations, lmcs_pivots_linear[i] can be the mapped values of the linear mapping. For the linear mapping, all segments of the PWL mapping between the minimum bin index and the maximum bin index can have the same LMCS codeword (lmcsCW). That is, lmcs_pivots_linear[lmcs_min_bin_idx + 1] can be the same as lmcsCW[lmcs_min_bin_idx].
[0331] In addition, in Formula 9 and Formula 10, CSCALE_FP_PREC represents a shift factor (bit shift factor), for example, CSCALE_FP_PREC can be a predetermined constant value. In one example, CSCALE_FP_PREC can be 11.
[0332] With a single chroma residual scaling factor (ChromaScaleCoeffSingle), there is no need to calculate the average of the corresponding luma block anymore, find the index in the PWL linear mapping to get the chroma residual scaling factor. Therefore, the coding efficiency with chroma residual scaling can be increased. This not only eliminates the dependency on the corresponding luma block, solves the delay problem, but also reduces the complexity.
[0333] The semantics related to LMCS data and / or the luma-dependent chroma residual scaling process of the chroma samples according to the above-described embodiments can be described in the standard document format as shown in the following table.
[0334] [Table 26]
[0335]
[0336] [Table 27]
[0337]
[0338]
[0339] In another embodiment of the present document, the encoder can determine the parameters related to the single chroma scaling factor and signal them to the decoder. With the signaling, the encoder can use other information available at the encoder to derive the chroma scaling factor. This embodiment aims to eliminate the chroma residual scaling delay problem.
[0340] For example, another example of identifying the linear mapping to be used for determining the chroma residual scaling factor is given as follows:
[0341] - Calculate lmcs_pivots_linear[i] of this linear line with input values in multiples of OrgW by connecting the pivot points of the PWL mapping at lmcs_min_bin_idx and lmcs_max_bin_idx + 1
[0342] - Use those of linear line 1 and the luma PWL mapping to get the weighted sum of the difference of the mapping values of the pivot points. The weights can be based on the encoder statistics (e.g., histogram of bins).
[0343] - Get the weighted average difference avgDiff.
[0344] - adjust the last pivot point of the linear line 1 according to the weighted average difference, e.g., 2*avgDiff
[0345] - use the adjusted linear line's inverse slope to calculate the chroma residual scale.
[0346] The following shows an example of signaling the syntax for the y value used for chroma scaling factor derivation.
[0347] [Table 28]
[0348]
[0349] In Table 28, the syntax element lmcs_chroma_scale can specify a single chroma (residual) scaling factor (ChromaScaleCoeffSingle = lmcs_chroma_scale) for LMCS chroma residual scaling. That is, the information about the chroma residual scaling factor can be directly signaled, and the signaled information can be derived as the chroma residual scaling factor. In other words, the value of the signaled information about the chroma residual scaling factor can be (directly) derived as the value of the single chroma residual scaling factor. Here, the syntax element lmcs_chroma_scale can be signaled together with other LMCS data (i.e., syntax elements related to the absolute value and sign of the codeword, etc.).
[0350] Alternatively, the encoder can signal only the necessary parameters to derive the chroma residual scaling factor at the decoder. To derive the chroma residual scaling factor at the decoder, the input value x and the mapped value y are needed. Since the x value is the bin length and is a known number, it does not need to be signaled. After all, only the y value needs to be signaled in order to derive the chroma residual scaling factor. Here, the y value can be the mapped value of any pivot point in the linear mapping (i.e., the mapped value of P2 or Ps in Figure 11
[0351] The following shows an example of signaling the syntax for the mapped value used for deriving the chroma residual scaling factor.
[0352] [Table 29]
[0353]
[0354] [Table 30]
[0355]
[0356] One of the above Tables 29 and 30 can be used to signal the y value at any linear pivot point specified by the encoder and the decoder. That is, the encoder and the decoder can use the same syntax to derive the y value.
[0357] First, an embodiment according to Table 29 is described. In Table 29, lmcs_cw_linear can represent a mapping value at Ps or P2. That is, in the embodiment according to Table 29, a fixed number can be signaled by lmcs_cw_linear.
[0358] In an example according to this embodiment, if lmcs_cw_linear represents a mapping value of one bin (i.e., lmcs_pivots_linear[lmcs_min_bin_idx + 1]) at Ps, the chroma scale factor can be derived based on the following equation. Figure 11
[0359] [Equation 11]
[0360] ChromaScaleCoeffSingle = OrgCW * (1 « CSCALE_FP_PREC) / lmcs_cw_linear
[0361] In another example according to this embodiment, if lmcs_cw_linear represents lmcs_max_bin_idx + 1 (i.e., lmcs_pivots_linear[lmcs_max_bin_idx + 1]) at P2, the chroma scale factor can be derived based on the following equation. Figure 11
[0362] [Equation 12]
[0363] ChromaScaleCoeffSingle = OrgCW * (lmcs_max_bin_idx - lmcs_max_bin_idx + 1) * (1 « CSCALE_FP_PREC) / lmcs_cw_linear
[0364] In the above equations, CSCALE_FP_PREC represents a shift factor (bit shift factor), for example, CSCALE_FP_PREC can be a predetermined constant value. In one example, CSCALE_FP_PREC can be 11.
[0365] Next, an embodiment according to Table 30 is described. In this embodiment, lmcs_cw_linear can be signaled as a delta value with respect to a fixed number (i.e., lmcs_delta_abs_cw_linear, lmcs_delta_sign_cw_linear_flag). In an example of this embodiment, when lmcs_cw_linear represents lmcs_pivots_linear[lmcs_min_bin_idx + 1] (i.e., Figure 11 lmcs_cw_linear_delta and lmcs_cw_linear can be derived based on the following equations.
[0366] [Equation 13]
[0367] lmcs_cw_linear_delta = (1 - 2 * lmcs_delta_sign_cw_linear_flag) * lmcs_delta_abs_linear_cw
[0368] [Equation 14]
[0369] lmcs_cw_linear = lmcs_cw_linear_delta + OrgCW
[0370] In another example of this embodiment, when lmcs_cw_linear represents the mapped value in P2) of the table of FIG. 19, lmcs_cw_linear_delta and lmcs_cw_linear can be derived based on the following equations. Figure 11
[0371] [Equation 15]
[0372] lmcs_cw_linear_delta = (1 - 2 * lmcs_delta_sign_cw_linear_flag) * lmcs_delta_abs_linear_cw
[0373] [Equation 16]
[0374] lmcs_cw_linear = lmcs_cw_linear_delta + OrgCW * (lmcs_max_bin_idx - lmcs_max_bin_idx + 1)
[0375] In the above equations, OrgCW can be a value derived based on the above Table 19 and / or Table 20.
[0376] The luma-dependent chroma residual scaling process for semantics and / or chroma samples related to LMCS data according to the above embodiment can be described in a standard document format as shown in the following table.
[0377] [Table 31]
[0378]
[0379] [Table 32]
[0380]
[0381] Figure 12 One example of linear shaping (or linear shaping, linear mapping) according to an embodiment of the present document is shown. That is, in this embodiment, it is proposed to use a linear shaper in LMCS. For example, Figure 12 This example in Table 33 can involve forward linear shaping (mapping).
[0382] Referring to Figure 12 , the linear shaper can include two pivot points, i.e., P1 and P2. P1 and P2 can represent input values and mapped values, e.g., P1 can be (min_input, 0) and P2 can be (max_input, max_mapped). Here, min_input represents a minimum input value, and max_input represents a maximum input value. Any input value less than or equal to min_input is mapped to 0, and any input value greater than max_input is mapped to max_mapped. Any input luminance value within min_input and max_input is linearly mapped to other values. Figure 12 An example of mapping is shown. The pivot points P1, P2 can be determined at the encoder, and a linear fit can be used to approximate the piecewise linear mapping.
[0383] In another embodiment according to the present document, another example of a method of signaling a linear shaper can be proposed. The pivot points P1, P2 of the linear shaper model can be explicitly signaled. The following shows an example of explicitly signaling the syntax and semantics of the linear shaper model according to this example.
[0384] [Table 33]
[0385]
[0386] [Table 34]
[0387]
[0388] Referring to Table 33 and Table 34, the input value of the first pivot point can be derived based on the syntax element lmcs_min_input, and the input value of the second pivot point can be derived based on the syntax element lmcs_max_input. The mapped value of the first pivot point can be a predetermined value (a value known by both the encoder and the decoder), e.g., the mapped value of the first pivot point is 0. The mapped value of the second pivot point can be derived based on the syntax element lmcs_max_mapped. That is, the linear shaper model can be explicitly (directly) signaled based on the information signaled in the syntax of Table 33.
[0389] Alternatively, lmcs max input and lmcs max mapped can be signaled as delta values. The following shows an example of syntax and semantics for signaling the linear reshaper model as delta values.
[0390] [Table 35]
[0391]
[0392] [Table 36]
[0393]
[0394] Referring to Table 36, the input value of the first pivot point can be derived based on the syntax element lmcs min input. For example, lmcs min input can have a mapped value of 0. lmcs max input delta can specify the difference between the input value of the second pivot point and the maximum luma value (i.e., (1 « bitdepthY) - 1). lmcs max mapped delta can specify the difference between the mapped value of the second pivot point and the maximum luma value (i.e., (1 « bitdepthY) - 1).
[0395] According to embodiments of the present document, the forward mapping of luma prediction samples, the inverse mapping of luma reconstructed samples, and the chroma residual scaling can be performed based on the above examples of linear reshapers. In one example, the inverse scaling for luma (reconstructed) samples (pixels) in the inverse mapping based on linear reshapers can only need one inverse scaling factor. The same is true for the forward mapping and the chroma residual scaling. That is, the steps of determining ScaleCoeff[i], InvScaleCoeff[i], and ChromaScaleCoeff[i] (i is the bin index) can be replaced by only one single factor. Here, one single factor is a fixed point representation of the (positive) slope or inverse slope of the linear mapping. In one example, the inverse luma mapping scaling factor (inverse scaling factor in the inverse mapping of luma reconstructed samples) can be derived based on at least one of the following equations.
[0396] [Equation 17]
[0397] InvScaleCoeffSingle = OrgCW / kmcsCWLinear
[0398] [Equation 18]
[0399] InvScaleCoeffSingle = OrgCW*(lmcs max bin idx - lmcs max bin idx + 1) / lmcsCWLinearAll
[0400] [Formula 19]
[0401] InvScaleCoeffSingle = (lmcs_max_input - lmcs_min_input) / lmcsCWLinearAll
[0402] lmcsCWLinear of Formula 17 can be derived from Table 31 described above. lmcsCWLinearALL of Formula 18 and Formula 19 can be derived from at least one of Table 33 to Table 36 described above. In Formula 17 or Formula 18, OrgCW can be derived from Table 19 and / or Table 20.
[0403] The following table describes a formula and a syntax (a conditional statement) indicating a forward mapping process of a luma sample (i.e., a luma prediction sample) in picture reconstruction. In the following table and formula, FP_PREC is a constant value of bit shift, and can be a predetermined value. For example, FP_PREC can be 11 or 15.
[0404] [Table 37]
[0405]
[0406] [Table 38]
[0407]
[0408] Table 37 can be used to derive a forward mapped luma sample based on Table 17 to Table 20 described above in a luma mapping process. That is, Table 37 can be described together with Table 19 and Table 20. In Table 37, a forward mapped luma (prediction) sample PredMapPSamples[i][j] as an output can be derived from a luma (prediction) sample predSamples[i][j] as an input. idxY of Table 37 can be referred to as a (forward) mapping index, which can be derived based on a prediction luma sample.
[0409] Table 38 can be used to derive a forward mapped luma sample in linear reshaper based luma mapping. For example, lmcs_min_input, lmcs_max_input, lmcs_max_mapped, and ScaleCoeffSingle of Table 38 can be derived by at least one of Tables 33 to 36. In Table 38, in case of “lmcs_min_input < predSamples[i][j] < lmcs_max_input”, a forward mapped luma (prediction) sample PredMapSamples[i][j] can be derived from an input luma (prediction) sample predSamples[i][j] as an output. Comparing between Table 37 and Table 38, a change from the existing LMCS to the application of linear reshaper can be seen from a forward mapping perspective.
[0410] The following formula and table describe an inverse mapping process of a luma sample (i.e., a luma reconstruction sample). In the following formula and table, “lumaSample” as input can be a luma reconstruction sample before inverse mapping (before modification). “invSample” as output can be an inverse mapped (modified) luma reconstruction sample. In other cases, the clipped invSample can be referred to as a modified luma reconstruction sample.
[0411] [Formula 20]
[0412] invSample = InputPivot[idxYInv] + (InvScaleCoeff[idxYInv] * (lumaSample - LmcsPivot[idxYInv]) + (1 << (FP_PREC - 1))) >> FP_PREC
[0413] [Formula 21]
[0414] invSample = lmcs_min_input + (InvScaleCoeffSingle * (lumaSample - lmcs_min_input) + (1 << (FP_PREC - 1))) >> FP_PREC
[0415] [Table 39]
[0416]
[0417] [Table 40]
[0418]
[0419] Formula 21 can be used to derive inverse mapped luma samples in luma mapping according to the present document. In formula 20, index idxInv can be derived based on table 50, table 51 or table 52 described later.
[0420] Formula 21 can be used to derive inverse mapped luma samples from luma mapping according to the application of linear reshaper. For example, lmcs_min_input of formula 21 can be derived from at least one of table 33 to table 36. By comparing between formula 20 and formula 21, the change of the application of linear reshaper with respect to existing LMCS can be seen from the perspective of forward mapping.
[0421] Table 39 can include examples of formulas to derive inverse mapped luma samples in luma mapping. For example, index idxInv can be derived based on table 50, table 51 or table 52 described later.
[0422] Table 40 can include other examples of formulas to derive inverse mapped luma samples in luma mapping. For example, lmcs_min_input and / or lmcs_max_mapped of table 40 can be derived by at least one of table 33 to table 36, and / or InvScaleCoeffSingle of table 40 is at least table 33 to table 36, and / or formula 17 to formula 19 can be derived by one.
[0423] Based on the above examples of linear reshaper, the segment index identification process can be omitted. That is, in the present example, since there is only one segment with valid reshaped luma pixels, the segment index identification process for inverse luma mapping and chroma residual scaling can be removed. Therefore, the complexity of inverse luma mapping can be reduced. In addition, the delay problem caused by the dependence on luma segment index identification during chroma residual scaling can be eliminated.
[0424] According to the embodiments using the above linear reshaper, the following advantages can be provided for LMCS: i) the encoder reshaper design can be simplified, preventing possible artifacts caused by sudden changes between segment linear segments; ii) the decoder inverse mapping process, which can remove the segment index identification process, can be simplified by eliminating the segment index identification process; iii) by removing the segment index identification process, the delay problem caused by the dependence on corresponding luma blocks in chroma residual scaling can be removed; iv) the signaling overhead can be reduced, and frequent updates of the reshaper can be made more feasible; v) for many places where a loop of 16 segments was needed in the past, the loop can be eliminated. For example, in order to derive InvScaleCoeff[i], the number of division operations according to lmcsCW[i] can be reduced to 1.
[0425] In another embodiment according to the present document, LMCS based on flexible bins is proposed. Here, flexible bins can refer to bins whose number is not fixed to a predetermined (predefined, specific) number. In the existing embodiment, the number of bins in LMCS is fixed to 16, which are equally distributed for input sample values. In this embodiment, a flexible number of bins is proposed, and these segments (bins) can not be equally distributed in terms of original pixel values.
[0426] The following table exemplarily shows the syntax of LMCS data (data field) according to this embodiment and semantics of the syntax elements included therein.
[0427] [Table 41]
[0428]
[0429] [Table 42]
[0430]
[0431] Referring to Table 41, information on the number of bins lmcs_num_bins_minus1 can be signaled. Referring to Table 42, lmcs_num_bins_minus1 + 1 can be equal to the number of bins, which can be in the range from 1 to (1 « BitDepthY) - 1. For example, lmcs_num_bins_minus1 or lmcs_num_bins_minus1 + 1 can be a multiple of a power of 2.
[0432] In the embodiment described together with Table 41 and Table 42, regardless of whether the reshaper is linear (signaling of lmcs_num_bins_minus1), the number of pivot points can be derived based on lmcs_num_bins_minus1 (information on the number of bins), and the input value and mapped value of the pivot point (LmcsPivot_input[i], LmcsPivot_mapped[i]) can be derived based on the sum of the signaled codeword values (lmcs_delta_input_cw[i], lmcs_delta_mapped_cw[i]) (here, the initial input value LmcsPivot_input[0] and the initial output value LmcsPivot_mapped[0] are 0).
[0433] Figure 13 An example of linear forward mapping in the embodiment of the present document is shown. Figure 14 An example of inverse forward mapping in the embodiment of the present document is shown.
[0434] In accordance with the present document, a method of encoding video data is proposed. Figure 13 and Figure 14In an embodiment, a method is proposed that supports both regular LMCS and linear LMCS. In an example according to this embodiment, regular LMCS and / or linear LMCS can be indicated based on a syntax element lmcs_is_linear. In an encoder, after determining the linear LMCS line, the mapping values (i.e., pL in Figure 13 and Figure 14 The mapping values in pL can be divided into equal segments (i.e., LmcsMaxBinIdx - lmcs_min_bin_idx + 1). The codewords in bin LmcsMaxBinIdx can be signaled using the syntax of the LMCS data or shaper mode described above.
[0435] The following table exemplarily shows the syntax of LMCS data (data field) and semantics of the syntax elements included therein according to an example of this embodiment.
[0436] [Table 43]
[0437]
[0438] [Table 44]
[0439]
[0440]
[0441] The following table exemplarily shows the syntax of LMCS data (data field) and semantics of the syntax elements included therein according to another example of this embodiment.
[0442] [Table 45]
[0443]
[0444] [Table 46]
[0445]
[0446]
[0447] Referring to Tables 43 to 46, when lmcs_is_linear_flag is true, all LMCSDeltaCW[i] between lmcs_min_bin_idx and LmcsMaxBinIdx can have the same value. That is, lmcsCW[i] for all slices between lmcs_min_bin_idx and LmcsMaxBinIdx can have the same value. The scale and de-scale and chroma scale for all slices between lmcs_min_bin_idx and LmcsMaxBinIdx can be the same. Then if linear reshaper is true, the slice index does not need to be derived, it can use scale, de-scale only from one slice.
[0448] The following shows an exemplary illustration of the identification process of the slice index according to the present embodiment.
[0449] [Table 47]
[0450]
[0451] According to another embodiment of the present document, the application of regular 16-slice PWL LMCS and linear LMCS can depend on a higher level syntax, i.e., sequence level.
[0452] The following shows an exemplary illustration of the syntax of SPS and semantics of the syntax elements included therein according to the present embodiment.
[0453] [Table 48]
[0454]
[0455] [Table 49]
[0456]
[0457] Referring to Tables 48 and 49, the enabling of regular LMCS and / or linear LMCS can be determined (signaled) by the syntax elements included in SPS. Referring to Table 48, based on the syntax element sps_linear_lmcs_enabled_flag, one of regular LMCS or linear LMCS can be used in sequence unit.
[0458] In addition, whether only linear LMCS or regular LMCS is enabled or both can also depend on the profile level. In one example, for a certain profile, i.e., SDR profile, only linear LMCS can be allowed, for another profile, i.e., HDR profile, only regular LMCS can be allowed, for another profile, both regular LMCS and / or linear LMCS can be allowed.
[0459] According to another embodiment of the present document, the LMCS segment index identification process can be used in inverse luminance mapping and chroma residual scaling. In this embodiment, the segment index identification process can be used for the blocks in which chroma residual scaling is enabled, and also invoked for all luminance samples in the shaping (mapping) domain. This embodiment aims to keep its complexity low.
[0460] The following table shows the existing segment function index identification process (derivation process).
[0461] [Table 50]
[0462]
[0463] In an example, in the segment index identification process, the input sample can be classified into at least two categories. For example, the input sample can be classified into three categories, a first, a second, and a third category. For example, the first category can represent samples (values) less than LmcsPivot[lmcs_min_bin_idx+1], the second category can represent samples (values) greater than or equal to LmcsPivot[LmcsMaxBinIdx] (values), and the third category can indicate samples (values) between LmcsPivot[lmcs_min_bin_idx+1] and LmcsPivot[LmcsMaxBinIdx].
[0464] In this embodiment, it is proposed to optimize the identification process by eliminating the category classification. This is because the input of the segment index identification process is the luminance value in the shaping (mapping) domain, and there should be no value exceeding the mapping value at the pivot points lmcs_min_bin_idx and LmcsMaxBinIdx+1. Therefore, the conditional process of classifying the sample into categories in the existing segment index identification process is unnecessary. For more details, specific examples will be described in tables below.
[0465] In an example according to this embodiment, the identification process included in Table 50 can be replaced by one of the following Tables 51 or 52. Referring to Tables 51 and 52, the first two categories of Table 50 can be removed, and for the last category, the boundary value (second boundary value or end point) in the iterative for loop is changed from LmcsMaxBinIdx to LmcsMaxBinIdx+1. That is, the identification process can be simplified, and the complexity of the segment index derivation can be reduced. Therefore, the LMCS related encoding can be efficiently performed according to this embodiment.
[0466] [Table 51]
[0467]
[0468] [Table 52]
[0469]
[0470] Referring to Table 51, a comparison process (formula corresponding to the condition of the if statement) corresponding to the condition of the if statement can be iteratively performed for all bin indices from the minimum bin index to the maximum bin index. In the case where the formula corresponding to the condition of the if statement is true, the bin index can be derived as a reverse mapping index of the reverse luma mapping (or a reverse scaling index of the chroma residual scaling). Based on the reverse mapping index, a modified reconstructed luma sample (or a scaled chroma residual sample) can be derived.
[0471] According to the embodiment of Table 51, a problem that can occur in a process of identifying a (reverse) piecewise function index can be solved. Due to the embodiment of Table 51, a calculation loophole can be eliminated and / or a redundant calculation process caused by an overlapping boundary condition can be omitted. If the present document does not follow the embodiment of Table 51, a mapping value used in the LMCS of the current block in a for syntax (loop syntax) for identifying a piecewise index (reverse mapping index) can exceed (deviate from) LmcsPivot[idxYInv+1]. According to this embodiment, a mapping value within an appropriate range used in the LMCS of the current block in the index identification process can be used.
[0472] The following table shows an effect of improving encoding performance by the embodiment of Table 51.
[0473] [Table 53]
[0474]
[0475] [Table 54]
[0476]
[0477] [Table 55]
[0478]
[0479] Referring to Tables 53 to 55, encoding performance can be improved according to the embodiment of Table 51. Alternatively, according to the embodiment of Table 51, a calculation loophole can be removed and / or a redundant calculation process can be omitted while maintaining encoding performance.
[0480] In an embodiment according to the present document, in the existing embodiment, for a slice (e.g., an intra slice) coded as a separate block tree, a delay due to chroma residual scaling dependency in a corresponding luma block is higher than that for a slice coded as a dual tree. Therefore, in the existing embodiment, LMCS chroma residual scaling is not applied to a slice coded as a separate block tree.
[0481] In this embodiment, chroma residual scaling can be applied even to slices coded as separate trees. When using the single chroma residual scaling factor described above, there can be no delay caused by the application of chroma residual scaling since there is no dependency between the chroma residual scaling in the corresponding luma block.
[0482] The following table shows the syntax and semantics of slice header according to the present embodiment.
[0483] [Table 56]
[0484]
[0485] [Table 57]
[0486]
[0487] Referring to Table 57, the chroma residual scaling flag can be signaled regardless of the conditional clause (or its flag) indicating whether the current block has a dual tree structure or a single tree structure.
[0488] In embodiments according to the present document, ALF data and / or LMCS data can be signaled in APS. For example, 32 APSs can be used. In one example, if all APSs are used for ALF and / or LMCS, the buffering of APSs can require about 10 KB of on-chip memory. To limit (reduce) the memory required to store ALF / LMCS parameters and the computational complexity required by LMCS, this embodiment proposes a method to limit the number of ALF and / or LMCS APSs.
[0489] In examples according to this embodiment, one LMCS model per picture can be used (allowed) regardless of the number of slices or tiles included in the picture. The number of APSs for LMCS can be less than 32. For example, the number of APSs for LMCS can be four.
[0490] The following table shows the semantics related to APS according to the present embodiment.
[0491] [Table 58]
[0492]
[0493] Referring to Table 58, the maximum number of LMCS APSs can be predetermined. For example, the maximum number of LMCS APSs can be four. Referring to Table 15 and Table 58, multiple APSs can include LMCS APSs. The (maximum) number of LMCS APSs can be four. In one example, the LMCS data field included in one of the LMCS APSs can be used in the LMCS process of the current block in the current picture.
[0494] The following table shows semantics of the syntax elements included in the slice header (or picture header).
[0495] [Table 59]
[0496]
[0497] The syntax elements described from Table 59 can be described with reference to Table 56. In an example, the syntax element slice_lmcs_aps_id can be included in the slice header. In another example, the syntax element slice_lmcs_aps_id of Table 56 can be included in the picture header, and in this case, slice_lmcs_aps_id can be modified to ph_lmcs_aps_id.
[0498] According to an embodiment of the present document, the total number of LMCS APSs can be less than or equal to four. In addition, in an example of this embodiment, only one LMCS model per picture can be used. In this case, only one LMCS model per picture can be used regardless of the resolution of the image / video. According to this embodiment, since there is a limitation on the LMCS APS, implementation can be made easy, and the problem of excessive consumption of resources (memory) can be solved.
[0499] The following table shows semantics of the syntax element examples disclosed in the present document.
[0500] [Table 60]
[0501]
[0502] The syntax elements described from Table 60 can be described with reference to Table 56. In an example, the syntax element slice_lmcs_aps_id can be included in the slice header. In another example, the syntax element slice_lmcs_aps_id of Table 60 can be included in the picture header, and in this case, slice_lmcs_aps_id can be modified to ph_lmcs_aps_id.
[0503] Referring to Table 60, adaptation_parameter_set_id of the LMCS APS referred to by the current slice (or the current picture) can be designated by the syntax element slice_lmcs_aps_id. That is, slice_lmcs_aps_id can indicate adaptation_parameter_set_id of the LMCS APS referred to by the current slice (or the current picture). The value of slice_lmcs_aps_id can be in the range of 0 to 3. That is, the value of slice_lmcs_aps_id can be 0, 1, 2, or 3. For example, among the 0th to 3rd LMCS APSs (or the 1st to 4th LMCS APSs), the 0th LMCS APS (or the 1st LMCS APS) indicated by slice_lmcs_aps_id can be referred to by the current slice (or the current picture), and the LMCS data (LMCS data field or shaper model) included in the 0th LMCS APS (or the 1st LMCS APS) can be used for the LMCS process of the current slice (or the current picture) (the LMCS process of the current slice (or the current picture) can be performed based on the LMCS data (LMCS data field or shaper model) included in the 0th LMCS APS (or the 1st LMCS APS)).
[0504] The encoding process related to LMCS can be cleaned up or simplified according to at least one of the embodiments disclosed in the present document.
[0505] The following drawings are created to illustrate specific examples of the present specification. Since the names of specific apparatuses described in the drawings or the names of specific signals / messages / fields are presented as examples, the technical features of the present specification are not limited to the specific names used in the following drawings.
[0506] Figure 15 and Figure 16 An example of a video / image encoding method and an associated component according to an embodiment of the present document is schematically represented. Figure 15 The method disclosed in Figure 2 may be performed by the encoding device disclosed in Figure 15 S1500 of the encoding device; S1510 and / or S1520 can be performed by the predictor 220 and / or the adder 250 of the encoding device; S1520 can be performed by the residual processor 230 or the adder 250 of the encoding device; S1530 and / or S1540 can be performed by the residual processor 230 or the adder 250 of the encoding device; S1550 to S1570 can be performed by the residual processor 230 of the encoding device; and S1580 can be performed by the entropy encoder 240 of the encoding device. Figure 15The methods disclosed herein can include the embodiments described above in this document.
[0507] Referring to Figure 15 The encoding device can generate predicted luma samples (S1500). Regarding the predicted luma samples, the encoding device can derive the predicted luma samples of the current block based on the prediction mode. In this case, various prediction methods disclosed in this document, such as inter prediction or intra prediction, can be applied.
[0508] The encoding device can derive predicted chroma samples. The encoding device can derive residual chroma samples based on the original chroma samples and the predicted chroma samples of the current block. For example, the encoding device can derive the residual chroma samples based on the difference between the predicted chroma samples and the original chroma samples.
[0509] The encoding device can derive bins for luma mapping and / or LMCS codewords (S1510). The encoding device can derive individual LMCS codewords for a plurality of bins. For example, the lmcsCW[i] described above can correspond to the LMCS codeword derived by the encoding device.
[0510] The encoding device can generate mapped predicted luma samples (S1520). For example, the encoding device can derive input values and mapped values (output values) for the pivot points of the luma mapping, and can generate the mapped predicted luma samples based on the input values and the mapped values. In an example, the encoding device can derive a mapping index (idxY) based on the first predicted luma samples, and can generate the first mapped predicted luma samples based on the mapped values and the input values of the pivot points corresponding to the mapping index. In another example, a linear mapping (linear reshaping, linear LMCS) can be used, and the mapped predicted luma samples can be generated based on a forward mapping scaling factor derived from two pivot points in the linear mapping, so that the index derivation process can be omitted due to the linear mapping.
[0511] The encoding device can generate reconstructed luma samples (S1530). The encoding device can generate the reconstructed luma samples based on the mapped predicted luma samples. Specifically, the encoding device can sum the residual luma samples described above and the mapped predicted luma samples, and generate the reconstructed luma samples based on the result of the summation.
[0512] The encoding device can generate modified reconstructed luma samples based on bins of the reconstructed luma samples, the LMCS codeword, and bins of the luma mapping (S1540). The encoding device can generate the modified reconstructed luma samples by an inverse mapping process on the recovered luma samples. For example, the encoding device can derive inverse mapping indices (e.g., invYldx) based on mapping values (e.g., LmcsPivot[i], i = lmcs min bin idx... LmcsMaxBinldx + 1) assigned to individual reconstructed luma samples and / or bin indices in the inverse mapping process. The encoding device can generate the modified reconstructed luma samples based on a mapping value (LmcsPivot[invYldx]) assigned to the inverse mapping indices.
[0513] The encoding device can generate scaled residual chroma samples. Specifically, the encoding device can derive chroma residual scaling factors and generate the scaled residual chroma samples based on the chroma residual scaling factors. Here, the chroma residual scaling of the encoding stage can be referred to as forward chroma residual scaling. Thus, the chroma residual scaling factors derived by the encoding device can be referred to as forward chroma residual scaling factors, and the scaled residual chroma samples can be generated as forward scaled residual chroma samples.
[0514] The encoding device can derive information on the LMCS data based on bins of the luma mapping and the LMCS codeword (S1550). Alternatively, the encoding device can generate the information on the LMCS data based on the mapped predicted luma samples and / or the scaled residual chroma samples. The encoding device can generate the information on the LMCS data of the reconstructed samples. The encoding device can derive LMCS related parameters that can be applied to filtering the reconstructed samples, and can generate the information on the LMCS data based on the LMCS related parameters. For example, the information on the LMCS data can include information on the above-described luma mapping (e.g., forward mapping, inverse mapping, linear mapping), information on the chroma residual scaling, and / or LMCS (or reshaping, reshaper) related indices (e.g., maximum bin index and minimum bin index).
[0515] The encoding device can generate residual luma samples based on the mapped predicted luma samples (S1560). For example, the encoding device can derive the residual luma samples based on a difference between the mapped predicted luma samples and the original luma samples.
[0516] The encoding device can derive residual information (S1570). The encoding device can derive the residual information based on the scaled residual chroma samples and / or the residual luma samples. The encoding device can derive transform coefficients based on a transform process on the scaled residual chroma samples and the luma residual samples. For example, the transform process can include at least one of DCT, DST, GBT, or CNT. The encoding device can derive quantized transform coefficients based on a quantization process on the transform coefficients. The quantized transform coefficients can have a one-dimensional vector form based on a coefficient scan order. The encoding device can generate residual information specifying the quantized transform coefficients. The residual information can be generated by various encoding methods such as exponential Golomb, CAVLC, CABAC, etc.
[0517] The encoding device can encode the image / video information (S1580). The image information can include LMCS-related information and / or residual information. For example, the LMCS-related information can include information on linear LMCS. In one example, at least one LMCS codeword can be derived based on the information on linear LMCS. The encoded video / image information can be output in the form of a bitstream. The bitstream can be transmitted to a decoding apparatus through a network or a storage medium.
[0518] According to embodiments of the present document, the image / video information can include various information. For example, the image / video information can include information disclosed in at least one of Tables 1 to 60 described above.
[0519] In embodiments, the image information can include an LMCS APS. The LMCS APS (one of the LMCS APSs) can include type information indicating that the LMCS APS is an APS including an LMCS data field and identifier information (ID information) of the LMCS APS. The LMCS data field can include information on LMCS data. The LMCS APS can include the LMCS data field based on the type information, and an LMCS codeword can be derived based on the LMCS data field. In one example, a value of the identifier information can be in a predetermined range. For example, the predetermined range can be in a range of 0 to 3. That is, the value of the identifier information can be 0, 1, 2, or 3. In another example, a maximum number of the LMCS APSs can be a predetermined value. For example, the maximum number (the predetermined value) of the LMCS APSs can be four.
[0520] In embodiments, the encoding device can generate a segment index for chroma residual scaling. The encoding device can derive a chroma residual scaling factor based on the segment index. The encoding device can generate scaled residual chroma samples based on the residual chroma samples and the chroma residual scaling factor.
[0521] In one embodiment, when the current block has a single tree structure or a dual tree structure (when the current block has a separate tree structure, when the current block is coded as a separate tree), a chroma residual scaling enabled flag indicating whether chroma residual scaling is applied for the current block can be generated by the encoding device. When the chroma residual scaling is applied for the current picture, the current slice, and / or the current block, the value of the chroma residual scaling enabled flag can be 1.
[0522] In an embodiment, the chroma residual scaling factor can be a single chroma residual scaling factor.
[0523] In an embodiment, based on the value of the type information being 1, the APS can contain a LMCS data field including LMCS parameters.
[0524] In an embodiment, the picture information can include header information. Here, the header information can be a picture header (or a slice header). The header information can include LMCS related APS ID information. The LMCS related APS ID information can indicate the identifier information of the LMCS APS of the current picture, the current block, or the current slice. That is, the value of the LMCS related APS ID information can be the same as the value of the identifier information of the LMCS APS. For example, the value of the LMCS related APS ID information can be in the range of 0 to 3. That is, the value of the LMCS related APS ID information can be 0, 1, 2, or 3.
[0525] In an embodiment, the picture information can include a sequence parameter set (SPS). The SPS can include a linear LMCS enabled flag indicating whether linear LMCS is enabled.
[0526] In an embodiment, a minimum bin index (e.g., lmcs_min_bin_idx) and / or a maximum bin index (e.g., LmcsMaxBinIdx) can be derived based on information about the LMCS data. A first mapping value (LmcsPivot[lmcs_min_bin_idx]) can be derived based on the minimum bin index. A second mapping value (LmcsPivot[LmcsMaxBinIdx] or LmcsPivot[LmcsMaxBinIdx+1]) can be derived based on the maximum bin index. The values of the reconstructed luma samples (e.g., lumaSample of Table 36 or 37) can be in the range of the first mapping value to the second mapping value. In one example, the values of all the reconstructed luma samples can be in the range of the first mapping value to the second mapping value. In another example, the values of some of the reconstructed luma samples can be in the range of the first mapping value to the second mapping value.
[0527] In an embodiment, the information on the LMCS data can include information on a linear LMCS and an LMCS data field. The information on the linear LMCS can be referred to as information on a linear mapping. The LMCS data field can include a linear LMCS flag indicating whether a linear LMCS is applied. When a value of the linear LMCS flag is 1, a predicted luma sample of the mapping can be generated based on the information on the linear LMCS.
[0528] In an embodiment, the information on the linear LMCS can include information on a first pivot point (e.g., P1 in Table 33) and information on a second pivot point (e.g., P2 in Table 33). For example, an input value and a mapping value of the first pivot point can be a minimum input value and a minimum mapping value, respectively. An input value and a mapping value of the second pivot point can be a maximum input value and a maximum mapping value, respectively. Input values between the minimum input value and the maximum input value can be linearly mapped. Figure 11 Figure 11 In an embodiment, the information on the linear mapping includes information on an input delta value of the second pivot point (i.e., lmcs_max_input_delta in Table 35) and information on a mapping delta value of the second pivot point (i.e., lmcs_max_mapped_delta in Table 35). The maximum input value can be derived based on the input delta value of the second pivot point, and the maximum mapping value can be derived based on the mapping delta value of the second pivot point.
[0529] In an embodiment, the information on the linear mapping includes information on an input delta value of the second pivot point (i.e., lmcs_max_input_delta in Table 35) and information on a mapping delta value of the second pivot point (i.e., lmcs_max_mapped_delta in Table 35). The maximum input value can be derived based on the input delta value of the second pivot point, and the maximum mapping value can be derived based on the mapping delta value of the second pivot point.
[0530] In an embodiment, the maximum input value and the maximum mapping value can be derived based on at least one formula included in Table 36 described above.
[0531] In an embodiment, the maximum input value and the maximum mapping value can be derived based on at least one formula included in Table 36 described above.
[0532] In an embodiment, generating the predicted luma sample of the mapping includes deriving a forward mapping scaling factor (i.e., ScaleCoeffSingle) of the predicted luma sample, and generating the predicted luma sample of the mapping based on the forward mapping scaling factor. The forward mapping scaling factor can be a single factor of the predicted luma sample.
[0533] In an embodiment, the forward mapping scaling factor can be derived based on at least one formula included in Table 36 and / or Table 38 described above.
[0534] In one embodiment, the mapped prediction luma sample can be derived based on at least one of the above-described Table 38.
[0535] In one embodiment, the encoding device can derive an inverse mapping scaling factor (i.e., InvScaleCoeffSingle) of the reconstructed luma sample (i.e., lumaSample). In addition, the encoding device can generate a modified reconstructed luma sample (i.e., invSample) based on the reconstructed luma sample and the inverse mapping scaling factor. The inverse mapping scaling factor can be a single factor of the reconstructed luma sample.
[0536] In one embodiment, the inverse mapping scaling factor can be derived using the segment index derived based on the reconstructed luma sample.
[0537] In one embodiment, the segment index can be derived based on the above-described Table 51. That is, the comparison process included in Table 51 (lumaSample < LmcsPivot[idxYInv + 1]) can be iteratively performed from the segment index that is the minimum bin index to the segment index that is the maximum bin index.
[0538] In one embodiment, the inverse mapping scaling factor can be derived based on at least one of the above-described Table 33, Table 34, Table 35, and Table 36 or Equation 11 or Equation 12.
[0539] In one embodiment, the modified reconstructed luma sample can be derived based on the above-described Equation 20, Equation 21, Table 39, and / or Table 40.
[0540] In one embodiment, the LMCS-related information can include information on the number of bins used to derive the mapped prediction luma sample (i.e., lmcs_num_bins_minus1 in Table 41). For example, the number of pivot points of the luma mapping can be set to be equal to the number of bins. In one example, the encoding device can generate the delta input value and the delta mapping value of the pivot points according to the number of bins, respectively. In one example, the input value and the mapping value of the pivot points are derived based on the delta input value (i.e., lmcs_delta_input_cw[i] in Table 41) and the delta mapping value (i.e., lmcs_delta_mapped_cw[i] in Table 41), and the mapped prediction luma sample can be generated based on the input value (i.e., LmcsPivot_input[i] in Table 42) and the mapping value (i.e., LmcsPivot_mapped[i] in Table 42).
[0541] In one embodiment, the encoding device can derive the LMCS delta codeword based on the at least one LMCS codeword and an original codeword (OrgCW) included in the LMCS-related information, and can derive the mapped luma prediction samples based on the at least one LMCS codeword and the original codeword. In one example, the information about the linear mapping can include information about the LMCS delta codeword.
[0542] In one embodiment, the at least one LMCS codeword can be derived based on a sum of the LMCS delta codeword and OrgCW, e.g., OrgCW is (1 « BitDepthY) / 16, where BitDepthY represents the luma bit depth. This embodiment can be based on Equation 12.
[0543] In one embodiment, the at least one LMCS codeword can be derived based on a sum of the LMCS delta codeword and OrgCW * (lmcs_max_bin_idx - lmcs_min_bin_idx + 1), e.g., lmcs_max_bin_idx and lmcs_min_bin_idx are the maximum bin index and the minimum bin index, respectively, and OrgCW can be (1 « BitDepthY) / 16. This embodiment can be based on Equation 15 and Equation 16.
[0544] In one embodiment, the at least one LMCS codeword can be a multiple of 2.
[0545] In one embodiment, the at least one LMCS codeword can be a multiple of 1 « (BitDepthY - 10) when the luma bit depth (BitDepthY) of the reconstructed luma samples is higher than 10.
[0546] In one embodiment, the at least one LMCS codeword can be in a range from (OrgCW » 1) to (OrgCW « 1) - 1.
[0547] Figure 17 and Figure 18 An example of an image / video decoding method and associated components according to embodiments of the present document is schematically represented. Figure 17 The method disclosed in Figure 3 may be performed by a decoding device. In particular, for example, Figure 17 S1700 of the method of FIG. 17 can be performed by the entropy decoder 310 of the decoding device, S1710 can be performed by the predictor 330 of the decoding device, S1720 can be performed by the residual processor 320, the predictor 330 and / or the adder 340 of the decoding device, and S1730 can be performed by the adder 340 of the decoding device. Figure 17 The method disclosed in
[0548] Referring to Figure 17 The decoding device can receive / obtain video / image information (S1700). The video / image information can include information on LMCS data and / or residual information. For example, the information on LMCS data related information can include information on luminance mapping (i.e., forward mapping, inverse mapping, linear mapping), information on chroma residual scaling, and / or indices (i.e., maximum bin index, minimum bin index, mapping index) related to LMCS (or reshaping, reshaper). The decoding device can receive / obtain the image / video information through a bitstream.
[0549] According to embodiments of the present document, the image / video information can include various information. For example, the image / video information can include information disclosed in at least one of Tables 1 to 60 described above.
[0550] The decoding device can generate predicted luma samples. The decoding device can derive predicted luma samples of the current block based on a prediction mode. In this case, various prediction methods disclosed in the present document, such as inter prediction or intra prediction, can be applied.
[0551] The image information can include residual information. The decoding device can generate residual chroma samples based on the residual information. Specifically, the decoding device can derive quantized transform coefficients based on the residual information. The quantized transform coefficients can have a one-dimensional vector form based on a coefficient scan order. The decoding device can derive transform coefficients based on a dequantization process on the quantized transform coefficients. The decoding device can derive residual chroma samples and / or residual luma samples based on the transform coefficients.
[0552] The decoding device can generate mapped predicted luma samples (S1710). For example, the decoding device can derive input values and mapping values (output values) of pivot points of luminance mapping, and can generate mapped predicted luma samples based on the input values and the mapping values. In one example, the decoding device can derive a (forward) mapping index (idxY) based on the first predicted luma samples, and can generate the first mapped predicted luma samples based on input values and mapping values of pivot points corresponding to the mapping index. In other examples, a linear mapping (linear reshaping, linear LMCS) can be used, and mapped predicted luma samples can be generated based on a forward mapping scaling factor derived from two pivot points in the linear mapping, so that an index derivation process can be omitted due to the linear mapping.
[0553] The decoding device can generate the residual luma samples based on the residual information (S1720). For example, the decoding device can derive quantized transform coefficients based on the residual information. The quantized transform coefficients can have a one-dimensional vector form based on a coefficient scan order. The decoding device can derive transform coefficients based on a dequantization process of the quantized transform coefficients. The decoding device can derive the residual samples based on an inverse transform process of the transform coefficients. The residual samples can include the residual luma samples and / or the residual chroma samples.
[0554] The decoding device can generate the reconstructed luma samples (S1730). The decoding device can generate the reconstructed luma samples based on the mapped prediction luma samples. Specifically, the decoding device can sum the residual luma samples with the mapped prediction luma samples, and can generate the reconstructed luma samples based on a result of the summing.
[0555] The decoding device can generate modified reconstructed luma samples based on the reconstructed luma samples and the information about the LMCS data (S1740). The decoding device can generate the modified reconstructed luma samples by a reverse mapping process of the recovered luma samples.
[0556] The decoding device can generate the scaled residual chroma samples. Specifically, the decoding device can derive chroma residual scaling factors and generate the scaled residual chroma samples based on the chroma residual scaling factors. Here, in contrast to the encoding side, the chroma residual scaling at the decoding side can be referred to as reverse chroma residual scaling. Thus, the chroma residual scaling factors derived by the decoding device can be referred to as reverse chroma residual scaling factors, and the reverse scaled residual chroma samples can be generated.
[0557] The decoding device can generate the reconstructed chroma samples. The decoding device can generate the reconstructed chroma samples based on the scaled residual chroma samples. Specifically, the decoding device can perform a prediction process for the chroma components and can generate predicted chroma samples. The decoding device can generate the reconstructed chroma samples based on a sum of the predicted chroma samples and the scaled residual chroma samples.
[0558] In an embodiment, the picture information can include an LMCS APS. The LMCS APS (one of the LMCS APSs) can include type information indicating that the LMCS APS is an APS including an LMCS data field and identifier information (ID information) of the LMCS APS. The LMCS data field can include information on LMCS data. The LMCS APS can include the LMCS data field based on the type information, and can derive an LMCS codeword based on the LMCS data field. The mapped prediction luma samples and the modified reconstructed luma samples can be generated based on the LMCS codeword. In one example, a value of the identifier information can be in a predetermined range. For example, the predetermined range can be in a range of 0 to 3. That is, the value of the identifier information can be 0, 1, 2, or 3. A maximum number of the LMCS APSs can be a predetermined value. In another example, the maximum number of the LMCS APSs can be a predetermined value. For example, the maximum number (the predetermined value) of the LMCS APSs can be four.
[0559] In an embodiment, a segment index (e.g., idxYInv in Table 35, Table 36, or Table 37) can be identified based on the information on the LMCS data. The decoding device can derive a chroma residual scaling factor based on the segment index. The decoding device can generate scaled residual chroma samples based on the residual chroma samples and the chroma residual scaling factor.
[0560] In one embodiment, when the current block has a single tree structure or a dual tree structure (when the current block has a separate tree structure, when the current block is coded as a separate tree), a chroma residual scaling enabled flag indicating whether chroma residual scaling is applied for the current block can be signaled. When the chroma residual scaling is applied for the current picture, the current slice, and / or the current block, a value of the chroma residual scaling enabled flag can be 1.
[0561] In an embodiment, the chroma residual scaling factor can be a single chroma residual scaling factor.
[0562] In an embodiment, based on a value of the type information being 1, the APS can contain an LMCS data field including LMCS parameters.
[0563] In an embodiment, the picture information can include header information. Here, the header information can be a picture header (or a slice header). The header information can include LMCS-related APS ID information. The LMCS-related APS ID information can indicate the identifier information of the LMCS APS for the current picture, the current block, or the current slice. That is, a value of the LMCS-related APS ID information can be the same as a value of the identifier information of the LMCS APS. For example, the value of the LMCS-related APS ID information can be in a range of 0 to 3. That is, the value of the LMCS-related APS ID information can be 0, 1, 2, or 3.
[0564] In an embodiment, a minimum bin index (e.g., lmcs_min_bin_idx) and / or a maximum bin index (e.g., LmcsMaxBinIdx) can be derived based on the information on the LMCS data. A first mapping value (LmcsPivot[lmcs_min_bin_idx]) can be derived based on the minimum bin index. A second mapping value (LmcsPivot[LmcsMaxBinIdx] or LmcsPivot[LmcsMaxBinIdx+1]) can be derived based on the maximum bin index. Values of the reconstructed luma samples (e.g., lumaSample of Table 51 or Table 52) can be within a range of the first mapping value to the second mapping value. In one example, values of all of the reconstructed luma samples can be within the range of the first mapping value to the second mapping value. In another example, values of some of the reconstructed luma samples can be within the range of the first mapping value to the second mapping value.
[0565] In an embodiment, the picture information can include a sequence parameter set (SPS). The SPS can include a linear LMCS enabled flag indicating whether a linear LMCS is enabled.
[0566] In an embodiment, the chroma residual scaling factor can be a single chroma residual scaling factor.
[0567] In an embodiment, the information on the LMCS data can include information on a linear LMCS and an LMCS data field. The information on the linear LMCS can be referred to as information on a linear mapping. The LMCS data field can include a linear LMCS flag indicating whether a linear LMCS is applied. When a value of the linear LMCS flag is 1, predicted luma samples mapped based on the information on the linear LMCS can be generated.
[0568] In an embodiment, the information on the linear LMCS can include information on a first pivot point (e.g., P1 in Table 34) and information on a second pivot point (e.g., P2 in Table 34). For example, an input value and a mapped value of the first pivot point can be a minimum input value and a minimum mapped value, respectively. An input value and a mapped value of the second pivot point can be a maximum input value and a maximum mapped value, respectively. Input values between the minimum input value and the maximum input value can be linearly mapped. Figure 11 Figure 11 In an embodiment, the information on the linear LMCS can include information on a first pivot point (e.g., P1 in Table 34) and information on a second pivot point (e.g., P2 in Table 34). For example, an input value and a mapped value of the first pivot point can be a minimum input value and a minimum mapped value, respectively. An input value and a mapped value of the second pivot point can be a maximum input value and a maximum mapped value, respectively. Input values between the minimum input value and the maximum input value can be linearly mapped.
[0569] In an embodiment, the picture information can include information on a maximum input value and information on a maximum mapped value. The maximum input value can be the same as a value of the information on the maximum input value (e.g., lmcs_max_input in Table 33). The maximum mapped value can be the same as a value of the information on the maximum mapped value (e.g., lmcs_max_mapped in Table 33).
[0570] In an embodiment, the information on the linear mapping can include information on an input delta value of the second pivot point (e.g., lmcs_max_input_delta of Table 35) and information on a mapped delta value of the second pivot point (e.g., lmcs_max_mapped_delta of Table 35). The maximum input value can be derived based on the input delta value of the second pivot point, and the maximum mapped value can be derived based on the mapped delta value of the second pivot point.
[0571] In an embodiment, the maximum input value and the maximum mapped value can be derived based on at least one formula included in Table 36 above.
[0572] In an embodiment, generating the mapped prediction luma sample includes deriving a forward mapping scaling factor (i.e., ScaleCoeffSingle) of the prediction luma sample, and generating the mapped prediction luma sample based on the forward mapping scaling factor. The forward mapping scaling factor can be a single factor of the prediction luma sample.
[0573] In an embodiment, the inverse mapping scaling factor can be derived using the segment index derived based on the reconstructed luma sample.
[0574] In an embodiment, the segment index can be derived based on Table 51 above. That is, the comparison process (lumaSample < LmcsPivot[idxYInv + 1]) included in Table 51 can be iteratively performed from the segment index that is the minimum bin index to the segment index that is the maximum bin index.
[0575] In an embodiment, the forward mapping scaling factor can be derived based on at least one formula included in Table 36 and / or Table 38 above.
[0576] In an embodiment, the mapped prediction luma sample can be derived based on at least one formula included in Table 38 above.
[0577] In an embodiment, the decoding device can derive an inverse mapping scaling factor (i.e., InvScaleCoeffSingle) of the reconstructed luma sample (i.e., lumaSample). In addition, the decoding device can generate a modified reconstructed luma sample (i.e., invSample) based on the reconstructed luma sample and the inverse mapping scaling factor. The inverse mapping scaling factor can be a single factor of the reconstructed luma sample.
[0578] In an embodiment, the inverse mapping scaling factor can be derived based on at least one formula included in Table 33, Table 34, Table 35, and Table 36 or Formula 11 or Formula 12 above.
[0579] In one embodiment, the modified reconstructed luma samples can be derived based on the above equation 20, equation 21, table 39 and / or table 40.
[0580] In one embodiment, the LMCS related information can include information on a number of bins used to derive the mapped prediction luma samples (i.e., lmcs num bins minusl in table 41). For example, the number of pivot points for luma mapping can be set equal to the number of bins. In one example, a decoding device can generate the delta input values and delta mapped values of the pivot points according to the number of bins, respectively. In one example, the input values and mapped values of the pivot points are derived based on the delta input values (i.e., lmcs delta input cw[i] in table 41) and the delta mapped values (i.e., lmcs delta mapped cw[i] in table 41), and the mapped prediction luma samples can be generated based on the input values (i.e., LmcsPivot_input[i] of table 42) and the mapped values (i.e., LmcsPivot_mapped[i] of table 42).
[0581] In one embodiment, a decoding device can derive the LMCS delta codewords based on the at least one LMCS codeword and an original codeword (OrgCW) included in the LMCS related information, and can derive the mapped luma prediction samples based on the at least one LMCS codeword and the original codeword. In one example, the information on the linear mapping can include information on the LMCS delta codewords.
[0582] In one embodiment, the at least one LMCS codeword can be derived based on a sum of the LMCS delta codeword and the OrgCW, e.g., OrgCW can be (1 « BitDepthY) / 16, where BitDepthY represents the luma bit depth. This embodiment can be based on equation 14.
[0583] In one embodiment, the at least one LMCS codeword can be derived based on a sum of the LMCS delta codeword and OrgCW*(lmcs max bin idx-lmcs min bin idx+1), e.g., lmcs max bin idx and lmcs min bin idx can be the maximum bin index and the minimum bin index, respectively, and OrgCW can be (1 « BitDepthY) / 16. This embodiment can be based on equation 15 and equation 16.
[0584] In one embodiment, the at least one LMCS codeword can be a multiple of 2.
[0585] In one embodiment, when the luma bit depth of the reconstructed luma sample (BitDepthY) is higher than 10, the at least one LMCS codeword can be a multiple of 1<<(BitDepthY-10).
[0586] In one embodiment, the at least one LMCS codeword can be in a range from (OrgCW>>1) to (OrgCW<<1)-1.
[0587] In the above paragraph, the information on the LMCS data can be the same as the information on the LMCS.
[0588] In the above embodiments, the method is described based on a flowchart having a series of steps or blocks. The disclosure is not limited to the order of the above steps or blocks. Some steps or blocks can occur at the same time or in a different order from other steps or blocks as described above. Furthermore, those skilled in the art will understand that the steps shown in the above flowchart are not exclusive, and additional steps can be included or one or more steps of the flowchart can be deleted without affecting the scope of the disclosure.
[0589] The method according to the above embodiments of the present document can be implemented in software form, and the encoding apparatus and / or the decoding apparatus according to the present document can be included, for example, in an apparatus that performs image processing of a TV, a computer, a smart phone, a set-top box, a display apparatus, etc.
[0590] When the embodiments in the present document are implemented in software, the above method can be implemented as a module (process, function, etc.) that performs the above functions. The module can be stored in a memory and executed by a processor. The memory can be located inside or outside the processor, and can be coupled to the processor by various well-known means. The processor can include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory can include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in the present document can be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in the respective drawings can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information on instructions or algorithms for implementation can be stored in a digital storage medium.
[0591] In addition, the decoding apparatus and the encoding apparatus to which the present disclosure is applied can be included in a multimedia broadcast transmitting / receiving apparatus, a mobile communication terminal, a home theater video apparatus, a digital theater video apparatus, a surveillance camera, a video chat apparatus, a real-time communication apparatus (e.g., video communication), a mobile streaming apparatus, a storage medium, a camcorder, a VoD service providing apparatus, an over-the-top (OTT) video apparatus, an Internet streaming service providing apparatus, a three-dimensional (3D) video apparatus, a teleconference video apparatus, a transportation user apparatus (i.e., a vehicle user apparatus, an airplane user apparatus, a ship user apparatus, etc.), and a medical video apparatus, and can be used to process a video signal and a data signal. For example, the over-the-top (OTT) video apparatus can include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.
[0592] In addition, the processing method to which the present disclosure is applied can be generated in the form of a program to be executed by a computer, and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the present disclosure can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices storing data readable by a computer system. For example, the computer-readable recording medium can include a BD, a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (i.e., transmission through the Internet). In addition, a bitstream generated by the encoding method can be stored in a computer-readable recording medium or can be transmitted via a wired / wireless communication network.
[0593] In addition, the embodiments of the present disclosure can be implemented as a computer program product according to program codes, and the program codes can be executed in a computer by the embodiments of the present disclosure. The program codes can be stored on a carrier readable by a computer.
[0594] Figure 19 An example of a content streaming system to which the embodiments disclosed in the present disclosure can be applied is shown.
[0595] Referring to Figure 19 The content streaming system to which the embodiments of the present disclosure are applied can mainly include an encoding server, a streaming server, a web server, a media storage device, a user device, and a multimedia input device.
[0596] The encoding server compresses content input from the multimedia input device (e.g., a smart phone, a camera, a camcorder, etc.) into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, when the multimedia input device (e.g., a smart phone, a camera, a camcorder, etc.) directly generates a bitstream, the encoding server can be omitted.
[0597] A bitstream can be generated by applying an encoding method or a bitstream generation method to which embodiments of the disclosure are applied, and a streaming server can temporarily store the bitstream in a process of transmitting or receiving the bitstream.
[0598] A streaming server transmits multimedia data to a user device based on a request of a user through a web server, and the web server serves as a medium that informs the user of a service. When the user requests a desired service to the web server, the web server transmits it to the streaming server, and the streaming server transmits multimedia data to the user. In this case, the content streaming system can include a separate control server. In this case, the control server is used to control commands / responses between devices in the content streaming system.
[0599] The streaming server can receive content from a media storage device and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store a bitstream for a predetermined time.
[0600] Examples of the user device can include a mobile phone, a smart phone, a laptop computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smart watch, a smart glass, a head-mounted display), a digital TV, a desktop computer, a digital signage, etc. The respective servers in the content streaming system can operate as a distributed server, in which case data received from the respective servers can be distributed.
[0601] The respective servers in the content streaming system can operate as a distributed server, and in this case, data received from the respective servers can be distributed and processed.
[0602] The claims described herein can be combined in various ways. For example, the technical features of the method claims of the present document can be combined and implemented as an apparatus, and the technical features of the apparatus claims of the present document can be combined and implemented as a method. In addition, the technical features of the method claims and the technical features of the apparatus claims of the present document can be combined to be implemented as an apparatus, and the technical features of the method claims and the technical features of the apparatus claims of the present document can be combined and implemented as a method.
Claims
1. An image decoding method performed by a decoding device, the image decoding method comprising the steps of: obtaining, from a bitstream, picture information including prediction mode information, residual information, and information on luminance mapping and chroma scaling (LMCS) data; deriving, based on the prediction mode information, an inter prediction mode for a current block; deriving, based on the inter prediction mode, motion information for the current block; generating, based on the motion information for the current block, a predicted luma sample for the current block; generating, based on the information on the LMCS data and the predicted luma sample, a mapped predicted luma sample for the current block; deriving, based on the residual information, a residual luma sample for the current block; generating, based on the mapped predicted luma sample and the residual luma sample, a reconstructed luma sample for the current block; and generating, based on the reconstructed luma sample and the information on the LMCS data, a modified reconstructed luma sample for the current block, wherein the picture information includes a LMCS adaptation parameter set (APS); wherein the LMCS APS includes identifier information of the LMCS APS and type information indicating that the LMCS APS is an APS including a LMCS data field; wherein the LMCS data field includes the information on the LMCS data; wherein the LMCS APS includes the LMCS data field based on the type information; wherein a LMCS codeword is derived based on the LMCS data field; wherein the mapped predicted luma sample and the modified reconstructed luma sample are generated based on the LMCS codeword; wherein a value of the identifier information is within a predetermined range; wherein the picture information includes header information; wherein the header information includes LMCS related APS ID information; wherein the LMCS related APS ID information represents the identifier information of the LMCS APS; and wherein a value of the LMCS related APS ID information is within a range of 0 to 3.
2. An image encoding method performed by an encoding device, the image encoding method comprising the steps of: deriving an inter prediction mode for a current block; generating, based on the inter prediction mode, motion information for the current block; generating, based on the motion information for the current block, a predicted luma sample for the current block; deriving a luminance mapping and chroma scaling (LMCS) codeword and bins for luminance mapping; generating, based on the LMCS codeword and the bins for luminance mapping, a mapped predicted luma sample for the current block; generating, based on the mapped predicted luma sample, a reconstructed luma sample for the current block; generating, based on the reconstructed luma sample, the LMCS codeword, and the bins for luminance mapping, a modified reconstructed luma sample for the current block; deriving, based on the LMCS codeword and the bins for luminance mapping, information on LMCS data; generating, based on the mapped predicted luma sample, a residual luma sample for the current block; generating, based on the residual luma sample, residual information; and encoding image information including the residual information and the information on the LMCS data, wherein the image information includes an LMCS adaptation parameter set (APS); wherein the LMCS APS includes identifier information of the LMCS APS and type information indicating that the LMCS APS is an APS including an LMCS data field; wherein the LMCS data field includes the information on the LMCS data; wherein the LMCS APS includes the LMCS data field based on the type information; wherein the LMCS codeword is derived based on the LMCS data field; wherein a value of the identifier information is within a predetermined range; wherein the image information includes header information; wherein the header information includes LMCS-related APS ID information; wherein the LMCS-related APS ID information represents the identifier information of the LMCS APS; and wherein a value of the LMCS-related APS ID information is within a range of 0 to 3.
3. A computer-readable storage medium storing a bitstream generated by an image encoding method performed by an encoding apparatus according to claim 2.
4. A method for transmission of data of an image, the method comprising the steps of: obtaining a bitstream for the image, wherein the bitstream is generated based on operations of deriving an inter prediction mode for a current block, generating motion information for the current block based on the inter prediction mode, generating predicted luma samples for the current block based on the motion information for the current block, deriving a luma mapping and chroma scaling (LMCS) codeword and bin for luma mapping, generating mapped predicted luma samples for the current block based on the LMCS codeword and the bin for luma mapping, generating reconstructed luma samples for the current block based on the mapped predicted luma samples, generating modified reconstructed luma samples for the current block based on the reconstructed luma samples, the LMCS codeword and the bin for luma mapping, deriving information on LMCS data based on the LMCS codeword and the bin for luma mapping, generating residual luma samples for the current block based on the mapped predicted luma samples, generating residual information based on the residual luma samples, and encoding image information including the residual information and the information on the LMCS data; and transmitting the data including the bitstream, wherein the image information includes an LMCS adaptation parameter set (APS); wherein the LMCS APS includes identifier information of the LMCS APS and type information indicating that the LMCS APS is an APS including an LMCS data field; wherein the LMCS data field includes the information on the LMCS data; wherein the LMCS APS includes the LMCS data field based on the type information; wherein the LMCS codeword is derived based on the LMCS data field; wherein a value of the identifier information is in a predetermined range; wherein the image information includes header information; wherein the header information includes LMCS-related APS ID information; wherein the LMCS-related APS ID information indicates the identifier information of the LMCS APS; and wherein a value of the LMCS-related APS ID information is in a range of 0 to 3.
Citation Information
Patent Citations
Systems and methods for coding video data using adaptive component scaling
CN109479133A
Largest magnitude indices selection for (run, level) encoding of a block coded picture
US20030118243A1