Video or image encoding method and apparatus therefor
By employing adaptive filtering and luminance mapping with chroma scaling techniques, the transmission and storage costs of high-resolution images/videos are addressed, improving compression efficiency and visual quality. This technology is suitable for image/video encoding in virtual reality, artificial reality, and immersive media.
Patent Information
- Application Number
- CN202410824124.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-03
- Filing Date
- 2020-03-19
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2040-03-19
AI Technical Summary
With the increasing demand for high-resolution, high-quality images/videos, existing technologies face challenges in terms of transmission and storage costs, and there is a lack of effective image/video compression technologies to handle the differences in characteristics of virtual reality, artificial reality, and immersive media.
By employing Adaptive Filtering Application (ALF) and Luminance Mapping and Chromatography Scaling (LMCS) techniques, relevant information is communicated through signals, filter coefficients are adaptively applied, and filters are derived within images or slices. Combined with entropy coding and entropy decoding, image/video compression efficiency is improved.
It improves image/video compression efficiency and subjective/objective visual quality, enabling more efficient image/video transmission and storage.
Smart Images

Figure CN118612428B_ABST
Abstract
Description
[0001] This application is a divisional application of patent application No. 202080031817.9 (PCT / KR2020 / 003787), filed on October 27, 2021, with an international application date of March 19, 2020, entitled "Video or Image Compilation Method and Apparatus Thereof". Technical Field
[0002] This disclosure relates to video or image encoding methods and apparatus. Background Technology
[0003] Recently, there has been a growing demand for high-resolution, high-quality images / videos, such as 4K or 8K Ultra High Definition (UHD) images / videos, across various fields. As image / video resolution or quality increases, relatively more information or bits are transmitted compared to traditional image / video data. Therefore, if image / video data is transmitted via media such as existing wired / wireless broadband lines or stored in traditional storage media, the costs of transmission and storage can easily increase.
[0004] In addition, there is growing interest and demand for virtual reality (VR) and artificial reality (AR) content, as well as immersive media such as holograms; and the broadcasting of images / videos that exhibit characteristics different from actual images / videos (e.g., game images / videos) is also increasing.
[0005] Therefore, efficient image / video compression technology is needed to effectively compress and send, store, or play high-resolution, high-quality images / videos that exhibit the various characteristics described above.
[0006] In addition, there is discussion about techniques such as Luminance Mapping and Chroma Scaling (LMCS) and Adaptive Loop Filtering (ALF) to improve compression efficiency and increase subjective / objective visual quality. To effectively apply these techniques, a method is needed to efficiently communicate relevant information using signals. Summary of the Invention
[0007] Technical solution
[0008] According to embodiments of this document, a method and apparatus for increasing image coding efficiency are provided.
[0009] According to the embodiments in this document, an effective filtering application method and device are provided.
[0010] According to the embodiments in this document, an effective LMCS application method and device are provided.
[0011] According to embodiments of this document, a method and apparatus are provided for adaptively / hierarchically signaling ALF-related information.
[0012] According to embodiments of this document, a method and apparatus are provided for adaptively / hierarchically signaling relevant information about LMCS.
[0013] According to the embodiments in this document, ALF data fields can be signaled by the APS, and APS ID information indicating the ID associated with the referenced APS can be signaled by header information (image header or slice header).
[0014] According to the embodiments in this document, information about the number of APS IDs associated with the ALF can be signaled via header information. In this case, the header information may include as many APS ID syntax elements as there are APS IDs associated with the ALF, and based on this, it is possible to derive adaptive filters (filter coefficients) and apply the ALF on a block / sub-block basis within the same image or slice.
[0015] According to the embodiments in this document, LMCS data fields can be signaled by the APS, and APS ID information indicating the ID of the referenced APS can be signaled by header information (image header or slice header).
[0016] According to the embodiments in this document, the APS can signal the APS type information, and the type information can indicate whether the corresponding APS is an APS that includes an ALF data field (or ALF parameter) or whether the corresponding APS includes an LMCS data field (or LMCS parameter).
[0017] According to embodiments of this document, a video / image decoding method performed by a decoding device is provided. The video / image decoding method may include the methods disclosed in the embodiments of this document.
[0018] According to embodiments of this document, a decoding device is provided for performing video / image decoding. The decoding device may include the methods disclosed in the embodiments of this document.
[0019] According to embodiments of this document, a video / image encoding method performed by an encoding device is provided. The video / image encoding method may include the methods disclosed in the embodiments herein.
[0020] According to embodiments of this document, an encoding apparatus for performing video / image encoding is provided. The encoding apparatus may include the methods disclosed in the embodiments of this document.
[0021] According to embodiments of this document, a computer-readable digital storage medium is provided that stores encoded video / image information generated according to a video / image encoding method disclosed in at least one embodiment of this document.
[0022] According to embodiments of this document, a computer-readable digital storage medium is provided that stores encoded information or encoded video / image information, said encoded information or encoded video / image information causing a decoding device to perform a video / image decoding method disclosed in at least one embodiment of this document.
[0023] Invention Effects
[0024] According to this disclosure, the overall image / video compression efficiency can be improved.
[0025] According to this disclosure, subjective / objective visual quality can be improved through effective filtering.
[0026] According to this disclosure, compression performance can be improved through LMCS.
[0027] According to this disclosure, ALF and / or LMCS can be adaptively applied on a per-picture, per-slice, and / or per-coded-block basis.
[0028] According to this disclosure, ALF parameters can be effectively communicated using signals.
[0029] According to this disclosure, LMCS parameters can be effectively communicated using signals. Attached Figure Description
[0030] Figure 1 Examples of video / image coding systems to which embodiments of the present disclosure may be applied are illustrated schematically.
[0031] Figure 2 This is a diagram schematically illustrating the configuration of a video / image encoding apparatus to which embodiments of the present disclosure may be applied.
[0032] Figure 3 This is a diagram schematically illustrating the configuration of a video / image decoding device to which embodiments of the present disclosure may be applied.
[0033] Figure 4 An example of a video / image coding method based on intra-frame prediction is shown.
[0034] Figure 5 An example of a video / image decoding method based on intra-frame prediction is shown.
[0035] Figure 6 An example of the intra-frame prediction process is shown.
[0036] Figure 7 An example of a video / image coding method based on inter-frame prediction is shown.
[0037] Figure 8 An example of a video / image decoding method based on inter-frame prediction is shown.
[0038] Figure 9 An example of the inter-frame prediction process is shown.
[0039] Figure 10 An example of a layered structure for encoded images / videos is shown.
[0040] Figure 11 This is a flowchart that schematically illustrates an example of an ALF process.
[0041] Figure 12 An example of the shape of an ALF filter is shown.
[0042] Figure 13 An example of a hierarchical structure for ALF data is shown.
[0043] Figure 14 Another example of the hierarchical structure of ALF data is shown.
[0044] Figure 15 The hierarchical structure of CVS according to an embodiment of this document is illustrated by way of example.
[0045] Figure 16 An exemplary LMCS structure is shown according to an embodiment of this document.
[0046] Figure 17 An exemplary LMCS structure according to another embodiment of this document is shown.
[0047] Figure 18 A diagram illustrating an exemplary forward mapping is shown.
[0048] Figure 19 and Figure 20 Examples of video / image encoding methods and related components according to embodiments of this document are illustrated schematically.
[0049] Figure 21 and Figure 22 Examples of video / image decoding methods and related components according to embodiments of this document are illustrated schematically.
[0050] Figure 23 An illustrative diagram illustrating the structure of the content flow system applying this disclosure. Detailed Implementation
[0051] This document's disclosure may be modified in various forms, and specific embodiments thereof will be described and illustrated in the accompanying drawings. However, the embodiments are not intended to limit this disclosure. The terminology used in this document is for describing particular embodiments only and is not intended to limit the embodiments of this document. Singular expressions include plural expressions, provided they are not clearly understood differently. Terms such as "comprising" and "having" are intended to indicate the presence of features, quantities, steps, operations, elements, components, or combinations thereof used in the document, and therefore should be understood that the possibility of having or adding one or more different features, quantities, steps, operations, elements, components, or combinations thereof is not excluded.
[0052] Furthermore, each configuration in the accompanying drawings described in this document is a separate illustration for explaining the functionality as distinct features, and does not imply that each configuration is implemented by different hardware or different software. For example, two or more configurations may be combined to form one configuration, and one configuration may be divided into multiple configurations. Embodiments of combined and / or separated configurations are included within the scope of this disclosure without departing from the spirit of the present disclosure.
[0053] This document relates to video / image coding. For example, the methods / exercises disclosed in this document can be applied to methods disclosed in the Universal Video Coding (VVC) standard. Furthermore, the methods / exercises disclosed in this document can be applied to methods disclosed in the Basic Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the Audio Video Coding 2 (AVS2) standard, or next-generation video / image coding standards (e.g., H.267, H.268, etc.).
[0054] This disclosure presents various embodiments of video / image coding, and unless otherwise specified, the above embodiments may also be performed in combination with each other.
[0055] In this disclosure, video can refer to a series of images over time. An image generally refers to a unit representing an image at a specific time frame, and a slice / tile refers to a unit that constitutes part of an image in terms of coding. A slice / tile may include one or more Coding Tree Units (CTUs). An image may consist of one or more slices / tiles. A tile is a rectangular region of a CTU within a specific tile column and row in an image. A tile column is a rectangular region of a CTU having a height equal to the height of the image and a width that can be specified by a syntax element in the image parameter set. A tile row is a rectangular region of a CTU having a width specified by a syntax element in the image parameter set and a height equal to the height of the image. A tile scan can represent a specific ordering of CTUs partitioning an image, where CTUs are sequentially ordered within a tile using CTU raster scans, and tiles within an image are sequentially ordered using the raster scans of the image's tiles (a tile scan is a specific ordering of CTUs partitioning an image, where CTUs are sequentially ordered within a tile using CTU raster scans, and tiles within an image are sequentially ordered using the raster scans of the image's tiles). A slice comprises an integer number of complete tiles or an integer number of consecutive complete CTU rows that can be exclusively contained within a single NAL unit.
[0056] Furthermore, an image can be divided into two or more sub-images. A sub-image can be a rectangular region of one or more slices within the image.
[0057] A pixel, or pel, can refer to the smallest unit that makes up a picture (or image). Additionally, the term "sample" can be used as the counterpart to a pixel. A sample can typically represent a pixel or a pixel value, and can represent pixel / pixel values for either the luminance component only or the chrominance component only.
[0058] A unit can represent a basic unit of image processing. A unit may include a specific region of an image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an M×N block may include a set (or array) of samples (or transform coefficients) with M columns and N rows.
[0059] In this document, “A or B” may mean “A only,” “B only,” or “both A and B.” In other words, “A or B” in this document can be interpreted as “A and / or B.” For example, in this document, “A, B, or C” means “A only,” “B only,” “C only,” or “any combination of A, B, and C.”
[0060] The forward slash ( / ) or comma (,) used in this document can mean "and / or". For example, "A / B" can mean "A and / or B". Therefore, "A / B" can mean "A only", "B only", or "both A and B". For example, "A, B, C" can mean "A, B, or C".
[0061] In this document, "at least one of A and B" may mean "A only", "B only", or "both A and B". Furthermore, in this document, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted as the same as "at least one of A and B".
[0062] Furthermore, in this document, "at least one of A, B, and C" means "A only", "B only", "C only", or "any combination of A, B, and C". Additionally, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".
[0063] Furthermore, the parentheses used in this document may mean "for example". Specifically, when indicating "prediction (intra-frame prediction)", "intra-frame prediction" may be cited as an example of "prediction". In other words, "prediction" in this document is not limited to "intra-frame prediction", and "intra-frame prediction" may be cited as an example of "prediction". Moreover, even when indicating "prediction (i.e., intra-frame prediction)", "intra-frame prediction" may be cited as an example of "prediction".
[0064] The technical features described individually in one of the accompanying figures in this document may be implemented individually or simultaneously.
[0065] In the following, embodiments of this document will be described in detail with reference to the accompanying drawings. Furthermore, throughout the drawings, the same reference numerals will be used to indicate the same elements, and the same descriptions of the same elements will be omitted.
[0066] Figure 1 Examples of video / image encoding systems to which embodiments of this document can be applied are shown.
[0067] Reference Figure 1 A video / image encoding system may include a source device and a receiving device. The source device may transmit encoded video / image information or data to the receiving device in the form of a file or stream via a digital storage medium or network.
[0068] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0069] Video sources can acquire video / images through processes that capture, synthesize, or generate video / images. Video sources may include video / image capture devices and / or video / image generation devices. For example, a video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. For example, a video / image generation device may include a computer, tablet computer, and smartphone, and may generate video / images (electronically). For example, virtual video / images may be generated via a computer, etc. In this case, the video / image capture process may be replaced by a process that generates related data.
[0070] Encoding devices can encode input video / images. For compression and encoding efficiency, encoding devices can perform a series of processes such as prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output as a bitstream.
[0071] The transmitter can send encoded images / image information or data, output as a bitstream, to the receiver of the receiving device in the form of a file or stream via a digital storage medium or network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating media files according to a predetermined file format and may include elements for transmission over a broadcast / communication network. The receiver can receive / extract the bitstream and send the received bitstream to a decoding device.
[0072] Decoding devices can decode video / images by performing a series of processes such as dequantization, inverse transform, and prediction, which correspond to the operations of encoding devices.
[0073] The renderer can render decoded video / images. The rendered video / images can be displayed on a monitor.
[0074] Figure 2 This diagram schematically illustrates the configuration of a video / image encoding apparatus to which embodiments of the present disclosure may be applied. Hereinafter, the term "encoding apparatus" may include image encoding apparatus and / or video encoding apparatus. Furthermore, an image encoding method / apparatus may include a video encoding method / apparatus. Alternatively, a video encoding method / apparatus may include an image encoding method / apparatus.
[0075] Reference Figure 2 The encoding device 200 includes an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transform 232, a quantizer 233, a dequantizer 234, and an inverse transform 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image segmenter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). Additionally, the memory 270 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware component may also include the memory 270 as an internal / external component.
[0076] Image segmenter 210 can segment an input image (or picture, frame) input to encoding device 200 into one or more processing units. As an example, a processing unit may be referred to as a coding unit (CU). In this case, coding units can be recursively segmented from coding tree units (CTUs) or maximum coding units (LCUs) according to a quadtree-binary-tritree (QTBTTT) structure. For example, a coding unit can be segmented into multiple deeper coding units based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, a quadtree structure is applied first, and a binary tree structure and / or a ternary tree structure can be applied later. Alternatively, a binary tree structure can be applied first. The encoding process according to this document can be performed based on the final coding units that are no longer segmented. In this case, based on encoding efficiency according to image characteristics, etc., the maximum coding unit can be directly used as the final coding unit, or, as needed, the coding unit can be recursively segmented into deeper coding units, so that coding units of optimal size can be used as the final coding units. Here, the encoding process may include processes such as prediction, transformation, and reconstruction, as described later. As another example, the processing unit may also include a prediction unit (PU) or a transformation unit (TU). In this case, each of the prediction and transformation units may be segmented or partitioned from the final encoding unit described above. The prediction unit may be a unit for predicting samples, and the transformation unit may be a unit for deriving transform coefficients and / or a unit for deriving the residual signal from the transform coefficients.
[0077] In some cases, a unit can be used interchangeably with terms such as a block or region. Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can typically represent a pixel or pixel value, either representing only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component. A sample can be used as a term corresponding to a picture (or image) of pixels or cells.
[0078] In the encoding device 200, a residual signal (residual block, residual sample array) is generated by subtracting the prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 from the input image signal (original block, original sample array), and the generated residual signal is sent to the converter 232. In this case, as shown, the unit in the encoder 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) may be called the subtractor 231. The predictor can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block including the prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. As described later in the description of the various prediction modes, the predictor can generate various types of information related to the prediction (e.g., prediction mode information) and send the generated information to the entropy encoder 240. The information about the prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0079] Intra-predictor 222 can refer to samples in the current image to predict the current block. Depending on the prediction mode, the referenced samples may be located near or separated from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. For example, non-directional modes can include DC mode and planar mode. For example, depending on the level of detail in the prediction direction, the directional modes can include 33 or 65 directional prediction modes. However, this is just an example, and more or fewer directional prediction modes can be used depending on the settings. Intra-predictor 222 can use the prediction modes applied to neighboring blocks to determine the prediction mode applied to the current block.
[0080] Inter-frame predictor 221 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. Here, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a juxtaposed reference block, juxtaposed CU (colCU), etc., and the reference image including the temporally neighboring block may be referred to as a juxtaposed image (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to deduce the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 221 can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be sent. Motion Vector Prediction (MVP) mode indicates the motion vector of the current block by using motion vectors from neighboring blocks as motion vector predictors and signaling the motion vector difference.
[0081] Predictor 200 can generate prediction signals based on various prediction methods described later. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but can also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined inter-frame and intra-frame prediction (CIIP). Furthermore, the predictor can perform prediction on blocks based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as screen content coding (SCC) in games, etc. IBC essentially performs prediction in the current frame, but it can be performed similarly to inter-frame prediction because it derives a reference block in the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this document. The palette mode can be considered as an example of intra-frame coding or intra-frame prediction. When a palette mode is applied, sample values in the image can be signaled based on information about the palette table and palette index.
[0082] The predicted signal generated by the predictor (including inter-frame predictor 221 and / or intra-frame predictor 222) can be used to generate a reconstructed signal or a residual signal. Transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of the following: Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Graphical Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, when the relationship information between pixels is illustrated as a graph, GBT refers to a transform obtained from a graph. CNT refers to a transform obtained based on the predicted signal generated using all previously reconstructed pixels. Furthermore, the transform process can be applied to pixel blocks with squares of the same size, and can also be applied to blocks of variable sizes other than squares.
[0083] Quantizer 233 quantizes the transform coefficients to send the quantized transform coefficients to entropy encoder 240, which encodes the quantized signal (information about the quantized transform coefficients) into a bitstream. This information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the quantized transform coefficients in block form as a one-dimensional vector based on the coefficient scan order, and also generates information about the quantized transform coefficients based on the one-dimensional vector form. Entropy encoder 240 can perform various encoding methods such as exponential Golomb coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). Entropy encoder 240 can also encode, either together or separately, information necessary for reconstructing the video / image (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / image information) can be sent or stored in units of Network Abstraction Layer (NAL) in bitstream form. The video / image information may also include information about various parameter sets, such as adaptation parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). Additionally, the video / image information may also include general constraint information. The information and / or syntax elements notified / transmitted by signals, as described later in this disclosure, can be encoded by the aforementioned encoding process and thus included in the bitstream. The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting the signal output from the entropy encoder 240 and / or a storage unit (not shown) for storing the signal may be configured as internal / external components of the encoding device 200, or the transmitter may also be included in the entropy encoder 240.
[0084] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, dequantizer 234 and inverse transform 235 apply dequantization and inverse transform to the quantized transform coefficients, making it possible to reconstruct the residual signal (residual block or residual sample). Adder 250 adds the reconstructed residual signal to the prediction signal output from inter-frame predictor 221 or intra-frame predictor 222, thereby generating a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). If no residual exists in the block to be processed, such as when a skip mode is applied, the prediction block can be used as a reconstructed block. Adder 250 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed within the current image, and, as described later, is also used for inter-frame prediction of the next image through filtering.
[0085] Additionally, Luminance Mapping and Chromatography Scaling (LMCS) can be applied during image encoding and / or reconstruction.
[0086] Filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 270 (specifically, the DPB of memory 270). Various filtering methods may include deblocking filtering, sample adaptive shifting, adaptive loop filtering, bilateral filtering, etc. Filter 260 can generate various types of filtering-related information and send the generated information to entropy encoder 240, as described later in the description of the various filtering methods. The filtering-related information can be encoded by entropy encoder 240 and output as a bitstream.
[0087] The modified reconstructed image sent to memory 270 can be used as a reference image in inter-frame predictor 221. When inter-frame prediction is applied by the encoding device, prediction mismatch between the encoding device 200 and the decoding device can be avoided and encoding efficiency can be improved.
[0088] The DPB of memory 270 can store a reconstructed picture modified for use as a reference picture in inter-frame predictor 221. Memory 270 can store motion information of blocks in the current picture that derive (or encode) motion information and / or motion information of already reconstructed blocks in the picture. The stored motion information can be sent to inter-frame predictor 221 and used as motion information for spatially or temporally neighboring blocks. Memory 270 can store reconstructed samples of reconstructed blocks in the current picture and can transmit the reconstructed samples to intra-frame predictor 222.
[0089] Figure 3This diagram is used to schematically explain the configuration of a video / image decoding apparatus to which embodiments of the present disclosure may be applied. Hereinafter, the term "decoding apparatus" may include an image decoding apparatus and / or a video decoding apparatus. Furthermore, an image encoding method / apparatus may include a video encoding method / apparatus. Alternatively, a video encoding method / apparatus may include an image encoding method / apparatus.
[0090] Reference Figure 3 The decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 331 and an intra-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. According to embodiments, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 may be configured by hardware components (e.g., a decoder chipset or processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0091] When the input includes a bitstream containing video / image information, the decoding device 300 can respond to... Figure 2 The encoding device shown reconstructs an image by processing video / image information. For example, the decoding device 300 can deduce units / blocks based on block segmentation information obtained from the bitstream. The decoding device 300 can perform decoding using processing units applied to the encoding device. Therefore, the processing unit for decoding can be, for example, an encoding unit, and the encoding unit can be segmented from the encoding tree unit or the maximum encoding unit according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the encoding unit. Furthermore, the reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproduction device.
[0092] Decoding device 300 can receive data in bitstream form from... Figure 2The signal output by the encoding device shown in the diagram can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can deduce the information (e.g., video / image information) required for image reconstruction (or picture reconstruction) by parsing the bitstream. The video / image information may also include information about various parameter sets, such as adaptation parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). In addition, the video / image information may also include general constraint information. The decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The information and / or syntax elements notified / received by signals, as described later in this disclosure, can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 310 can decode the information within the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, and output the values of the syntax elements required for image reconstruction and the quantized values of the residual correlation transform coefficients. More specifically, the CABAC entropy decoding method can receive a bin (binary bits) corresponding to each syntax element from the bitstream, determine a context model using information about the syntax element to be decoded, decoding information of adjacent blocks or blocks to be decoded, or information about symbols / bins decoded in previous stages, and generate symbols corresponding to the values of each syntax element by performing arithmetic decoding on the bins based on the determined context model to predict the probability of generating bins. At this point, the CABAC entropy decoding method can determine the context model and then update the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin. Prediction-related information from the information decoded by the entropy decoder 310 can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values (i.e., quantized transform coefficients and related parameter information) from the entropy decoding performed by the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). Additionally, filtering-related information from the information decoded by the entropy decoder 310 can be provided to the filter 350. Meanwhile, the receiver (not shown) for receiving the signal output from the encoding device can also be configured as an internal / external component of the decoding device 300, or the receiver can be a component of the entropy decoder 310. Furthermore, the decoding device according to this disclosure can be referred to as a video / image / picture decoding device, and the decoding device can be classified as an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of the following: a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0093] Dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. Dequantizer 321 can rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device. Dequantizer 321 can use quantization parameters (e.g., quantization step size information) to perform dequantization on the quantized transform coefficients and obtain the transform coefficients.
[0094] The inverse transformer 322 performs inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).
[0095] Predictor 330 can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on information about the prediction output from entropy decoder 310, and can determine a specific intra-frame / inter-frame prediction mode.
[0096] The predictor can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Alternatively, the predictor can predict blocks based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding, such as screen content coding (SCC), in games and similar applications. IBC essentially performs prediction within the current frame, but can be performed similarly to inter-frame prediction, such that a reference block is derived within the current frame. That is, IBC can use at least one inter-frame prediction technique described herein. The palette mode can be considered an example of intra-frame coding or intra-frame prediction. When a palette mode is applied, information about the palette table and palette index can be included in the video / image information and signaled.
[0097] Intra-predictor 331 can refer to samples in the current image to predict the current block. Depending on the prediction mode, the referenced samples may be located near or separated from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. Intra-predictor 331 can use the prediction modes applied to neighboring blocks to determine the prediction mode applied to the current block.
[0098] Inter-frame predictor 332 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. For example, inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference image index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the inter-frame prediction mode of the current block.
[0099] Adder 340 generates a reconstruction signal (reconstructed image, reconstruction block, reconstruction sample array) by adding the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including inter-frame predictor 332 and / or intra-frame predictor 331). If the block to be processed has no residual, such as when a skip mode is applied, the prediction block can be used as the reconstruction block.
[0100] Adder 340 can be referred to as a reconstructor or reconstruction block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, and can be output through filtering as described below, or it can be used for inter-frame prediction of the next image.
[0101] Simultaneously, Luminance Mapping and Chromatography Scaling (LMCS) can be applied in image decoding processing.
[0102] Filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 360 (specifically, the DPB of memory 360). For example, various filtering methods may include deblocking filtering, sample adaptive shifting, adaptive loop filtering, bilateral filtering, etc.
[0103] The (modified) reconstructed image stored in the DPB of memory 360 can be used as a reference image in inter-frame predictor 332. Memory 360 can store motion information of blocks in which motion information within the current image is derived (decoded) and / or motion information of blocks within previously reconstructed images. The stored motion information can be transmitted to inter-frame predictor 260 to be used as motion information for spatially adjacent blocks or temporally adjacent blocks. Memory 360 can store reconstructed samples of reconstructed blocks within the current image and transmit the stored reconstructed samples to intra-frame predictor 331.
[0104] In this document, the exemplary embodiments described in the filter 260, inter-frame predictor 221 and intra-frame predictor 222 of the encoding device 200 can be equally applied to or correspond to the filter 350, inter-frame predictor 332 and intra-frame predictor 331 of the decoding device 300, respectively.
[0105] Simultaneously, as described above, prediction is performed during video encoding to improve compression efficiency. This generates a prediction block that includes prediction samples for the current block, which is to be encoded (i.e., the target block). Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived in the same manner in both the encoding and decoding devices, and the encoding device can signal information about the residual between the original block and the prediction block (residual information) instead of the original sample values of the original block to the decoding device, thereby improving image encoding efficiency. The decoding device can derive a residual block including residual samples based on the residual information, add the residual block and the prediction block to generate a reconstructed block including reconstructed samples, and generate a reconstructed image including the reconstructed block.
[0106] Residual information can be generated through transformation and quantization procedures. For example, the encoding device can derive a residual block between the original block and the prediction block, perform a transformation process on the residual samples (residual sample array) included in the residual block to derive transform coefficients, perform a quantization process on the transform coefficients to derive quantized transform coefficients, and signal the relevant residual information (via bitstream) to the decoding device. Here, the residual information may include the value information, position information, transform technique, transform kernel, quantization parameters, etc., of the quantized transform coefficients. The decoding device can perform a dequantization / inverse transform process based on the residual information and derive residual samples (or residual blocks). The decoding device can generate a reconstructed image based on the prediction block and the residual block. Furthermore, for reference in inter-frame prediction of subsequent images, the encoding device can also perform dequantization / inverse transform on the quantized transform coefficients to derive residual blocks and generate a reconstructed image based on these residual blocks.
[0107] Intra-frame prediction can instruct the generation of prediction samples for the current block based on reference samples in the image to which the current block belongs (hereinafter referred to as the current image). When intra-frame prediction is applied to the current block, neighboring reference samples to be used for intra-frame prediction of the current block can be derived. The neighboring reference samples of the current block may include: samples adjacent to the left boundary and a total of 2xnH samples adjacent to the lower left of the current block of size nWxnH; samples adjacent to the upper boundary of the current block and a total of 2xnW samples adjacent to the upper right of the current block; and one sample adjacent to the upper left of the current block. Alternatively, the neighboring reference samples of the current block may include multiple columns of upper neighboring samples and multiple rows of left neighboring samples. In addition, the neighboring reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block of size nWxnH, a total of nW samples adjacent to the lower boundary of the current block, and one sample adjacent to the lower right of the current block.
[0108] However, some of the neighboring reference samples of the current block may not have been decoded or may be unavailable. In this case, the decoding device can construct neighboring reference samples to be used for prediction by replacing unavailable samples with available samples. Alternatively, neighboring reference samples to be used for prediction can be constructed by interpolation of available samples.
[0109] When deriving neighboring reference samples, (i) the predicted sample can be derived based on the average or interpolation of the neighboring reference samples of the current block, or (ii) the predicted sample can be derived based on a reference sample among the neighboring reference samples of the current block that is located in a specific (predictive) direction relative to the predicted sample. Case (i) can be referred to as non-directional mode or non-angular mode, and case (ii) can be referred to as directional mode or angular mode.
[0110] Alternatively, prediction samples can be generated by interpolating the first neighboring sample in the prediction direction of the intra-prediction mode within the current block and the second neighboring sample in the opposite direction from the neighboring reference sample, based on the prediction samples of the current block. This is known as Linear Interpolation Intra-Prediction (LIP). Additionally, a linear model (LM) can be used to generate chroma prediction samples based on luminance samples. This is known as LM mode or CCLM (Chromatic Component LM) mode.
[0111] Alternatively, a provisional prediction sample for the current block can be derived using filtered neighboring reference samples, or it can be derived by weighting the provisional prediction sample with at least one reference sample derived from the existing neighboring reference samples (i.e., unfiltered neighboring reference samples) according to the intra-prediction mode. This process is known as Position-Related Intra-Prediction (PDPC).
[0112] Furthermore, among multiple reference sample lines surrounding the current block, the reference sample line with the highest prediction accuracy is selected, and the predicted samples are derived using reference samples located in the prediction direction of the selected reference sample line. Intra-frame prediction coding can be performed by indicating the reference sample line used (by signaling) to the decoding device. This process can be referred to as multi-reference line intra-frame prediction or MRL-based intra-frame prediction.
[0113] Alternatively, the current block can be divided into vertical or horizontal sub-partitions to perform intra-prediction based on the same intra-prediction mode, but neighboring reference samples can be derived and used on a sub-partition basis. That is, in this case, the intra-prediction mode for the current block can be applied to all sub-partitions in the same way, but in some cases, intra-prediction performance can be improved by deriving and using neighboring reference samples on a sub-partition basis. This prediction method can be called intra-prediction based on intra-partition sub-partitions (ISP).
[0114] The intra-prediction methods described above can be referred to as intra-prediction types to distinguish them from intra-prediction modes. Intra-prediction types can be referred to by various terms such as intra-prediction techniques or additional intra-prediction modes. For example, an intra-prediction type (or additional intra-prediction mode) can include at least one of LIP, PDPC, MRL, and ISP mentioned above. General intra-prediction methods other than specific intra-prediction types such as LIP, PDPC, MRL, and ISP can be referred to as normal intra-prediction types. Normal intra-prediction types are typically applied when no specific intra-prediction type is applied, and prediction can be performed based on intra-prediction modes. Additionally, post-processing filtering can be performed on the derived prediction samples if necessary.
[0115] Specifically, the intra-frame prediction process may include an intra-frame prediction mode / type determination step, a neighboring reference sample derivation step, and a prediction sample derivation step based on the intra-frame prediction mode / type. Additionally, if necessary, a post-filtering step may be performed on the derived prediction samples.
[0116] Figure 4 An example of a video / image coding method based on intra-frame prediction is shown.
[0117] Reference Figure 4The encoding device performs intra-prediction on the current block (S400). The encoding device can deduce the intra-prediction mode / type for the current block, deduce neighboring reference samples for the current block, and generate prediction samples in the current block based on the intra-prediction mode / type and neighboring reference samples. Here, the processes of determining the intra-prediction mode / type, deduce neighboring reference samples, and generate prediction samples can be performed simultaneously, and any one process can be performed before the others. The encoding device can determine the mode / type applicable to the current block from multiple intra-prediction modes / types. The encoding device can compare the RD costs of the intra-prediction modes / types and determine the optimal intra-prediction mode / type for the current block.
[0118] Simultaneously, the encoding device can also perform a prediction sample filtering process. Prediction sample filtering can be referred to as post-filtering. Some or all of the prediction samples can be filtered according to the prediction sample filtering process. In some cases, the prediction sample filtering process can be omitted.
[0119] The encoding device generates residual samples for the current block based on the (filtered) prediction samples (S410). The encoding device can compare the prediction samples with the phase in the original samples of the current block and derive the residual samples.
[0120] The encoding device can encode image information including information about intra-frame prediction (prediction information) and residual information about residual samples (S420). The prediction information may include intra-frame prediction mode information and intra-frame prediction type information. The encoding device can output the encoded image information in the form of a bitstream. The output bitstream can be transmitted to the decoding device via a storage medium or a network.
[0121] Residual information may include the residual coding syntax to be described. The encoding device can derive the quantized transform coefficients by transforming / quantizing residual samples. The residual information may include information about the quantized transform coefficients.
[0122] Simultaneously, as described above, the encoding device can generate a reconstructed image (including reconstructed samples and reconstructed blocks). To this end, the encoding device can derive (modified) residual samples by dequantizing / inverse transforming the quantized transform coefficients again. As described above, the reason for transforming / quantizing the residual samples and then dequantizing / inverse transforming them again is to derive the same residual samples as those derived by the decoding device as described above. The encoding device can generate a reconstructed block including reconstructed samples for the current block based on the predicted samples and the (modified) residual samples. A reconstructed image of the current image can be generated based on the reconstructed block. As described above, in-loop filtering processes, etc., can be further applied to the reconstructed image.
[0123] Figure 5An example of a video / image decoding method based on intra-frame prediction is shown.
[0124] The decoding device can perform operations corresponding to those performed by the encoding device.
[0125] Prediction and residual information can be obtained from the bitstream. Residual samples for the current block can be derived based on the residual information. Specifically, the decoding device can derive transform coefficients by performing dequantization based on the quantized transform coefficients derived from the residual information, and perform an inverse transform on the transform coefficients to derive residual samples for the current block.
[0126] Specifically, the decoding device can deduce the intra-prediction mode / type of the current block based on the received prediction information (intra-prediction mode / type information) (S500). The decoding device can deduce the neighboring reference samples of the current block (S510). The decoding device generates prediction samples in the current block based on the intra-prediction mode / type and the neighboring reference samples (S520). In this case, the decoding device can perform a prediction sample filtering process. Prediction sample filtering can be referred to as post-filtering. Some or all of the prediction samples can be filtered according to the prediction sample filtering process. In some cases, the prediction sample filtering process can be omitted.
[0127] The decoding device generates residual samples for the current block based on the received residual information (S530). The decoding device can generate reconstructed samples for the current block based on the predicted samples and residual samples, and derive a reconstructed block including the reconstructed samples (S540). A reconstructed image of the current image can be generated based on the reconstructed block. As described above, in-loop filtering processes, etc., can also be applied to the reconstructed image.
[0128] Intra-luma_mpm_flag may include, for example, flag information indicating whether the most probable mode (MPM) is applied to the current block or whether the remaining modes are applied to the current block. When an MPM is applied to the current block, the prediction mode information may also include index information (e.g., intra_luma_mpm_idx) indicating one of the intra-luma_mpm_candidates. The intra-luma_mpm_candidates may include a list of MPM candidates or a list of MPMs. Additionally, if no MPM is applied to the current block, the intra-luma_mpm_remainder may also include residual mode information indicating one of the remaining intra-luma_mpm_remainders besides the intra-luma_mpm_candidates. The decoding device may determine the intra-luma_mpm_remainder for the current block based on the intra-luma_mpm_intra prediction mode information.
[0129] Furthermore, intra-prediction type information can be implemented in various forms. As an example, intra-prediction type information may include intra-prediction type index information indicating one of the intra-prediction types. As another example, intra-prediction type information may include at least one of the following: reference sample line information indicating whether MRL is applied to the current block and, if so, which reference sample line to use (e.g., intra_luma_ref_idx); ISP flag information indicating whether ISP is applied to the current block (e.g., intra_subpartitions_mode_flag); ISP type information indicating the segmentation type of the subpartition if ISP is applied (e.g., intra_subpartitions_split_flag); flag information indicating whether PDCP is applied; or flag information indicating whether LIP is applied. Additionally, intra-prediction type information may include a MIP flag indicating whether MIP is applied to the current block.
[0130] Intra-prediction mode information and / or intra-prediction type information can be encoded / decoded using the encoding methods described in this document. For example, intra-prediction mode information and / or intra-prediction type information can be encoded / decoded using entropy coding (e.g., CABAC, CAVLC).
[0131] Figure 6 An example of the intra-frame prediction process is shown.
[0132] Reference Figure 6 As described above, the intra-frame prediction process may include steps for determining the intra-frame prediction mode / type, deriving neighboring reference samples, and performing intra-frame prediction (generating prediction samples). The intra-frame prediction process may be performed in the encoding and decoding devices described above. In this document, the encoding device may include an encoding apparatus and / or a decoding apparatus.
[0133] Reference Figure 6 The coding device determines the intra-frame prediction mode / type (S600).
[0134] The encoding device can determine the intra prediction mode / type applicable to the current block from the various intra prediction modes / types described above, and can generate prediction-related information. The prediction-related information may include intra prediction mode information indicating the intra prediction mode applied to the current block and / or intra prediction type information indicating the intra prediction type applied to the current block. The decoding device can determine the intra prediction mode / type applicable to the current block based on the prediction-related information.
[0135] Intra-luma_mpm_flag may include, for example, flag information indicating whether the most probable mode (MPM) or a remaining mode is applied to the current block. When an MPM is applied to the current block, the prediction mode information may also include index information (e.g., intra_luma_mpm_idx) indicating one of the intra-luma_mpm_candidates. Intra-luma_mpm_candidates may include a list of MPM candidates or a list of MPMs. Additionally, when no MPM is applied to the current block, the intra-luma_mpm_remainder may include remaining mode information indicating one of the remaining intra-luma_mpm_remainders besides the intra-luma_mpm_candidates. The decoding device may determine the intra-luma_mpm_remainder for the current block based on the intra-luma_mpm_intra prediction mode information.
[0136] Furthermore, intra-prediction type information can be implemented in various forms. As an example, intra-prediction type information may include intra-prediction type index information indicating one of the intra-prediction types. As another example, intra-prediction type information may include at least one of the following: reference sample line information indicating whether MRL is applied to the current block and which reference sample line is used when applying MRL (e.g., intra_luma_ref_idx); ISP flag information indicating whether ISP is applied to the current block (e.g., intra_subpartitions_mode_flag); ISP type information indicating the segmentation type of the subpartition when ISP is applied (e.g., intra_subpartitions_split_flag); flag information indicating whether PDCP is applied; or flag information indicating whether LIP is applied. Additionally, intra-prediction type information may include a MIP flag indicating whether matrix-based intra-prediction (MIP) is applied to the current block.
[0137] For example, when applying intra-prediction, the intra-prediction mode to be applied to the current block can be determined by using the intra-prediction modes of neighboring blocks. For instance, the coding device can select one of the MPM candidates from a list of most probable modes (MPMs) derived from the intra-prediction modes and additional candidate modes of the current block's neighboring blocks (e.g., the left neighboring block and / or the upper neighboring block), or one of the remaining intra-prediction modes not included in the MPM candidates (and planar modes) based on the received MPM index. This MPM list can be configured to include or exclude planar modes as candidates. For example, the MPM list can have 6 candidates when it includes planar modes as candidates, and 5 candidates when it excludes planar modes as candidates. When the MPM list excludes planar modes as candidates, a signal indicating whether the intra-prediction mode of the current block is not a planar mode (e.g., intra_luma_not_planar_flag) can be used. For example, the MPM flag is signaled first, and the MPM index and non-planar flag can be signaled when the MPM flag is 1. Additionally, the MPM index can be signaled when the non-planar flag is 1. The reason the MPM list is configured not to include planar patterns as candidates is not because planar patterns are not MPMs, but because planar patterns are always considered MPMs; therefore, it is first checked whether an MPM is a planar pattern by signaling the flag (non-planar flag).
[0138] For example, MPM flags (e.g., `intra_luma_mpm_flag`) can be used to indicate whether the intra-prediction mode applied to the current block is included in the MPM candidates (and planar modes) or in the remaining modes. An MPM flag value of 1 indicates that the intra-prediction mode applied to the current block is included in the MPM candidates (and planar modes), and an MPM flag value of 0 indicates that the intra-prediction mode applied to the current block is not in the MPM candidates (and planar modes). A non-planar flag (e.g., `intra_luma_not_planar_flag`) value of 0 indicates that the intra-prediction mode applied to the current block is planar mode, and a non-planar flag value of 1 indicates that the intra-prediction mode applied to the current block is not planar mode. The MPM index can be signaled in the form of `mpm_idx` or `intra_luma_mpm_idx` syntax elements, and the remaining intra-prediction mode information can be signaled in the form of `rem_intra_luma_pred_mode` or `intra_luma_mpm_remainder` syntax elements. For example, the remaining intra-prediction mode information can be signaled using the syntax elements `rem_intra_luma_pred_mode` or `intra_luma_mpm_remainder`. For example, the remaining intra-prediction mode information can be indexed by prediction mode number to indicate one of the remaining intra-prediction modes not included in the MPM candidates (and planar modes) among all intra-prediction modes. The intra-prediction mode can be an intra-prediction mode for the luma component (samples). In the following, the intra-prediction mode information may include at least one of the following: an MPM flag (e.g., `intra_luma_mpm_flag`), a non-planar flag (e.g., `intra_luma_not_planar_flag`), an MPM index (e.g., `mpm_idx` or `intra_luma_mpm_idx`), and remaining intra-prediction mode information (e.g., `rem_intra_luma_pred_mode` or `intra_luma_mpm_remainder`). In this document, the MPM list may be referred to by various terms such as MPM candidate list or `candModeList`.
[0139] When MIP is applied to the current block, individual MPM flags (e.g., intra_mip_mpm_flag), MPM indexes (e.g., intra_mip_mpm_idx), and remaining intra-prediction mode information for MIP (e.g., intra_mip_mpm_remainder) can be signaled separately, without signaling non-plane flags.
[0140] In other words, typically when an image is divided into blocks, the current block to be encoded and its neighboring blocks have similar image characteristics. Therefore, there is a high probability that the current block and its neighboring blocks have the same or similar intra-prediction modes. Thus, the encoding device can use the intra-prediction modes of neighboring blocks to encode the intra-prediction mode of the current block.
[0141] The encoding device can be configured with a list of most probable modes (MPMs) for the current block. This MPM list may be referred to as the MPM candidate list. Here, MPM can refer to a mode used to improve coding efficiency by considering the similarity between the current block and neighboring blocks during intra-frame prediction mode coding. As mentioned above, the MPM list can be configured to include planar modes or to exclude planar modes. For example, when the MPM list includes planar modes, the number of candidates in the MPM list can be six. And when the MPM list does not include planar modes, the number of candidates in the MPM list can be five.
[0142] The encoding device can perform prediction based on various intra-prediction modes and determine the optimal intra-prediction mode based on rate distortion optimization (RDO) based on that prediction. In this case, the encoding device can determine the optimal intra-prediction mode by using only the MPM candidates and planar modes configured in the MPM list, or by further using the remaining intra-prediction modes along with the MPM candidates and planar modes configured in the MPM list. Specifically, for example, if the intra-prediction type of the current block is a specific type other than the normal intra-prediction type (e.g., LIP, MRL, or ISP), the encoding device can determine the optimal intra-prediction mode by considering only the MPM candidates and planar modes as intra-prediction mode candidates for the current block. That is, in this case, the intra-prediction mode for the current block can be determined only from the MPM candidates and planar modes, and in this case, the encoding / signaling of the MPM flag can be omitted. In this case, the decoding device can estimate the MPM flag as 1 without separately receiving the signal of the MPM flag.
[0143] Simultaneously, typically, when the intra-prediction mode of the current block is not a planar mode and is one of the MPM candidates in the MPM list, the encoding device generates an MPM index (mpm idx) indicating one of the MPM candidates. When the intra-prediction mode of the current block is not in the MPM list, the encoding device generates MPM remainder information (remaining intra-prediction mode information), which indicates the same mode as the intra-prediction mode of the current block among the remaining intra-prediction modes not included in the MPM list (and planar modes). The MPM remainder information may include, for example, the intra_luma_mpm_remainder syntax element.
[0144] The decoding device obtains intra-prediction mode information from the bitstream. As described above, the intra-prediction mode information may include at least one of an MPM flag, a non-planar flag, an MPM index, and remaining MPM information (remaining intra-prediction mode information). The decoding device can configure an MPM list. The MPM list is configured in the same manner as the MPM list configured in the encoding device. That is, the MPM list may include intra-prediction modes of adjacent blocks, and may also include specific intra-prediction modes according to a predetermined method.
[0145] The decoding device can determine the intra-prediction mode of the current block based on the MPM list and intra-prediction mode information. For example, when the MPM flag is 1, the decoding device can deduce a planar mode as the intra-prediction mode of the current block (based on the non-planar flag), or it can deduce the candidate indicated by the MPM index from the MPM candidates in the MPM list as the intra-prediction mode of the current block. Here, the MPM candidates may only indicate candidates included in the MPM list, or they may include not only candidates included in the MPM list, but also planar modes that are applicable when the MPM flag is 1.
[0146] As another example, when the MPM flag value is 0, the decoding device can deduce the intra-prediction mode and planar mode indicated by the remaining intra-prediction mode information (which may be referred to as MPM remaining information) in the remaining intra-prediction modes not included in the MPM list as the intra-prediction mode for the current block. Simultaneously, as another example, when the intra-prediction type of the current block is a specific type (e.g., LIP, MRL, or ISP), the decoding device can deduce the planar mode or a candidate indicated by the MPM flag in the MPM list as the intra-prediction mode for the current block without parsing / decoding / checking the MPM flag.
[0147] The encoding device derives the peripheral reference samples of the current block (S610). When intra-frame prediction is applied to the current block, neighboring reference samples to be used for intra-frame prediction of the current block can be derived. The neighboring reference samples of the current block may include: samples adjacent to the left boundary and a total of 2xnH samples adjacent to the lower left of the current block of size nWxnH; samples adjacent to the upper boundary of the current block and a total of 2xnW samples adjacent to the upper right of the current block; and one sample adjacent to the upper left of the current block. Alternatively, the neighboring reference samples of the current block may include multiple columns of upper neighboring samples and multiple rows of left neighboring samples. In addition, the neighboring reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block of size nWxnH, a total of nW samples adjacent to the lower boundary of the current block, and one sample adjacent to the lower right of the current block.
[0148] On the other hand, when MRL is applied (i.e., when the MRL index value is greater than 0), adjacent reference samples can be located on rows 1 and 2, instead of on row 0 adjacent to the current block on the left / top. In this case, the number of adjacent reference samples can be further increased. Meanwhile, when ISP is applied, adjacent reference samples can be derived on a sub-partition basis.
[0149] The encoding device derives prediction samples by performing intra-frame prediction on the current block (S620). The encoding device can derive prediction samples based on intra-frame prediction mode / type and neighboring samples. The encoding device can derive reference samples from neighboring reference samples of the current block according to the intra-frame prediction mode of the current block, and can derive prediction samples of the current block based on the reference samples.
[0150] Simultaneously, when applying inter-frame prediction, the predictor of the encoding / decoding device can derive prediction samples by performing inter-frame prediction on a block-by-block basis. Inter-frame prediction can be applied when performing prediction on the current block. That is, the predictor of the encoding / decoding device (more specifically, the inter-frame predictor) can derive prediction samples by performing inter-frame prediction on a block-by-block basis. Inter-frame prediction can represent a prediction derived by a method that depends on data elements (e.g., sample values or motion information) of images other than the current image. When inter-frame prediction is applied to the current block, the prediction block (prediction sample array) for the current block can be derived based on the reference block (reference sample array) on the reference image indicated by the reference image index, specified by the motion vector. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information of the current block can be predicted on a block, sub-block, or sample-by-sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information can include motion vectors and reference image indices. Motion information can also include inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the application of inter-frame prediction, neighboring blocks can include spatially adjacent blocks existing in the current image and temporally adjacent blocks existing in a reference image. The reference image including the reference block and the reference image including the temporally adjacent blocks can be the same as or different from each other. Temporally adjacent blocks can be referred to by names such as juxtaposed reference blocks, juxtaposed CUs (colCU), etc., and the reference image including the temporally adjacent blocks can be referred to as a juxtaposed image (colPic). For example, a candidate list of motion information can be configured based on the neighboring blocks of the current block, and a signal can be used to indicate which candidate's flag or index information is selected (used) to derive the motion vector of the current block and / or the reference image index. Inter-frame prediction can be performed based on various prediction modes, and for example, in skip mode and merge mode, the motion information of the current block can be the same as the motion information of the selected neighboring blocks. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, the motion vector of the selected neighboring block can be used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived by using the sum of the motion vector predictor and the motion vector difference.
[0151] Motion information can also include L0 motion information and / or L1 motion information, depending on the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). The L0 direction motion vector can be referred to as the L0 motion vector or MVL0, while the L1 direction motion vector can be referred to as the L1 motion vector or MVL1. Prediction based on the L0 motion vector is called L0 prediction, prediction based on the L1 motion vector is called L1 prediction, and prediction based on both L0 and L1 motion vectors is called bidirectional prediction. Here, the L0 motion vector can indicate the motion vector associated with the reference image list L0, and the L1 motion vector can indicate the motion vector associated with the reference image list L1. The reference image list L0 can include images preceding the current image in output order, and the reference image list L1 can include images following the current image in output order as reference images. The previous image can be referred to as the forward (reference) image, and the subsequent image can be referred to as the backward (reference) image. The reference image list L0 can also include images following the current image in output order as reference images. In this case, previous images can be indexed first in the reference image list L0, and then subsequent images can be indexed. The reference image list L1 can also include images that precede the current image in the output order as reference images. In this case, subsequent images can be indexed first in the reference image list L1, and then previous images can be indexed. Here, the output order can correspond to the image order count (POC) order.
[0152] The video / image coding process based on inter-frame prediction may schematically include, for example, the following.
[0153] Figure 7 An example of a video / image coding method based on inter-frame prediction is shown.
[0154] The encoding device performs inter-frame prediction on the current block (S700). The encoding device can deduce the inter-frame prediction mode and motion information of the current block, and generate prediction samples for the current block. Here, the inter-frame prediction mode determination process, the motion information derivation process, and the prediction sample generation process can be executed simultaneously, and any one process can be executed earlier than the others. For example, the inter-frame prediction unit of the encoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit can determine the prediction mode for the current block, the motion information derivation unit can deduce the motion information of the current block, and the prediction sample derivation unit can deduce the prediction samples for the current block. For example, the inter-frame prediction unit of the encoding device can search for blocks similar to the current block in a predetermined region (search region) of a reference image through motion estimation, and deduce a reference block whose difference from the current block is the smallest or equal to or less than a predetermined standard. Based on this, a reference image index indicating the reference image where the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding device can determine the mode applied to the current block from various prediction modes. The encoding device can compare the RD costs of various prediction modes and determine the best prediction mode for the current block.
[0155] For example, when a skip mode or merge mode is applied to the current block, the encoding device can configure a merge candidate list, as described below, and deduce a reference block among the reference blocks indicated by the merge candidates included in the merge candidate list whose difference from the current block is the smallest, equal to, or less than a predetermined standard. In this case, a merge candidate associated with the deduced reference block can be selected, and merge index information indicating the selected merge candidate can be generated and signaled to the decoding device. The motion information of the current block can be deduced using the motion information of the selected merge candidate.
[0156] As another example, when the (A)MVP mode is applied to the current block, the encoding device can configure an (A)MVP candidate list, as described below, and use the motion vector of the selected MVP candidate from the motion vector prediction sub-candidates included in the (A)MVP candidate list as the MVP of the current block. In this case, for example, it indicates that the motion vector of the reference block derived through motion estimation can be used as the motion vector of the current block, and the MVP candidate with the motion vector having the smallest difference from the motion vector of the current block can become the selected MVP candidate. The motion vector difference (MVD) can be derived, which is the difference obtained by subtracting the MVP from the motion vector of the current block. In this case, information about the MVD can be signaled to the decoding device. Furthermore, when the (A)MVP mode is applied, the value of the reference picture index can be configured as reference picture index information and signaled separately to the decoding device.
[0157] The encoding device can derive residual samples based on the predicted samples (S710). The encoding device can derive residual samples by comparing the original samples and the predicted samples of the current block.
[0158] The encoding device encodes image information, including prediction information and residual information (S720). The encoding device can output the encoded image information in the form of a bitstream. The prediction information may include information about prediction mode information (e.g., skip flag, merge flag, or mode index, etc.) and information about motion information as information related to the prediction process. The information about motion information may include candidate selection information (e.g., merge index, MVP flag, or MVP index), which is information used to derive motion vectors. In addition, the information about motion information may include information about MVD and / or reference image index information. Furthermore, the information about motion information may include information indicating whether L0 prediction, L1 prediction, or bidirectional prediction is applied. The residual information is information about residual samples. The residual information may include information about the quantized transform coefficients of the residual samples.
[0159] The output bitstream can be stored in a (digital) storage medium and transmitted to a decoding device or via a network.
[0160] Simultaneously, as described above, the encoding device can generate a reconstructed image (including reconstructed samples and reconstructed blocks) based on the reference samples and residual samples. This is to derive the same prediction result as that performed by the decoding device, and as a result, encoding efficiency can be improved. Therefore, the encoding device can store the reconstructed image (or reconstructed samples or reconstructed blocks) in memory and use the reconstructed image as a reference image. An in-loop filtering process can be further applied to the reconstructed image as described above.
[0161] The video / image decoding process based on inter-frame prediction may schematically include, for example, the following processes.
[0162] Figure 8 An example of a video / image decoding method based on inter-frame prediction is shown.
[0163] Reference Figure 8 The decoding device can perform operations corresponding to those performed by the encoding device. The decoding device can perform predictions on the current block based on the received prediction information and derive prediction samples.
[0164] Specifically, the decoding device can determine the prediction mode for the current block based on the received prediction information (S800). The decoding device can determine which inter-frame prediction mode to apply to the current block based on the prediction mode information in the prediction information.
[0165] For example, the merge flag can be used to determine whether to apply the merge mode or the (A)MVP mode to the current block. Alternatively, a variety of inter-frame prediction mode candidates can be selected based on the mode index. Inter-frame prediction mode candidates may include skip mode, merge mode, and / or (A)MVP mode, or may include various inter-frame prediction modes described below.
[0166] The decoding device derives motion information for the current block based on the determined inter-frame prediction mode (S810). For example, when a skip mode or merge mode is applied to the current block, the decoding device can configure a merge candidate list as described below and select a merge candidate from among the merge candidates included in the merge candidate list. Here, the selection can be performed based on selection information (merge index). The motion information for the current block can be derived using the motion information of the selected merge candidate. The motion information of the selected merge candidate can be used as the motion information for the current block.
[0167] As another example, when the (A)MVP mode is applied to the current block, the decoding device can be configured with an (A)MVP candidate list, as described below, and use the motion vector of the selected MVP candidate from the motion vector prediction sub-candidates included in the (A)MVP candidate list as the MVP of the current block. Here, selection can be performed based on selection information (MVP flag or MVP index). In this case, the MVD of the current block can be derived based on information about the MVD, and the motion vector of the current block can be derived based on the MVD and the MVP of the current block. Furthermore, the reference image index of the current block can be derived based on reference image index information. The image indicated by the reference image index in the reference image list for the current block can be derived as the reference image referenced for inter-frame prediction of the current block.
[0168] Furthermore, as described below, motion information for the current block can be derived without a candidate list configuration, and in this case, the motion information for the current block can be derived based on the process disclosed in the prediction mode. In this case, the candidate list configuration can be omitted.
[0169] The decoding device can generate prediction samples for the current block based on the motion information of the current block (S820). In this case, a reference image can be derived based on the reference image index of the current block, and prediction samples for the current block can be derived by using samples of the reference block indicated by the motion vector of the current block on the reference image. In some cases, a filtering process for all or some of the prediction samples for the current block can be further performed.
[0170] For example, the inter-frame prediction unit of the decoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit may determine the prediction mode for the current block based on the received prediction mode information. The motion information derivation unit may derive the motion information (motion vector and / or reference image index) of the current block based on information about the received motion information. The prediction sample derivation unit may derive the prediction samples of the current block.
[0171] The decoding device generates residual samples for the current block based on the received residual information (S830). The decoding device can generate reconstruction samples for the current block based on the predicted samples and residual samples, and generate a reconstructed image based on the generated reconstruction samples (S840). Thereafter, as described above, an in-loop filtering process can be further applied to the reconstructed image.
[0172] Figure 9 An example of the inter-frame prediction process is shown.
[0173] Reference Figure 9 As described above, the inter-frame prediction process may include an inter-frame prediction mode determination step, a motion information derivation step based on the determined prediction mode, and a prediction processing step (prediction sample generation) based on the derivation of motion information. The inter-frame prediction process may be performed by the encoding and decoding devices described above. In this document, the encoding device may include an encoding device and / or a decoding device.
[0174] Reference Figure 9 The encoding device determines the inter-frame prediction mode (S900) for the current block. Various inter-frame prediction modes can be used for prediction of the current block in the image. For example, various modes can be used, such as merge mode, skip mode, motion vector prediction (MVP) mode, affine mode, sub-block merge mode, merge with MVD (MMVD) mode, and history motion vector prediction (HMVP) mode. Decoder-side motion vector correction (DMVR) mode, adaptive motion vector resolution (AMVR) mode, bidirectional prediction with CU-level weights (BCW), bidirectional optical flow (BDOF), etc., can also be used as additional modes. Affine mode can also be referred to as affine motion prediction mode. MVP mode can also be referred to as advanced motion vector prediction (AMVP) mode. In this document, some modes and / or motion information candidates derived from some modes can also be included in one of the motion information related candidates in other modes. For example, HMVP candidates can be added to the merge candidates of merge / skip mode, or also to the MVP candidates of MVP mode. If an HMVP candidate is used as a motion information candidate for merge mode or skip mode, then the HMVP candidate can be referred to as an HMVP merge candidate.
[0175] Prediction mode information, indicating the inter-frame prediction mode of the current block, can be signaled from the encoding device to the decoding device. In this case, the prediction mode information can be included in the bitstream and received by the decoding device. The prediction mode information may include index information indicating one of a plurality of candidate modes. Alternatively, the inter-frame prediction mode can be indicated by hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, whether to apply a skip mode can be indicated by signaling a skip flag, whether to apply a merge mode can be indicated by signaling a merge flag when a skip mode is not applied, and whether to apply the MVP mode, or additional flags for differentiation can be signaled when the merge mode is not applied. Affine modes can be signaled as independent modes or as dependent modes depending on the merge mode or the MVP mode. For example, affine modes may include an affine merge mode and an affine MVP mode.
[0176] The encoding device derives motion information for the current block (S910). Motion information derivation can be based on inter-frame prediction modes.
[0177] The encoding device can use motion information from the current block to perform inter-frame prediction. The encoding device can derive the optimal motion information for the current block through a motion estimation process. For example, the encoding device can search for highly correlated similar reference blocks in a predetermined search range within a reference image, using original blocks from the original image as the current block, on a fractional-pixel basis, and derive motion information from the searched reference blocks. Block similarity can be derived based on the difference in phase-based sample values. For example, block similarity can be calculated based on the sum of absolute differences (SAD) between the current block (or its template) and a reference block (or its template). In this case, motion information can be derived based on the reference block with the minimum SAD in the search region. The derived motion information can be signaled to the decoding device based on the inter-frame prediction mode using various methods.
[0178] The encoding device performs inter-frame prediction based on motion information for the current block (S920). The encoding device can derive prediction samples for the current block based on the motion information. The current block, which includes the prediction samples, can be referred to as the prediction block.
[0179] Figure 10 An example of a layered structure for encoded images / videos is shown.
[0180] refer to Figure 10 Encoded images / videos are divided into the VCL (Video Coding Layer), which handles image / video decoding and processing, the subsystem for sending and storing encoded information, and the Network Abstraction Layer (NAL), which exists between the VCL and the subsystems and is responsible for network adaptation functions.
[0181] VCL can generate VCL data that includes compressed image data (slice data), or parameter sets that include Picture Parameter Set (PPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), etc., or additional supplementary enhancement information (SEI) messages necessary for the image decoding process.
[0182] In NAL, NAL units are generated by adding header information (NAL unit header) to the Raw Byte Sequence Payload (RBSP) generated in VCL. In this case, RBSP refers to slice data, parameter sets, SEI messages, etc., generated in VCL. The NAL unit header may include NAL unit type information specified according to the RBSP data included in the corresponding NAL unit.
[0183] As shown in the figure, NAL units can be divided into VCL NAL units and non-VCL NAL units based on the RBSP generated in the VCL. A VCL NAL unit can mean an NAL unit that includes information about the image (slice data), while a non-VCL NAL unit can mean an NAL unit that contains information needed to decode the image (parameter set or SEI message).
[0184] The aforementioned VCL NAL units and non-VCL NAL units can be transmitted over the network by adding header information according to the subsystem's data standard. For example, NAL units can be converted into predetermined standard data formats, such as H.266 / VVC file format, Real-time Transport Protocol (RTP), Transport Stream (TS), etc., and transmitted over various networks.
[0185] As described above, in a NAL unit, the NAL unit type can be specified according to the RBSP data structure included in the corresponding NAL unit, and information about the NAL unit type can be stored in the NAL unit header and notified by a signal in the NAL unit header.
[0186] For example, NAL units can be roughly classified into VCL NAL unit types and non-VCL NAL unit types depending on whether the NAL unit includes information about the image (slice data). VCL NAL unit types can be classified according to the attributes and type of the image included in the VCL NAL unit, and non-VCL NAL unit types can be classified according to the type of parameter set.
[0187] Below is an example of a NAL cell type specified based on the type of the parameter set included in a non-VCL NAL cell type.
[0188] -APS (Adaptive Parameter Set) NAL Unit: The type of NAL unit including APS.
[0189] -DPS (Decoding Parameter Set) NAL Unit: The type of NAL unit including DPS.
[0190] -VPS (Video Parameter Set) NAL Unit: Includes the type of NAL unit for the VPS.
[0191] -SPS (Sequence Parameter Set) NAL Unit: The type of NAL unit that includes SPS.
[0192] -PPS (Image Parameter Set) NAL Unit: The type of NAL unit including PPS.
[0193] -PH (Picture Header) NAL Unit: Types of NAL units including PH.
[0194] The aforementioned NAL unit type contains syntax information for the NAL unit type, and this syntax information can be stored in the NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified by the nal_unit_type value.
[0195] Furthermore, as mentioned above, an image can include multiple slices, and a slice can include a slice header and slice data. In this case, an image header can also be added to multiple slices within an image (slice header and slice dataset). The image header (image header syntax) can include information / parameters typically applicable to images. In this document, tile groups can be mixed with or replaced by slices or images. Additionally, in this document, tile group headers can be mixed with or replaced by slice headers or image headers.
[0196] A slice header (slice header syntax) may include information / parameters typically applicable to a slice. APS (APS syntax) or PPS (PPS syntax) may include information / parameters typically applicable to one or more slices or images. SPS (SPS syntax) may include information / parameters typically applicable to one or more sequences. VPS (VPS syntax) may include information / parameters typically applicable to multiple layers. DPS (DPS syntax) may include information / parameters typically applicable to the entire video. DPS may include information / parameters related to the concatenation of encoded video sequences (CVS). In this document, the High-Level Syntax (HLS) may include at least one of the following: APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, image header syntax, and slice header syntax.
[0197] In this document, the image / video information encoded in the encoding device and signaled to the decoding device in the form of a bitstream may include, in addition to image segmentation information in the image, intra / inter-frame prediction information, residual information, in-loop filtering information, information included in the slice header, information included in the image header, information included in the APS, information included in the PPS, information included in the SPS, information included in the VPS, and / or information included in the DPS. Additionally, the image / video information may also include information from the NAL unit header.
[0198] Simultaneously, to compensate for differences between the original and reconstructed images caused by errors in compression coding processes such as quantization, an in-loop filtering process can be performed on the reconstructed samples or reconstructed images as described above. As mentioned above, in-loop filtering can be performed by filters from the encoding device and the decoding device, and deblocking filters, SAO, and / or adaptive loop filters (ALF) can be applied. For example, the ALF process can be performed after the deblocking filtering and / or SAO processing is completed. However, even in this case, the deblocking filtering and / or SAO processing can be omitted.
[0199] Figure 11 This is a flowchart that schematically illustrates an example of an ALF process. Figure 8 The ALF procedures disclosed herein can be executed in both encoding and decoding devices. In this document, encoding devices may include encoding devices and / or decoding devices.
[0200] refer to Figure 11 The encoding device derives a filter for the ALF (S1100). The filter may include filter coefficients. The encoding device can determine whether to apply the ALF, and when it is determined that the ALF is applied, it can derive a filter including the filter coefficients for the ALF. The filter (coefficients) of the ALF or the information about the derivation of the filter (coefficients) of the ALF may be referred to as ALF parameters. Information about whether the ALF is applied (e.g., an ALF availability flag) and ALF data for deriving the filter can be signaled from the encoding device to the decoding device. The ALF data may include information about the filter used to derive the ALF. Furthermore, for example, for layered control of the ALF, the ALF availability flag can be signaled at the SPS, picture header, slice header, and / or CTB levels, respectively.
[0201] To derive the filters for the ALF, the activity and / or directivity of the current block (or the ALF target block) are derived, and the filters can be derived based on the activity and / or directivity. For example, the ALF process can be applied in 4×4 block units (based on the luma component). The current block or the ALF target block can be, for example, a CU, or a 4x4 block within a CU. Specifically, for example, the filters for the ALF can be derived based on a first filter derived from information included in the ALF data and a predefined second filter, and the encoding device can select one of the filters based on the activity and / or directivity. The encoding device can use filter coefficients included in the selected filters of the ALF.
[0202] The encoding device performs filtering based on filters (S1110). Modified reconstructed samples can be derived based on the filtering. For example, filter coefficients in the filters can be arranged or assigned according to the filter shape, and filtering can be performed on the reconstructed samples in the current block. Here, the reconstructed samples in the current block can be reconstructed samples after the deblocking filter processing and SAO processing are completed. For example, a single filter shape can be used, or a filter shape selected from a plurality of predetermined filter shapes can be used. For example, the filter shape applied to the luma component and the filter shape applied to the chroma component can be different. For example, a 7×7 diamond filter shape can be used for the luma component, and a 5×5 diamond filter shape can be used for the chroma component.
[0203] Figure 12 An example of the shape of an ALF filter is shown.
[0204] Figure 12 (a) shows the shape of a 7×7 rhombus filter. Figure 6 (b) shows the shape of a 5×5 rhombus filter. Figure 6In this document, Cn in the filter shape represents the filter coefficient. When n in Cn is the same, this indicates that the same filter coefficient can be assigned. The position and / or unit where filter coefficients are assigned according to the filter shape of the ALF can be called a filter tap. In this case, a filter coefficient can be assigned to each filter tap, and the arrangement of the filter taps can correspond to the filter shape. The filter tap located at the center of the filter shape can be called the center filter tap. The same filter coefficient can be assigned to two filter taps with the same n value, and these two filter taps exist in positions corresponding to each other relative to the center filter tap. For example, in the case of a 7×7 diamond filter shape, there are 25 filter taps, and since filter coefficients C0 to C11 are assigned in a centrally symmetric manner, only 13 filter coefficients can be used to assign filter coefficients to the 25 filter taps. Furthermore, for example, in the case of a 5×5 diamond filter shape, there are 13 filter taps, and since filter coefficients C0 to C5 are assigned in a centrally symmetric manner, only 7 filter coefficients can be used to assign filter coefficients to the 13 filter taps. For example, to reduce the amount of data regarding the information about the filter coefficients notified by signals, 12 of the 13 filter coefficients in a 7×7 diamond filter shape are (explicitly) notified by signals, and one filter coefficient can be (implicitly) derived. Furthermore, for example, 6 of the 7 filter coefficients in a 5×5 diamond filter shape can be (explicitly) notified by signals, and one filter coefficient can be (implicitly) derived.
[0205] According to the embodiments in this document, the ALF parameters used in the ALF process can be signaled via an Adaptive Parameter Set (APS). The ALF parameters can be derived from filter information or ALF data used for ALF.
[0206] ALF is an in-loop filtering technique applicable to image / video coding as described above. ALF can be performed using a Wiener-based adaptive filter. This minimizes the mean square error (MSE) between the original sample and the decoded sample (or reconstructed sample). Advanced design features of ALF tools can incorporate syntax elements accessible in the SPS and / or slice header (or tile group header).
[0207] Figure 13 An example of a hierarchical structure for ALF data is shown.
[0208] refer to Figure 13A encoded video sequence (CVS) may include a sequence of images (SPS), one or more picture pins (PPS), and one or more subsequent encoded images. Each encoded image can be divided into rectangular regions. These rectangular regions may be referred to as tiles. One or more tiles may be aggregated to form a tile group or slice. In this case, the tile group header may be linked to the PPS, and the PPS may be linked to the SPS. According to existing methods, ALF data (ALF parameters) is included in the tile group header. Considering that a video consists of multiple images and each image includes multiple tiles, frequent ALF data (ALF parameter) signaling in units of tile groups reduces coding efficiency.
[0209] According to the embodiments presented in this document, ALF parameters can be signaled by including them in the APS as follows.
[0210] Figure 14 Another example of the hierarchical structure of ALF data is shown.
[0211] refer to Figure 14 An APS is defined, and the APS can carry the necessary ALF data (ALF parameters). Additionally, the APS can have self-identification parameters and ALF data. The self-identification parameters of the APS can include an APS ID. That is, in addition to the ALF data field, the APS can also include information indicating the APS ID. The tile group header or slice header can use the APS index information to reference the APS. In other words, the tile group header or slice header can include APS index information, and the ALF procedure for the target block can be performed based on the ALF data (ALF parameters) included in the APS with the APS ID indicated by the APS index information. Here, the APS index information can be referred to as the APS ID information.
[0212] Additionally, the SPS can include flags that allow the use of ALF. For example, the SPS can be checked when CVS starts, and flags can be checked within the SPS. For example, the SPS can include the syntax in Table 1 below. The syntax in Table 1 can be part of the SPS.
[0213] [Table 1]
[0214]
[0215] For example, the semantics of grammatical elements included in the grammar in Table 1 can be specified as shown in the table below.
[0216] [Table 2]
[0217]
[0218] That is, the `sps_alf_enabled_flag` syntax element can indicate whether ALF is available based on a value of 0 or 1. The `sps_alf_enabled_flag` syntax element can be referred to as the ALF availability flag (which can be called the first ALF availability flag) and can be included in the SPS. That is, the ALF availability flag can be signaled at the SPS (or at the SPS level). When the value of the ALF availability flag signaled by the SPS is 1, it can be determined that ALF is essentially available for use with images in the CVS referencing the SPS. Meanwhile, as mentioned above, ALF can be individually enabled / disabled by signaling additional availability flags at a level lower than the SPS.
[0219] For example, if the ALF tool is available for CVS, an additional availability flag (which may be referred to as the second ALF availability flag) can be signaled in the tile group header or slice header. For instance, when the ALF is available at the SPS level, the second ALF availability flag can be parsed / signed. If the value of the second ALF availability flag is 1, the ALF data can be parsed via the tile group header or slice header. For example, the second ALF availability flag can specify ALF availability conditions for the luma and chroma components. ALF data can be accessed via APS ID information.
[0220] [Table 3]
[0221]
[0222] [Table 4]
[0223]
[0224] For example, the semantics of grammatical elements included in the grammars in Table 3 or Table 4 can be specified as shown in the following tables.
[0225] [Table 5]
[0226]
[0227] [Table 6]
[0228]
[0229]
[0230] The second ALF availability flag can include either the tile_group_alf_enabled_flag syntax element or the slice_alf_enabled_flag syntax element.
[0231] Based on APS ID information (such as the tile_group_aps_id syntax element or the slice_aps_id syntax element), the APS referenced by the corresponding tile group or the corresponding slice can be identified. APS may include ALF data.
[0232] Additionally, the structure of the APS, including ALF data, can be described based on, for example, the following syntax and semantics. The syntax in Table 7 can be part of the APS.
[0233] [Table 7]
[0234]
[0235] [Table 8]
[0236]
[0237]
[0238] As mentioned above, the `adaptation_parameter_set_id` syntax element indicates the identifier of the corresponding APS. That is, the APS can be identified based on the `adaptation_parameter_set_id` syntax element. The `adaptation_parameter_set_id` syntax element can be referred to as the APS ID information. Furthermore, the APS may include ALF data fields. The ALF data fields can be parsed / signaled after the `adaptation_parameter_set_id` syntax element.
[0239] Furthermore, for example, APS extension flags (e.g., the aps_extension_flag syntax element) can be resolved / signed in the APS. APS extension flags can indicate the presence of the APS extension data flag (aps_extension_data_flag) syntax element. APS extension flags can be used, for example, to provide extension points for later versions of the VVC standard.
[0240] The core processing / handling of ALF information can be performed in the slice header or tile group header. The aforementioned ALF data fields can include information about the processing of ALF filters. For example, information that can be extracted from ALF data fields includes information about the number of filters used, information indicating whether ALF is applied only to the luminance component, information about the color components, and information about the Exponential Columbus (EG) parameter and / or information about the delta values of the filter coefficients, etc.
[0241] Additionally, ALF data fields can include, for example, the following ALF data syntax.
[0242] [Table 9]
[0243] For example, the semantics of grammatical elements included in the grammar in Table 9 can be specified as shown in the table below.
[0244] [Table 10]
[0245]
[0246]
[0247]
[0248]
[0249] For example, parsing ALF data via a tile group header or slice header can begin by first parsing / signaling the `alf_chroma_idc` syntax element. The `alf_chroma_idc` syntax element can have values ranging from 0 to 3. The value indicates whether the ALF-based filtering process is applied only to the luma component or to a combination of luma and chroma components. Once the availability (availability parameter) of each component is determined, information about the number of luma (component) filters can be parsed. As an example, the maximum number of available filters can be set to 25. If at least one luma filter is signaled, the filter index information can be parsed / signaled for each filter ranging from 0 to the maximum number of filters (e.g., 25, which may alternatively be referred to as a category). This may imply that each category (i.e., from 0 to the maximum number of filters) is associated with a filter index. When a filter index flag is used for the filters in each category, a flag (e.g., `alf_luma_coeff_delta_flag`) can be parsed / signaled. Flags can be used to explain whether flag information related to the prediction of ALF brightness filter coefficient increment values (such as alf_luma_coeff_delta_prediction_flag) exists in the slice header or tile group header.
[0250] If the number of luminance filters signaled by the `alf_luma_num_filters_signalled_minus1` syntax element is greater than 0, and the value of the `alf_luma_coeff_delta_flag` syntax element is 0, then it means that the `alf_luma_coeff_delta_prediction_flag` syntax element exists in the slice header or tile group header, and its state can be evaluated. If the state of the `alf_luma_coeff_delta_prediction_flag` syntax element is 1, then it means that luminance filter coefficients are predicted from previous luminance (filter) coefficients. If the state of the `alf_luma_coeff_delta_prediction_flag` syntax element is 0, then it means that luminance filter coefficients are not predicted from the increment of previous luminance (filter) coefficients.
[0251] When delta filter coefficients (e.g., alf_luma_coeff_delta_abs) are encoded based on exponential Golomb code, the order k of the exponential Golomb (EG) code may need to be determined in order to decode the delta luminance filter coefficients (e.g., alf_luma_coeff_delta_abs). This information may be required to decode the filter coefficients. The order of the exponential Golomb code can be represented as EG(k). To determine EG(k), the alf_luma_min_eg_order_minus1 syntax element can be parsed / signaled. The alf_luma_min_eg_order_minus1 syntax element can be an entropy-coded syntax element. The alf_luma_min_eg_order_minus1 syntax element can indicate the minimum order of the EG used to decode the delta luminance filter coefficients. For example, the value of the alf_luma_min_eg_order_minus1 syntax element can be in the range of 0 to 6. After parsing / signaling the `alf_luma_min_eg_order_minus1` syntax element, the `alf_luma_eg_order_increase_flag` syntax element can be parsed / signaled. If the value of the `alf_luma_eg_order_increase_flag` syntax element is 1, this indicates that the order of the EG indicated by the `alf_luma_min_eg_order_minus1` syntax element increases by 1. If the value of the `alf_luma_eg_order_increase_flag` syntax element is 0, this indicates that the order of the EG indicated by the `alf_luma_min_eg_order_minus1` syntax element does not increase. The order of the EG can be represented by the index of the EG. For example, the EG order (or EG index) based on the `alf_luma_min_eg_order_minus1` syntax element and the `alf_luma_eg_order_increase_flag` syntax element (related to the luma component) can be determined as follows.
[0252] [Table 11]
[0253]
[0254] Based on the above determination process, expGoderY can be derived as xpGoderY = KminTab. In this way, an array including the EG order can be derived, which can be used by the decoding device. expGoOrderY indicates the EG order (or EG index).
[0255] A predefined Columbus order index (i.e., golombOrderIdxY) may exist. The predefined Columbus order can be used to determine the final Columbus order used to encode the coefficients.
[0256] For example, the predefined Columbus order can be configured as shown in the table below.
[0257] golombOrderIdxY[]={0, 0, 1, 0, 1, 2, 1, 0, 0, 1, 2}
[0258] Here, the order k = expGoderY[GolombOrderIdxY[j]], and j can be indicated as the j-th filter coefficient notified by the signal. For example, if j = 2, that is, the third filter coefficient golomborderIdxY[2] = 1, then k = expGoderY[1].
[0259] In this case, for example, if the value of the alf_luma_coeff_delta_flag syntax element is true, i.e., 1, then for each filter notified by a signal, the alf_luma_coeff_flag syntax element can be used to signal the luminance filter coefficients. The alf_luma_coeff_flag syntax element indicates whether the luminance filter coefficients are (explicitly) notified by a signal.
[0260] When the EG order and the states of the aforementioned related flags (e.g., alf_luma_coeff_delta_flag, alf_luma_coeff_flag, etc.) are determined, the difference and sign information of the luminance filter coefficients can be parsed / signed (i.e., when alf_luma_coeff_flag indicates true). The absolute value of the increment for each of the 12 filter coefficients can be parsed / signed (alf_luma_coeff_delta_abs syntax element). Additionally, if the alf_luma_coeff_delta_abs syntax element has a value, the sign information can be parsed / signed (alf_luma_coeff_delta_sign syntax element). This information, including the difference and sign information of the luminance filter coefficients, can be referred to as information about the luminance filter coefficients.
[0261] The increments of the filter coefficients can be determined and stored along with their signs. In this case, the increments of the signed filter coefficients can be stored as an array, which can be represented as `filterCoefficient`. The increments of the filter coefficients can be called increment luminance coefficients, and the increments of the signed filter coefficients can be called signed increment luminance coefficients.
[0262] To determine the final filter coefficients from the signed incremental luminance coefficients, the (luminance) filter coefficients can be updated as follows.
[0263] filterCoefficients[sigFiltIdx][j]+=filterCoefficients[sigFiltIdx][j]
[0264] Here, j can indicate the filter coefficient index, and sigFiltIdx can indicate the filter index notified by a signal. j = {0, ..., 11} and sigFiltIdx = {0, ..., alf_luma_filters_signaled_minus1}.
[0265] The coefficients can be copied into the final AlfCoeffl[filtIdx][j]. Here, filtidx = 0, ..., 24, and j = 0, ..., 11.
[0266] The signed incremental luminance coefficients of a given filter index can be used to determine the first 12 filter coefficients. For example, the thirteenth filter coefficient of a 7×7 filter can be determined based on the following equation. The thirteenth filter coefficient can represent the center-tap filter coefficients mentioned above.
[0267] [Equation 1]
[0268] AlfCoeff L [filtIdx]
[12] =128-∑ k AlfCoefff L [filtIdx][k]<<1
[0269] Here, filter coefficient index 12 indicates the thirteenth filter coefficient. For reference, since filter coefficient indices start from 0, the value 12 indicates the thirteenth filter coefficient.
[0270] For example, to ensure bitstream consistency, when k is 0,…,11, the final filter coefficient AlfCoeffl[filtIdx][k] ranges from -2. 7 to 2 7 -1, and when k is 12, it can range from 0 to 2. 8 -1. Here, k can be replaced with j.
[0271] If processing is performed for the luma component, processing for the chroma component can be performed based on the `alf_chroma_idc` syntax element. If the value of the `alf_chroma_idc` syntax element is greater than 0, the minimum EG order information for the chroma component can be parsed / signed (e.g., the `alf_chroma_min_eg_order_minus1` syntax element). According to the embodiments described above in this document, a 5×5 diamond filter shape can be used for the chroma component, and in this case, the maximum Golomb index can be 2. In this case, the EG order (or EG index) of the chroma component can be determined, for example, as follows.
[0272] [Table 12]
[0273]
[0274] Based on the deterministic process, expGoOrderC can be derived as expGoOrderC = KminTab. Therefore, it is possible to deduce the array containing EG orders that the decoding device can use. expGoOrderC indicates the EG order (or EG index) of the chroma components.
[0275] A predefined Golomb order index (golombOrderIdxC) may exist. The predefined Golomb order can be used to determine the final Golomb order used to encode the coefficients.
[0276] For example, the predefined Columbus order can be configured as shown in the table below.
[0277] golombOrderIdxC[]={0, 0, 1, 0, 0, 1}
[0278] Here, the order k = expGoOrderC[golombOrderIdxC[j]], and j can represent the j-th filter coefficient notified by the signal. For example, if j = 2, that is, the third filter coefficient golomboorderIdxY[2] = 1, then k = expGoOrderC[1].
[0279] Based on this, the absolute value information and sign information of the chroma filter coefficients can be parsed / signed. This information, including the absolute value information and sign information of the chroma filter coefficients, can be referred to as information about the chroma filter coefficients. For example, a 5×5 diamond filter shape can be applied to the chroma components, and in this case, the absolute increment information (alf_chroma_coeff_abs syntax element) of each of the six (chroma component) filter coefficients can be parsed / signed. Additionally, if the value of the alf_chroma_coeff_abs syntax element is greater than 0, the sign information (alf_chroma_coeff_abs syntax element) can be parsed / signed. For example, the six chroma filter coefficients can be derived based on the information about the chroma filter coefficients. In this case, for example, the seventh chroma filter coefficient can be determined based on the following equation. The seventh filter coefficient can represent the aforementioned center-tap filter coefficient.
[0280] [Equation 2]
[0281] AlfCoeff C [6]=128-∑ k AlfCoeff C [filtIdx][k]<<1
[0282] Here, filter coefficient index 6 indicates the seventh filter coefficient. For reference, since filter coefficient indices start from 0, the value 6 indicates the seventh filter coefficient.
[0283] For example, to ensure bitstream consistency, when k is 0, ..., 5, the final filter coefficient AlfCoeffc[filtIdx][k] ranges from -2. 7 to 2 7 -1, while when k is 6, it can range from 0 to 2. 8 -1. Here, k can be replaced with j.
[0284] If the (luminance / chrominance) filter coefficients are derived, ALF-based filtering can be performed based on the filter coefficients or a filter including the filter coefficients. As described above, this is how modified reconstructed samples can be derived. Alternatively, multiple filters can be derived, and the filter coefficients of one of these filters can be used in the ALF process. As an example, one of the multiple filters can be indicated based on filter selection information notified by a signal. Or, for example, one of the multiple filters can be selected based on the activity and / or directionality of the current block or the ALF target block, and the filter coefficients of the selected filter can be used in the ALF process.
[0285] Furthermore, as mentioned above, Luminance Mapping and Chroma Scaling (LMCS) can be applied to improve coding efficiency. LMCS can be referred to as a loop shaper (shaper). To increase coding efficiency, LMCS control and / or LMCS-related signaling can be performed in layers.
[0286] Figure 15 An exemplary hierarchical structure of a CVS according to an embodiment of this document is illustrated. A coded video sequence (CVS) may include a sequence parameter set (SPS), a picture parameter set (PPS), a tile group header, tile data, and / or a CTU. Here, the tile group header and tile data may be referred to as a slice header and slice data, respectively.
[0287] The SPS can originally include flags for enabling the tool to be used in CVS. Additionally, the SPS can be a PPS reference including information about parameters that change for each image. Each encoded image can include tiles of one or more encoded rectangular fields. Tiles can be grouped by raster scanning to form tile groups. Each tile group is encapsulated with header information called a tile group header. Each tile consists of a CTU containing encoded data. Here, the data can include raw sample values, predicted sample values, and their luminance and chrominance components (luminance predicted sample values and chrominance predicted sample values).
[0288] Figure 16 An exemplary LMCS structure is shown according to an embodiment of this document. Figure 16 The LMCS structure 1600 may include an in-loop mapping portion 1610 for the luminance component based on an adaptive piecewise linear (adaptive PWL) model and a luminance-dependent chrominance residual scaling portion 1620 for the chrominance component. The blocks of dequantization and inverse transform 1611, reconstruction 1612, and intra-frame prediction 1613 of the in-loop mapping portion 1610 represent the processing applied in the mapped (shaping) domain. The blocks of loop filter 1615 and motion compensation or inter-frame prediction 1617 of the in-loop mapping portion 1610, and the blocks of reconstruction 1622, intra-frame prediction 1623, motion compensation or inter-frame prediction 1624, and loop filter 1625 of the chrominance residual scaling portion 1620 represent the processing applied in the original (unmapped, unshaping) domain.
[0289] like Figure 16As shown, when LMCS is enabled, at least one of inverse shaping (mapping) processing 1614, forward shaping (mapping) processing 1618, and chroma scaling processing 1621 can be applied. For example, inverse shaping processing can be applied to (reconstructed) luminance samples (or multiple luminance samples or arrays of luminance samples) of a reconstructed image. Inverse shaping processing can be performed based on the piecewise (inverse) index of the luminance samples. The piecewise (inverse) index identifies the segment (or portion) to which the luminance sample belongs. The output of inverse shaping processing is a modified (reconstructed) luminance sample (or multiple modified luminance samples or arrays of modified luminance samples). LMCS can be enabled or disabled at tile groups (or slices), images, or higher levels.
[0290] Forward shaping and / or chroma scaling can be applied to generate reconstructed images. Images can include luminance samples and chroma samples. A reconstructed image with luminance samples can be called a reconstructed luminance image, and a reconstructed image with chroma samples can be called a reconstructed chroma image. A combination of a reconstructed luminance image and a reconstructed chroma image can be called a reconstructed image. Reconstructed luminance images can be generated based on forward shaping. For example, if inter-frame prediction is applied to the current block, forward shaping is applied to luminance prediction samples derived from (reconstructed) luminance samples of a reference image. Since the (reconstructed) luminance samples of the reference image are generated based on inverse shaping, forward shaping can be applied to the luminance prediction samples to derive shaped (mapped) luminance prediction samples. Forward shaping can be performed based on the piecewise function index of the luminance prediction samples. The piecewise function index can be derived based on the value of the luminance prediction samples or the value of the luminance samples of the reference image used for inter-frame prediction. Reconstructed samples can be generated based on (shaped / mapped) luminance prediction samples. Inverse shaping (mapping) can be applied to the reconstructed samples. Reconstructed samples that have undergone inverse shaping (mapping) processing can be referred to as inverse-shaping (mapping) reconstructed samples. Furthermore, inverse-shaping (mapping) reconstructed samples can be simply referred to as shaping (mapping) reconstructed samples. When intra-frame prediction (or intra-block copying (IBC)) is applied to the current block, a forward mapping of the current block's prediction samples may not be necessary because inverse shaping has not yet been applied to the reconstructed samples referencing the current image. In the reconstructed luminance image, (reconstructed) luminance samples can be generated based on (shaping) luminance prediction samples and the corresponding luminance residual samples.
[0291] A reconstructed chroma image can be generated based on chroma scaling. For example, it can be generated based on chroma prediction samples and chroma residual samples (c) in the current block. res This is used to derive the (reconstructed) chroma samples in the reconstructed chroma image. Based on the (scaled) chroma residual samples (c) used for the current block... resScale The chromaticity residual sample (c) is derived using the chromaticity residual scaling factor (cScaleInv, which can be called varScale) and the chromaticity residual scaling factor (cScaleInv, which can be called varScale). resThe chromaticity residual scaling factor can be calculated based on the shaped luminance prediction sample values in the current block. For example, it can be based on the shaped luminance prediction sample values (Y'). pred The average brightness value of (ave(Y') pred To calculate the scaling factor. For reference, based on... Figure 16 The (scaled) chromaticity residual sample derived from the inverse transform / dequantization in C can be referred to as c. rescale Furthermore, the chromaticity residual sample derived by performing (inverse) scaling on the (scaled) chromaticity residual sample can be referred to as c. res .
[0292] Figure 17 This document illustrates an LMCS structure according to another embodiment of this document. Reference will be made to... Figure 16 describe Figure 17 Here, we will mainly describe Figure 17 LMCS structure and Figure 16 The differences between the LMCS structures. Figure 17 The in-ring mapping part and the luminance-dependent chrominance residual scaling part can be compared with... Figure 16 The in-ring mapping part and the luminance-dependent chrominance residual scaling part operate in the same way.
[0293] refer to Figure 17 The chromaticity residual scaling factor can be derived based on the luminance reconstruction samples. In this case, the average luminance value (avgY) can be obtained based on the luminance reconstruction samples of neighboring luminance outside the reconstruction block rather than the luminance reconstruction samples inside the reconstruction block. r ), and can be based on average brightness value (avgY) r The chroma residual scaling factor is derived using a forward mapping method. Here, the adjacent luminance reconstruction sample can be an adjacent luminance reconstruction sample of the current block, or it can be an adjacent luminance reconstruction sample that includes the Virtual Pipeline Data Unit (VPDU) of the current block. For example, when intra-frame prediction is applied to the target block, the reconstruction sample can be derived based on the prediction sample, where the prediction sample is derived based on intra-frame prediction. Furthermore, for example, when inter-frame prediction is applied to the target block, the forward mapping is applied to the prediction sample derived based on inter-frame prediction, and the reconstruction sample can be generated based on the shaped (or forward-mapped) luminance prediction sample.
[0294] Motion picture / image information signaled via a bitstream may include LMCS parameters (information about the LMCS). LMCS parameters can be configured as a high-level syntax (HLS, including slice header syntax), etc. A detailed description of LMCS parameters and configuration will be described later. As described above, the syntax table described in this document (and the embodiments below) can be constructed / encoded on the encoder (encoding device) side and signaled to the decoder (decoding device) via a bitstream. The decoder can parse / decode the LMCS information in the syntax table (in the form of syntax components). One or more embodiments described below can be combined. The encoder can encode the current image based on the information about the LMCS, and the decoder can decode the current image based on the information about the LMCS.
[0295] Intra-loop mapping of the luma component can improve compression efficiency by redistributing codewords across the dynamic range to adjust the dynamic range of the input signal. For luma mapping, a forward mapping (shaping) function (FwdMap) and its corresponding inverse mapping (shaping) function (InvMap) can be used. The forward mapping function (FwdMap) can be signaled using a partially linear model, for example, a partially linear model with 16 segments or bins. These segments can have the same length. In one example, the inverse mapping function (InvMap) may not be signaled separately, but can be derived from the forward mapping function (FwdMap). That is, the inverse mapping can be a function of the forward mapping. For example, the inverse mapping function can be a function where the forward mapping function is symmetric about y = x.
[0296] In-loop (luminance) shaping can be used to map input luminance values (samples) to changed values in the shaping domain. After reconstruction, the shaped values can be encoded and mapped back to the original (unmapped, unshaped) domain. Chroma residual scaling can be applied to compensate for the difference between the luminance and chrominance signals. In-loop shaping can be performed by specifying a high-level syntax for the shaper model. The shaper model syntax can be signaled to the partially linear model (PWL model). Forward lookup tables (FwLUTs) and / or backward lookup tables (InvLUTs) can be derived based on the partially linear model. As an example, when deriving a forward lookup table (FwLUT), a backward lookup table (InvLUT) can be derived based on the forward lookup table (FwLUT). The forward lookup table (FwLUT) maps the input luminance value Yi to the changed value Yr, and the backward lookup table (InvLUT) maps the reconstructed value Yr to the reconstructed value Yr based on the changed value. ' i.
[0297] In one example, SPS may include the syntax shown in Table 13 below. The syntax in Table 13 may include `sps_reshaper_enabled_flag` as a tool enabling flag. Here, `sps_reshaper_enabled_flag` can be used to specify whether a shaper is used in the encoded video sequence (CVS). That is, `sps_reshaper_enabled_flag` can be a flag used to enable shaping in SPS. In one example, the syntax in Table 13 may be part of SPS.
[0298] [Table 13]
[0299]
[0300] In one example, the semantics that can be indicated by sps_seq_parameter_set_id and sps_reshaper_enabled_flag are shown in Table 14 below.
[0301] [Table 14]
[0302]
[0303] In one instance, the tile group header or slice header may include the syntax of Table 15 or Table 16 below.
[0304] [Table 15]
[0305]
[0306] [Table 16]
[0307]
[0308] The semantics of the grammatical elements included in the grammar of Table 15 or Table 16 may include, for example, what is disclosed in the following tables.
[0309] [Table 17]
[0310]
[0311] [Table 18]
[0312]
[0313]
[0314] As an example, if `sps_reshaper_enabled_flag` is parsed, additional data for configuring lookup tables (FwdLUT and / or InvLUT) (e.g., information included in Tables 15 or 16 above) can be parsed in the tile group header. For this purpose, the state of the SPS shaper flag can be checked in the slice header or tile group header. When `sps_reshaper_enabled_flag` is true (or 1), the additional flag `tile_group_reshaper_model_present_flag` (or `slice_reshaper_model_present_flag`) can be parsed. The purpose of `tile_group_reshaper_model_present_flag` (or `slice_reshaper_model_present_flag`) can be to indicate the presence of a shaper model. For example, when `tile_group_reshaper_model_present_flag` (or `slice_reshaper_model_present_flag`) is true (or 1), it can indicate that a shaper exists for the current tile group (or the current slice). When tile_group_reshaper_model_present_flag (or slice_reshaper_model_present_flag) is false (or 0), it indicates that there is no shaper for the current tile group (or the current slice).
[0315] If a shaper exists and is enabled in the current tile group (or slice), the shaper model (e.g., `tile_group_reshaper_model()` or `slice_reshaper_model()`) can be processed, resolving the `tile_group_reshaper_enable_flag` (or `slice_reshaper_enable_flag`) in addition to additional flags. `tile_group_reshaper_enable_flag` (or `slice_reshaper_enable_flag`) indicates whether the shaper model is used in the current tile group (or slice). For example, if `tile_group_reshaper_enable_flag` (or `slice_reshaper_enable_flag`) is 0 (or false), it indicates that the shaper model is not used in the current tile group (or slice). If tile_group_reshaper_enable_flag (or slice_reshaper_enable_flag) is 1 (or true), it indicates that the shaper model can be used for the current tile group (or the current slice).
[0316] As an example, `tile_group_reshaper_model_present_flag` (or `slice_reshaper_model_present_flag`) can be true (or 1), and `tile_group_reshaper_enable_flag` (or `slice_reshaper_enable_flag`) can be false (or 0). This means that an integer model exists but is not used in the current tile group (or slice). In this case, the integer model can be used in the following tile groups (or slices). As another example, `tile_group_reshaper_enable_flag` can be true (or 1), and `tile_group_reshaper_model_present_flag` can be false (or 0).
[0317] When resolving the reshaper model (e.g., `tile_group_reshaper_model()` or `slice_reshaper_model()`) and `tile_group_reshaper_enable_flag` (or `slice_reshaper_enable_flag`), the conditions required for chroma scaling can be determined (evaluated). These conditions can include condition 1 (the current tile group / slice is not intra-coded) and / or condition 2 (the current tile group / slice is not split into two separate coded quadtree structures for luma and chroma, i.e., the current tile group / slice is not a binary tree structure). If condition 1 and / or condition 2 are true and / or `tile_group_reshaper_enable_flag` (or `slice_reshaper_enable_flag`) is true (or 1), then `tile_group_reshaper_chroma_residual_scale_flag` (or `slice_reshaper_chroma_residual_scale_flag`) can be resolved. When `tile_group_reshaper_chroma_residual_scale_flag` (or `slice_reshaper_chroma_residual_scale_flag`) is enabled (if 1 or true), chroma residual scaling is enabled for the current tile group (or the current slice). When `tile_group_reshaper_chroma_residual_scale_flag` (or `slice_reshaper_chroma_residual_scale_flag`) is disabled (if 0 or false), chroma residual scaling is disabled for the current tile group (or the current slice).
[0318] The purpose of the above shaping is to parse the data needed to construct the lookup table (FwdLUT and / or InvLUT). In the example, the lookup table constructed based on the parsed data can divide the distribution of the allowable range of luminance values into multiple bins (e.g., 16). Therefore, luminance values in a given bin can be mapped to changed luminance values.
[0319] Figure 18 A diagram illustrating an exemplary forward mapping is shown. Figure 12 The five bins are shown here only as examples.
[0320] refer to Figure 18The x-axis represents the input luminance value, and the y-axis represents the changed output luminance value. The x-axis is divided into 5 bins or segments, each of length L; that is, the five bins mapped to the changed luminance value have the same length. A forward lookup table (FwdLUT) can be constructed using data available from the tile group header that facilitates the mapping (e.g., shaper data).
[0321] In an embodiment, an output pivot point associated with a bin index can be calculated. The output pivot point sets the minimum and maximum boundaries of the output range for luminance codeword shaping. The calculation of the output pivot point can be performed based on a piecewise cumulative distribution function of the number of codewords. The output pivot range can be partitioned based on the maximum number of bins to be used and the size of the lookup table (FwdLUT or InvLUT). As an example, the output pivot range can be partitioned based on the product between the maximum number of bins and the size of the lookup table. For example, if the product between the maximum number of bins and the size of the lookup table is 1024, the output pivot range can be partitioned into 1024 entries. The partitioning of the output pivot range can be performed (applied or implemented) based on a scaling factor. In one example, the scaling factor can be derived based on Equation 3 below.
[0322] [Equation 3]
[0323] SF=(y2-y1)*(1< <FP_PREC)+c
[0324] In Equation 3, SF represents the scaling factor, and y1 and y2 represent the output pivot points corresponding to each bin. Furthermore, FP_PREC and c can be predetermined constants. The scaling factor determined based on Equation 3 can be referred to as the scaling factor used for forward shaping.
[0325] In another embodiment, for inverse integer shaping (inverse mapping), for a range of bins (e.g., reshaper_model_min_bin_idx to reshaper_model_max_bin_idx), the pivot point of the input integer corresponding to the pivot point of the mapping in the forward lookup table (FwdLUT) and the inverse output pivot point of the mapping are obtained (given as the initial codeword number * bin index). In another example, the scaling factor SF can be derived based on Equation 4 below.
[0326] [Equation 4]
[0327] SF=(y2-y1)*(1< <FP_PREC) / (x2-x1)
[0328] In Equation 4, SF represents the scaling factor, x1 and x2 represent the input pivot points, and y1 and y2 represent the output pivot points corresponding to each bin. Here, the input pivot points can be pivot points mapped based on a forward lookup table (FwdLUT), while the output pivot points can be pivot points mapped based on a reverse lookup table (InvLUT). Furthermore, FP_PREC can be a predetermined constant. The FP_PREC of Equation 4 can be the same as or different from the FP_PREC of Equation 3, and the scaling factor determined based on Equation 4 can be referred to as the scaling factor used for inverse shaping. During inverse shaping, the input pivot points can be partitioned based on the scaling factor of Equation 4, and based on the partitioned input pivot points, pivot values corresponding to the minimum and maximum bin values from 0 to the minimum bin index (reshaper_model_min_bin_idx) and / or from the minimum bin index (reshaper_model_min_bin_idx) to the maximum bin index (reshaper_model_max_bin_idx) are specified.
[0329] Table 19 below illustrates the syntax of the shaper model according to an embodiment. This shaper model may be referred to as an LMCS model. Here, the shaper model is exemplarily described as a piece group shaper, but this specification is not necessarily limited to this embodiment. For example, the shaper model may be included in an APS, or the piece group shaper model may be referred to as a slice shaper model.
[0330] [Table 19]
[0331]
[0332] The semantics of the grammatical elements included in the grammar of Table 19 may include, for example, what is disclosed in the following table.
[0333] [Table 20]
[0334]
[0335] The shaper model comprises components as elements, such as reshape_model_min_bin_idx, reshape_model_delta_max_bin_idx, reshaper_model_bin_delta_abs_cw_prec_minus1, reshape_model_bin_delta_abs_CW[i], and reshaper_model_bin_delta_sign_CW. Each component will be described in detail below.
[0336] `reshape_model_min_bin_idx` indicates the minimum bin (or segment) index used during shaper configuration. The value of `reshape_model_min_bin_idx` can range from 0 to `MaxBinIdx`. For example, `MaxBinIdx` can be 15.
[0337] In an embodiment, the tile group shaper model can preferably resolve two indices (or parameters), namely, `reshaper_model_min_bin_idx` and `reshaper_model_delta_max_bin_idx`. The maximum bin index (`reshaper_model_max_bin_idx`) can be derived (determined) based on these two indices. `reshaper_model_delta_max_bin_idx` can represent the maximum allowed bin index `MaxBinIdx` subtracted from the actual maximum bin index (`reshaper_model_max_bin_idx`) used during shaper configuration. The value of the maximum bin index (`reshaper_model_max_bin_idx`) can be in the range from 0 to `MaxBinIdx`. For example, `MaxBinIdx` can be 15. As an example, the value of `reshaper_model_max_bin_idx` can be derived based on Equation 5 below.
[0338] [Equation 5]
[0339] reshape_model_max_bin_idx=MaxBinldx-reshape_model_delta_max_bin_idx.
[0340] The maximum bin index (reshaper_model_max_bin_idx) can be greater than or equal to the minimum bin index (reshaper_model_min_bin_idx). The minimum bin index can be referred to as the minimum allowed bin index or the minimum allowed bin index, and the maximum bin index can also be referred to as the maximum allowed bin index or the maximum allowed bin index.
[0341] If the maximum bin index (reshaper_model_max_bin_idx) is derived, the syntax element reshaper_model_bin_delta_abs_cw_prec_minus1 can be resolved. The number of bits used to represent the syntax reshape_model_bin_delta_abs_CW[i] can be determined based on reshaper_model_bin_delta_abs_cw_prec_minus1. For example, the number of bits used to represent reshape_model_bin_delta_abs_CW[i] can be the same as reshaper_model_bin_delta_abs_cw_prec_minus1 plus 1.
[0342] `reshape_model_bin_delta_abs_CW[i]` indicates information related to the absolute increment codeword value (the absolute value of the increment codeword) of the i-th bin. In one example, if the absolute increment codeword value of the i-th bin is greater than 0, then `reshaper_model_bin_delta_sign_CW_flag[i]` can be resolved. The sign of `reshape_model_bin_delta_abs_CW[i]` can be determined based on `reshaper_model_bin_delta_sign_CW_flag[i]`. In one example, if `reshaper_model_bin_delta_sign_CW_flag[i]` is 0 (or false), then the corresponding variable `RspDeltaCW[i]` can be positive. In other cases (if reshaper_model_bin_delta_sign_CW_flag[i] is not 0, or if reshaper_model_bin_delta_sign_CW_flag[i] is 1 (or true)), the corresponding variable RspDeltaCW[i] can be negative. If reshaper_model_bin_delta_sign_CW_flag[i] does not exist, it can be inferred to be 0 (or false).
[0343] In the embodiment, the variable RspDeltaCW[i] can be derived based on the above reshape_model_bin_delta_abs_CW[i] and reshape_model_bin_delta_sign_CW_flag[i]. RspDeltaCW[i] can be referred to as the value of the incremental codeword. For example, RspDeltaCW[i] can be derived based on Equation 6 below.
[0344] [Equation 6]
[0345] RspDeltaCW[i] = (1 - 2 * reshape_model_bin_delta_sign_CW[i]) * reshape_model_bin_delta_abs_CW[i]
[0346] In Equation 6, reshape_model_bin_delta_sign_CW[i] can be information related to the sign of RspDeltaCW[i]. For example, reshape_model_bin_delta_sign_CW[i] can be the same as the above-mentioned reshaper_model_bin_delta_sign_CW_flag[i]. Here, i can range from the minimum bin index (reshaper_model_min_bin_idx) to the maximum bin index (reshape_model_max_bin_idx).
[0347] The variable (or array) RspCW[i] can be derived based on RspDeltaCW[i]. Here, RspCW[i] can represent the number of codewords allocated (distributed) to the i-th bin. That is, the number of codewords allocated (distributed) to each bin can be stored in array form. In one example, if i is less than the above-mentioned reshaper_model_min_bin_idx or greater than reshaper_model_max_bin_idx, (i < reshaper_model_min_bin_idx or reshaper_model_max_bin_idx < i), then RspCW[i] can be 0. Otherwise (if i is greater than or equal to the above-mentioned reshaper_model_min_bin_idx and less than or equal to reshaper_model_max_bin_idx, (reshaper_model_min_bin_idx <= i <= reshaper_model_max_bin_idx)), then RspCW[i] can be derived based on the above RspDeltaCW[i], the luminance bit depth (BitDepthsy), and / or MaxBinIdx. In this case, for example, RspCW[i] can be derived based on the following Equation 7.
[0348] [Equation 7]
[0349] RspCW[i] = OrgCW + RspDeltaCW[i]
[0350] In Equation 7, OtgCW can be a predetermined value, for example, it can be determined based on Equation 8 below.
[0351] [Equation 8]
[0352] OrgCW = (1 < <BitDepth Y ) / (MaxBinIdx+1)
[0353] In Equation 8, BitDepthY is the luminance bit depth, and MaxBinIdx represents the maximum allowed bin index. In one example, if BitDepthY is 10, then RspCW[i] can have values from 32 to 2*OrgCW-1.
[0354] The variable InputPivot[i] can be derived based on the OrgCW above. For example, InputPivot[i] can be derived based on Equation 9 below.
[0355] [Equation 9]
[0356] InputPivot[i] = i * OrgCW
[0357] The variables ReshapePivot[i], ScaleCoef[i], and / or InvScaleCoeff[i] can be derived based on the above RspCW[i], InputPivot[i], and / or OrgCW. For example, ReshapePivot[i], ScaleCoef[i], and / or InvScaleCoeff[i] can be derived based on Table 21 below.
[0358] [Table 21]
[0359]
[0360] In Table 21, a for loop syntax can be used, where i increments from 0 to MaxBinIdx, and shiftY can be a predetermined constant for bit shifting. Whether InvScaleCoeff[i] is based on RspCW[i] can be deduced based on a conditional clause stating whether RspCW[i] is 0.
[0361] The ChromaScaleCoef[i] used to derive the chroma residual scaling factor can be derived based on Table 22 below.
[0362] [Table 22]
[0363]
[0364] In Table 22, ShiftC can be a predetermined constant used for bit shifting. Referring to Table 22, it can be deduced whether ChromaScaleCoef[i] is based on the array ChromaResidualScalLut based on the conditional clause of whether RspCW[i] is 0. Here, ChromaResidualScalLut can be a predetermined array. However, the ChromaResidualScalLut array is merely exemplary, and this embodiment is not necessarily limited to Table 22.
[0365] The method for deriving the i-th variable has been described above. The (i+1)-th variable can be based on ReshapePivot[i+1], and for example, ReshapePivot[i+1] can be derived based on Equation 10.
[0366] [Equation 10]
[0367] ReshapePivot[i+1]=ReshapePivot[i]+RspCW[i]
[0368] In Equation 10, RspCW[i] can be derived based on Equations 7 and / or 8 above. Luminance mapping can be performed based on the above embodiments and examples, and the above syntax and the components included therein are merely exemplary representations, and the embodiments are not limited to the above tables or equations. Hereinafter, a method for performing chroma residual scaling (scaling of the chroma components of the residual samples) based on luminance mapping will be described.
[0369] (Luminance-dependent) Chroma residual scaling is used to compensate for the difference between luminance samples and their corresponding chrominance samples. For example, chroma residual scaling can be signaled at the tile group or slice level to indicate whether it is enabled. In one example, if luminance mapping is enabled and dual-tree segmentation is not applied to the current tile group, a signal can be used to attach a flag indicating whether luminance-dependent chroma residual scaling is enabled. In another example, luminance-dependent chroma residual scaling can be disabled if luminance mapping is not used, or if dual-tree segmentation is not used for the current tile group. In yet another example, chroma residual scaling can always be disabled for chrominance tiles with a size less than or equal to 4.
[0370] Chroma residual scaling can be performed based on the average luminance value of a reference sample. For example, the reference sample may include samples corresponding to a luminance prediction block (the luminance component of the prediction block to which intra-frame prediction and / or inter-frame prediction are applied). When inter-frame prediction is applied, the reference sample may include samples after the forward mapping is applied to the luminance component prediction samples. Alternatively, the reference sample may include neighboring samples of the current block or neighboring samples of the VPDU including the current block. In this case, when inter-frame prediction is applied to neighboring blocks including neighboring samples, the neighboring samples may include luminance component reconstruction samples derived based on the luminance component prediction samples to which the forward mapping of the neighboring blocks is applied. The scaling operation on the encoder side and / or decoder side can be implemented, for example, based on fixed-point integer calculations according to Equation 11 below.
[0371] [Equation 11]
[0372] c'=sign(c)*((abs(c)*s+2CSCALE_FP_PREC-1)>>>CSCALE_FP_PREC)
[0373] In Equation 11 above, c' represents the scaled chroma residual sample (the scaled chroma component of the residual sample), c represents the chroma residual sample (the chroma component of the residual sample), s can represent the chroma residual scaling factor, and CSCALE_FP_PREC can represent a predetermined constant.
[0374] As mentioned above, the average luminance value of the reference sample can be obtained, and the chromaticity residual scaling factor can be derived based on the average luminance value. As stated above, the chromaticity component residual sample can be scaled based on the chromaticity residual scaling factor, and the chromaticity component reconstructed sample can be generated based on the scaled chromaticity component residual sample.
[0375] The embodiments in this document propose a signaling structure for effectively applying the aforementioned ALF and / or LMCS. According to the embodiments in this document, for example, ALF data and / or LMCS data can be included in an HLS (e.g., an APS), and a filter for the ALF and / or LMCS model (shaper model) can be adaptively derived by signaling the referenced APS ID via header information (e.g., image header, slice header) that serves as a lower level of the APS. The LMCS model can be derived based on LMCS parameters. Furthermore, for example, multiple APS IDs can be signaled via header information, and in this way, different ALF and / or LMCS models can be applied on a block-by-block basis within the same image / slice.
[0376] For example, according to an embodiment of this document, an APS may carry ALF data and / or LMCS data. In this case, whether a corresponding APS carries ALF data or LMCS data can be indicated by APS type information (e.g., the aps_params_type syntax element). As mentioned above, LMCS data may be mixed or referred to as shaper data, and may carry LMCS / shaper parameters for deriving the LMCS / shaper model. In this case, multiple APSs can be signaled that the first APS carries ALF data, while the second APS carries LMCS (shaper) data. The first and second APSs can be identified at a lower level (e.g., header information, CTU, or CU, etc.) based on the APS ID.
[0377] For example, the following shows an example of an APS according to one embodiment.
[0378] [Table 23]
[0379]
[0380] Referring to Table 23, as an example, APS type information (e.g., aps_params_type) can be parsed / signed out in the APS. APS type information can be referred to as APS parameter type information. APS type information can indicate whether the corresponding APS carries ALF data or LMCS data. As an example, APS type information can be parsed / signed out after adaptation_parameter_set_id.
[0381] For example, the aps_params_type, ALF_APS, and LMCS_APS included in Table 23 can be described according to the following table. That is, ALF_APS, determined by aps_params_type, can indicate the APS carrying ALF data, and LMCS_APS can indicate the APS carrying LMCS data.
[0382] [Table 24]
[0383]
[0384] Referring to Table 24, for example, the APS type information (aps_params_type) can be a syntax element used to classify the type of the corresponding APS. When the value of the APS type information (aps_params_type) is 0, the type of the corresponding APS can be ALF_APS, the corresponding APS can carry ALF data, and the ALF data can include ALF parameters for deriving the filter / filter coefficients. When the value of the APS type information (aps_params_type) is 1, the type of the corresponding APS can be LMCS_APS, the corresponding APS can carry LMCS data, and the LMCS data can include LMCS parameters for deriving the LMCS model / bin / mapping index.
[0385] Furthermore, for example, as described above, ALF or LMCS data included in the APS can be accessed via header information. In this way, filters for ALF and / or LMCS models can be adaptively derived on a per-image / slice / CTU / CU basis. Additionally, multiple APS IDs can be signaled for the header information, thereby enabling the application of different ALF and / or LMCS models on a per-block basis within the same image / slice.
[0386] For example, the following shows an example of header information according to an embodiment. The header information may include a slice header and / or an image header.
[0387] [Table 25]
[0388]
[0389] [Table 26]
[0390]
[0391] The semantics of the grammatical elements included in the grammars in Tables 25 and 26 may include, for example, what is disclosed in the following table.
[0392] [Table 27]
[0393]
[0394]
[0395] [Table 28]
[0396]
[0397]
[0398] For example, referring to Tables 23 to 28, the header information may include ALF-related information. ALF-related information includes ALF availability flags (e.g., the slice_ALF_enabled_flag syntax element or the ph_ALF_enabled_flag syntax element), ALF-related APS ID quantity information (e.g., the num_ALF_APS_IDs_minus1 syntax element or the ph_num_ALF_APS_IDs_luma syntax element), and ALF-related APS ID syntax elements (e.g., slice_ALF_APS_ID[i] or ph_ALF_APS_ID_luma[i]), the number of which is the same as the number of ALF-related APS IDs derived from the ALF-related APS ID quantity information.
[0399] Additionally, the header information may include LMCS-related information. LMCS-related information may include at least one of the following: LMCS availability flag information (e.g., the slice_lmcs_enabled_flag syntax element or the ph_lmcs_enabled_flag syntax element), LMCS-related APS ID information (the slice_lmcs_aps_id syntax element or the ph_lmcs_aps_id syntax element), and chroma residual scaling flag information (the slice_chroma_residual_scale_flag syntax element or the ph_chroma_residual_scale_flag syntax element).
[0400] According to this embodiment, hierarchical control of ALF and / or LMCS is possible.
[0401] For example, as described above, the availability of ALF tools can be determined by the ALF availability flag in the SPS (e.g., the `sps_alf_enabled_flag` syntax element), and whether ALF is available in the current image or slice can be indicated by the ALF availability flag in the header information (e.g., the `slice_alf_enabeld_flag` syntax element or the `ph_alf_enabeld_flag` syntax element). When the value of the ALF availability flag in the header information is 1, the ALF-related APSID count syntax element can be parsed / signaled. Additionally, an ALF-related APS ID syntax element can be parsed / signaled as many times as the number of ALF-related APS IDs derived from the ALF-related APS ID count syntax element. That is, this indicates that multiple APSs can be parsed or referenced through a single header information.
[0402] Furthermore, for example, the availability of the LMCS (or reshaper) tool can be determined via an LMCS availability flag in the SPS (e.g., sps_reshaper_enabled_flag). sps_reshaper_enabled_flag can be referred to as sps_lmcs_enabled_flag. Whether LMCS is available in the current image or slice can be indicated by LMCS availability flag information in the header (e.g., the slice_lmcs_enabled_flag syntax element or the ph_lmcs_enabled_flag syntax element). When the value of the LMCS availability flag in the header is 1, the associated APS ID syntax element can be parsed / signaled. The LMCS model (shaper model) can be derived from the APS indicated by the associated APS ID syntax element. For example, the APS may also include an LMCS data field, and the LMCS data field may include the aforementioned LMCS model (shaper model) information.
[0403] Additionally, multiple APSs can be signaled, with the first APS carrying ALF data and the second APS carrying LMCS (Shaper) data. The first and second APSs can be identified by their APS IDs via header information. Furthermore, for example, the first APS can carry first ALF data, the second APS can carry second ALF data, and the third APS can carry LMCS (Shaper) data. The first and / or second APS can be referenced based on the ALF-related APS ID syntax elements in the header information, and the third APS can be referenced based on the LMCS-related APS ID syntax elements.
[0404] The following figures were created to illustrate specific examples of this specification. Since the names of particular devices or signals / messages / fields described in the figures are presented by way of example, the technical features of this specification are not limited to the specific names used in the following figures.
[0405] Figure 19 and Figure 20 Examples of video / image encoding methods and related components according to embodiments of this document are illustrated schematically. Figure 19 The method disclosed in the article can be derived from Figure 2 The encoding device shown in the diagram performs this operation. Specifically, for example, Figure 19 S1900 can be executed by the predictor 220 of the encoding device, S1910 can be executed by the residual processor 230 of the encoding device, S1920 can be executed by the adder 250 of the encoding device, S1930 can be executed by the filter 260 of the encoding device, and S1940 can be executed by the entropy encoder 240 of the encoding device. Figure 19The methods disclosed herein may include the embodiments described above.
[0406] Reference Figure 19 The encoding device determines the intra-prediction mode of the current block in the current image and derives prediction samples (S1900). The encoding device can derive prediction samples for the current block based on the prediction mode. In this case, various prediction methods disclosed in this document, such as inter-frame prediction or intra-frame prediction, can be applied. In this case, the encoding device can generate prediction mode information. The encoding device can generate prediction mode information indicating the prediction mode to be applied to the current block.
[0407] The encoding device derives residual samples based on the predicted samples (S1910). The encoding device can derive residual samples based on both the predicted samples and the original samples. In this case, residual information can be generated based on the residual samples. The residual information may include information about the aforementioned (quantized) transformation coefficients.
[0408] The encoding device generates reconstructed samples based on predicted samples (S1920). The encoding device can derive (modified) residual samples based on residual information. Reconstructed samples can be generated based on the (modified) residual samples and predicted samples. Reconstructed blocks and reconstructed images can be derived based on the reconstructed samples.
[0409] The encoding device derives ALF parameters (S1930). The encoding device can derive ALF parameters associated with the ALF, which can be applied to filter the reconstructed samples. ALF parameters can be included in the ALF data field. For example, the ALF data field can include the ALF parameters described above in this document. The ALF data field can include, for example, information indicating the filter or filter coefficients of the ALF.
[0410] Furthermore, although not publicly disclosed, the encoding device may derive LMCS parameters for the LMCS process. LMCS parameters may be included in the LMCS (Shaper) data field. The LMCS data field may include the LMCS parameters described above in this document. The LMCS data field may include, for example, information for deriving the LMCS (Shaper) model / bin / codeword / mapping index used for LMCS (Shaper).
[0411] The encoding device encodes the image / video information (S1940). The image / video information may include prediction-related information (prediction mode information) and / or ALF data fields. Furthermore, the image / video information may include LMCS data fields. Additionally, the image / video information may include residual information. The prediction-related information may include information about various prediction modes (e.g., merge mode, MVP mode, etc.), MVD information, etc.
[0412] Encoded image / video information can be output as a bitstream. The bitstream can be sent to a decoding device via a network or (digital) storage medium.
[0413] According to embodiments of this document, image / video information may include various types of information. For example, image / video information may include information disclosed in at least one of Tables 1, 3, 4, 7, 9, 13, 15, 16, 19, 23, 24, 25, 26, 27 and / or 28 above.
[0414] For example, image / video information may include at least one of the adaptation parameter set (APS).
[0415] For example, image / video information may include a first adaptation parameter set (APS), which includes an ALF data field, and the ALF data field may include ALF parameters for deriving filter coefficients for the ALF process.
[0416] For example, image / video information includes header information, which includes image headers or slice headers, and the header information includes ALF-related APS ID information, and the first APS including ALF data fields can be identified based on the ALF-related APS ID information.
[0417] For example, the header information includes the number of ALF-related APS IDs, and the number of ALF-related APS IDs is specified based on the value of the number of ALF-related APS IDs. ALF-related APS ID syntax elements with the same number of ALF-related APS IDs can be included in the header information.
[0418] For example, the header information includes an ALF availability flag indicating whether the ALF is available in the image or slice, and when the value of the ALF availability flag is 1, the header information includes the number of ALF-related APS IDs.
[0419] For example, image / video information may include an SPS, and the SPS may include a first ALF availability flag indicating whether the ALF is available.
[0420] For example, when the value of the first ALF availability flag is 1, the header information may include a second ALF availability flag indicating whether the ALF is available in the image or slice.
[0421] For example, the prediction mode could be an inter-frame prediction mode. In this case, generating reconstructed samples could include deriving mapped prediction samples based on prediction samples derived from the inter-frame prediction mode, and generating reconstructed samples based on the mapped prediction samples. In this case, the image / video information includes a second APS, which includes a Luminance Mapping and Chroma Scaling (LMCS) data field. The LMCS data field includes LMCS parameters indicating integer codewords, and the mapped prediction samples used for the prediction samples can be derived based on the integer codewords.
[0422] For example, the first APS includes first type information indicating that the first APS is an APS that includes an ALF data field, and the second APS includes second type information indicating that the second APS is an APS that includes an LMCS data field.
[0423] Additionally, the header information may include LMCS-related information. LMCS-related information may include at least one of the following: LMCS availability flag information (e.g., the slice_lmcs_enabled_flag syntax element or the ph_lmcs_enabled_flag syntax element), LMCS-related APS ID information (the slice_lmcs_aps_id syntax element or the ph_lmcs_aps_id syntax element), and chroma residual scaling flag information (the slice_chroma_residual_scale_flag syntax element or the ph_chroma_residual_scale_flag syntax element).
[0424] For example, the header information may include LMCS-related APS ID information, and a second APS including LMCS data fields may be identified based on the LMCS-related APS ID information.
[0425] For example, the header information may include an LMCS enable flag, and when the value of the LMCS enable flag is 1, the header information may include LMCS-related APS ID information.
[0426] Figure 21 and Figure 22 Examples of image / video decoding methods and related components according to embodiments of this document are illustrated schematically. Figure 21 The method disclosed in the article can be derived from Figure 3 The decoding device disclosed in the document performs the operation. Specifically, for example, Figure 21S2100 can be executed by the entropy decoder 310 of the decoding device, S2110 and S2120 can be executed by the predictor 330 of the decoding device, S2130 can be executed by the residual processor 320 of the decoding device, S2140 can be executed by the adder 340 of the decoding device, and S2150 and S2160 can be executed by the filter 350 of the decoding device. Figure 21 The methods disclosed herein may include the embodiments described above.
[0427] Reference Figure 21 The decoding device receives / acquires image / video information (S2100). The decoding device can receive / acquire image / video information via bitstream.
[0428] Image / video information may include prediction-related information (including prediction mode information) and residual information. Additionally, according to embodiments of this document, image / video information may include various types of information. For example, image / video information may include information disclosed in at least one of Tables 1, 3, 4, 7, 9, 13, 15, 16, 19, 23, 24, 25, 26, 27 and / or 28 above.
[0429] The decoding device can deduce the prediction mode of the current block based on image / video information (S2110). The decoding device can determine the prediction mode of the current block based on the prediction mode information.
[0430] The decoding device derives the prediction sample for the current block (S2120). The decoding device can perform prediction and derive the prediction sample based on the prediction mode.
[0431] The decoding device derives the residual samples (S2130). The decoding device can derive the residual samples based on the residual information. For example, the residual information may include information about the (quantized) transform coefficients. The quantized transform coefficients can be derived based on the information about the (quantized) transform coefficients, and the transform coefficients can be derived through a dequantization process. Subsequently, the residual samples can be derived through an inverse transform process of the transform coefficients.
[0432] The decoding device generates reconstructed samples (S2140). The decoding device can generate reconstructed samples based on predicted samples and residual samples. Reconstructed blocks and reconstructed images can be derived based on the reconstructed samples.
[0433] The decoding device derives the filter coefficients (S2150) for the ALF process used to reconstruct the samples. The decoding device can derive the filter or filter coefficients used for ALF based on the ALF parameters.
[0434] The decoding device can perform the ALF process based on filters or filter coefficients (S2160). The decoding device can generate modified reconstructed samples based on reconstructed samples and filters or filter coefficients. A filter in the ALF can include a set of filter coefficients. Filters or filter coefficients can be derived based on ALF parameters.
[0435] For example, image / video information may include at least one of the adaptation parameter sets (APS).
[0436] For example, image / video information may include a first adaptation parameter set (APS), which includes an ALF data field, and the ALF data field may include ALF parameters for deriving filter coefficients for the ALF process.
[0437] For example, image / video information includes header information, which includes image headers or slice headers, and the header information includes ALF-related APS ID information, and the first APS including ALF data fields can be identified based on the ALF-related APS ID information.
[0438] For example, the header information includes the number of ALF-related APS IDs, and the number of ALF-related APS IDs is specified based on the value of the number of ALF-related APS IDs. The header information may include as many ALF-related APS ID syntax elements as there are ALF-related APS IDs.
[0439] For example, the header information includes an ALF availability flag indicating whether the ALF is available in the image or slice, and when the ALF availability flag is 1, the header information includes the number of ALF-related APS IDs.
[0440] For example, image / video information may include an SPS, and the SPS may include a first ALF availability flag indicating whether the ALF is available.
[0441] For example, when the value of the first ALF availability flag is 1, the header information may include a second ALF availability flag indicating whether the ALF is available in the image or slice.
[0442] For example, the prediction mode could be an inter-frame prediction mode. In this case, generating reconstructed samples could include deriving mapped prediction samples based on prediction samples derived from the inter-frame prediction mode, and generating reconstructed samples based on the mapped prediction samples. In this case, the image / video information includes a second APS, which includes a Luminance Mapping and Chroma Scaling (LMCS) data field. The LMCS data field includes LMCS parameters indicating integer codewords, and the mapped prediction samples used for the prediction samples can be derived based on the integer codewords.
[0443] For example, the first APS includes first type information indicating that the first APS is an APS that includes an ALF data field, and the second APS includes second type information indicating that the second APS is an APS that includes an LMCS data field.
[0444] Additionally, the header information may include LMCS-related information. LMCS-related information may include at least one of the following: LMCS availability flag information (e.g., the slice_lmcs_enabled_flag syntax element or the ph_lmcs_enabled_flag syntax element), LMCS-related APS ID information (the slice_lmcs_aps_id syntax element or the ph_lmcs_aps_id syntax element), and chroma residual scaling flag information (the slice_chroma_residual_scale_flag syntax element or the ph_chroma_residual_scale_flag syntax element).
[0445] For example, the header information may include LMCS-related APS ID information, and a second APS including LMCS data fields may be identified based on the LMCS-related APS ID information.
[0446] For example, the header information may include an LMCS enable flag, and when the value of the LMCS enable flag is 1, the header information may include LMCS-related APS ID information.
[0447] Although the method has been described based on a flowchart listing the steps or blocks in the above embodiments, the steps of the embodiments are not limited to a specific order, and may be performed with or without the steps described above, or in a different order or simultaneously. Furthermore, those skilled in the art will understand that the steps in the flowchart are not exclusive, and that one or more steps in the flowchart may be included or may be removed without affecting the scope of the embodiments described herein.
[0448] The methods described above according to the embodiments of this document can be implemented in software, and the encoding and / or decoding devices according to this document can be included in devices performing image processing such as TVs, computers, smartphones, set-top boxes, and display devices.
[0449] When the embodiments are implemented using the software described in this document, the aforementioned methods may be implemented by modules (processes or functions) that perform the aforementioned functions. Modules may be stored in memory and executed by a processor. Memory may be installed inside or outside the processor and may be connected to the processor via various known means. The processor may include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. Memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. In other words, the embodiments disclosed herein may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in the figures may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information about the implementation (e.g., information about instructions) or algorithms may be stored in a digital storage medium.
[0450] Additionally, the decoding and encoding devices to which the embodiments of this document apply may include: multimedia broadcast transceivers, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, and real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, video-on-demand (VoD) service providers, over-the-top (OTT) video devices, internet streaming service providers, 3D video devices, virtual reality (VR) devices, augmented reality (AR) devices, image telephony video devices, vehicle terminals (e.g., vehicle (including autonomous vehicle) terminals, aircraft terminals, or ship terminals), and medical video devices; and may be used to process image signals or data. For example, OTT video devices may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).
[0451] Furthermore, the processing methods to which the embodiments of this document are applied can be generated in the form of a computer-executable program and stored in a computer-readable recording medium. Multimedia data having data structures according to the embodiments of this document can also be stored in a computer-readable recording medium. Computer-readable recording media include all kinds of storage devices and distributed storage devices in which computer-readable data is stored. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. Computer-readable recording media also include media embodied in the form of a carrier wave (e.g., transmission over the Internet). Additionally, bitstreams generated by encoding methods can be stored in a computer-readable recording medium or transmitted via wired or wireless communication networks.
[0452] Additionally, the embodiments of this document may be embodied as a computer program product based on program code, and the program code may be executed on a computer according to the embodiments of this document. The program code may be stored on a computer-readable medium.
[0453] Figure 23 Examples of content streaming systems to which the embodiments described in this document can be applied.
[0454] refer to Figure 23 Content streaming systems using embodiments of this document typically include encoding servers, streaming servers, web servers, media storage, user devices, and multimedia input devices.
[0455] An encoding server is used to compress content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generate a bitstream, and transmit it to a streaming server. As another example, if the multimedia input device, such as a smartphone, camera, or camcorder, directly generates the bitstream, the encoding server can be omitted.
[0456] Bitstreams can be generated using the encoding methods or bitstream generation methods applied in the embodiments of this document. Furthermore, the streaming server can temporarily store the bitstream during transmission or reception.
[0457] A streaming server sends multimedia data to a user's device via a web server based on the user's request. The web server acts as a tool to notify the user of available services. When a user requests a desired service, the web server forwards the request to the streaming server, which then sends the multimedia data to the user. In this respect, the content streaming system may include a separate control server, which in this case controls the commands / responses between the various devices within the content streaming system.
[0458] A streaming server can receive content from media storage devices and / or encoding servers. For example, if content is received from an encoding server, it can be received in real time. In this case, the streaming server can store the bitstream for a predetermined period of time to provide a smooth streaming service.
[0459] For example, user equipment may include mobile phones, smartphones, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, board PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc.
[0460] Each server in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be processed in a distributed manner.
[0461] The claims described herein can be combined in various ways. For example, the technical features of the method claims of this specification can be combined and implemented as an apparatus, and the technical features of the apparatus claims of this specification can be combined and implemented as a method. Furthermore, the technical features of the method claims and the apparatus claims of this specification can be combined to implement an apparatus, and the technical features of the method claims and the apparatus claims of this specification can be combined and implemented as a method.
Claims
1. A decoding device for image decoding, the decoding device comprising: Memory; as well as At least one processor, connected to the memory, is configured to: Image information, including prediction mode information and residual information, is obtained through bitstream; Based on the prediction pattern information, deduce the prediction pattern of the current block; Based on the prediction model, derive the prediction sample; Based on the residual information, the residual samples are derived; Based on the predicted samples and the residual samples, reconstructed samples are generated; Derive the filter coefficients for the adaptive loop filter (ALF) used for the reconstructed samples; as well as Based on the reconstructed samples and the filter coefficients, modified reconstructed samples are generated. The image information includes a sequence parameter set (SPS), slice header information, and a first adaptation parameter set (APS). The first APS includes a first type of information related to whether the first APS is an APS that includes ALF data fields. Wherein, based on the first type information indicating that the first APS is an APS including the ALF data field, the first APS includes the ALF data field. The ALF data field includes ALF parameters used to derive the filter coefficients. The SPS includes ALF enable flag information related to whether the ALF is enabled. Specifically, based on the ALF enable flag information in the SPS, the slice header information includes ALF enable flag information related to whether the ALF is enabled in the slice. The slice header information includes ALF-related APSID quantity information based on the ALF enable flag information in the slice header. Specifically, based on the value of the ALF-related APS ID quantity information, the quantity of ALF-related APS IDs is derived, and Among them, the number of ALF-related APS ID syntax elements, which is equal to the number of ALF-related APSIDs, is included in the slice header information.
2. An encoding device for image encoding, the encoding device comprising: Memory; as well as At least one processor, connected to the memory, is configured to: The prediction sample for the current block is derived based on inter-frame prediction or intra-frame prediction. Generate prediction mode information representing the prediction mode of the current block; Derive the residual samples based on the predicted samples; Residual information is generated based on the residual samples; Reconstructed samples are generated based on the predicted samples; Derive the ALF parameters for the adaptive loop filter (ALF) used for the reconstructed samples; and The image information, including the prediction mode information, the residual information, and the ALF parameters, is encoded. The image information includes a sequence parameter set (SPS), slice header information, and a first adaptation parameter set (APS). The first APS includes a first type of information related to whether the first APS is an APS that includes ALF data fields. Wherein, based on the first type information indicating that the first APS is an APS including the ALF data field, the first APS includes the ALF data field. The ALF data field includes ALF parameters used to derive filter coefficients. The SPS includes ALF enable flag information related to whether the ALF is enabled. Specifically, based on the ALF enable flag information in the SPS, the slice header information includes ALF enable flag information related to whether the ALF is enabled in the slice. The slice header information includes ALF-related APSID quantity information based on the ALF enable flag information in the slice header. Specifically, based on the value of the ALF-related APS ID quantity information, the quantity of ALF-related APS IDs is derived, and Among them, the number of ALF-related APS ID syntax elements, which is equal to the number of ALF-related APSIDs, is included in the slice header information.
3. An apparatus for transmitting image data, the apparatus comprising: At least one processor is configured to acquire a bitstream for the image, wherein the bitstream is generated based on: deriving prediction samples for the current block based on inter-frame prediction or intra-frame prediction; generating prediction mode information representing the prediction mode of the current block; deriving residual samples based on the prediction samples; generating residual information based on the residual samples; generating reconstructed samples based on the prediction samples; deriving ALF parameters for an adaptive loop filter (ALF) for the reconstructed samples; and encoding the image information to generate the bitstream, wherein the image information includes the prediction mode information, the residual information, and the ALF parameters. A transmitter configured to transmit the data comprising the bit stream. The image information includes a sequence parameter set (SPS), slice header information, and a first adaptation parameter set (APS). The first APS includes a first type of information related to whether the first APS is an APS that includes ALF data fields. Wherein, based on the first type information indicating that the first APS is an APS including the ALF data field, the first APS includes the ALF data field. The ALF data field includes ALF parameters used to derive filter coefficients. The SPS includes ALF enable flag information related to whether the ALF is enabled. Specifically, based on the ALF enable flag information in the SPS, the slice header information includes ALF enable flag information related to whether the ALF is enabled in the slice. The slice header information includes ALF-related APSID quantity information based on the ALF enable flag information in the slice header. Specifically, based on the value of the ALF-related APS ID quantity information, the quantity of ALF-related APS IDs is derived, and Among them, the number of ALF-related APS ID syntax elements, which is equal to the number of ALF-related APSIDs, is included in the slice header information.