Image encoding / decoding methods and apparatus using lossless color transformation, and bitstream transmission methods

By employing lossless color transformation technology during image encoding/decoding and resetting the values ​​of residual samples, the problem of low encoding/decoding efficiency in high-resolution and high-quality image transmission is solved, achieving more efficient image information transmission and storage.

CN114930827BActive Publication Date: 2025-11-14NOKIA TECHNOLOGIES OY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080092618.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-22
Filing Date
2020-11-20
Publication Date
2025-11-14
Estimated Expiration
2040-11-20

AI Technical Summary

Technical Problem

Existing technologies suffer from low encoding/decoding efficiency in the transmission of high-resolution and high-quality images, leading to increased transmission and storage costs.

Method used

Lossless color transformation technology is adopted to improve encoding/decoding efficiency by determining the residual sample of the current block and resetting the value of the residual sample based on half of the chroma residual sample value.

Benefits of technology

It improves the efficiency of image encoding/decoding, reduces transmission and storage costs, and achieves efficient image information transmission and storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114930827B_ABST
    Figure CN114930827B_ABST
Patent Text Reader

Abstract

Image encoding / decoding methods and apparatus are provided. The image decoding method performed by the image decoding apparatus according to this disclosure may include the steps of: determining a residual sample of a current block; and resetting the value of the residual sample based on whether a color space transformation has been applied. The step of resetting the value of the residual block may be performed based on half the value of the chromaticity residual sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to image encoding / decoding methods and apparatus. More specifically, this disclosure relates to image encoding / decoding methods and apparatus using color transformation, and a method for transmitting a bitstream generated by the image encoding method / apparatus of this disclosure. Background Technology

[0002] Recently, there has been an increasing demand for high-resolution and high-quality images, such as high-definition (HD) and ultra-high-definition (UHD) images, across various fields. With the improvement in image data resolution and quality, the amount of information or bits transmitted increases relative to existing image data. This increase in the amount of information or bits transmitted leads to increased transmission and storage costs.

[0003] Therefore, efficient image compression techniques are needed to effectively transmit, store, and reproduce information about high-resolution and high-quality images. Summary of the Invention

[0004] Technical issues

[0005] The purpose of this disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0006] In addition, this disclosure aims to provide an image encoding / decoding method and apparatus for improving encoding / decoding efficiency by performing lossless color transformation.

[0007] In addition, this disclosure aims to provide a method for transmitting a bit stream generated by an image encoding method or apparatus according to this disclosure.

[0008] In addition, this disclosure aims to provide a recording medium in which a bitstream generated by an image encoding method or apparatus according to this disclosure is stored.

[0009] Furthermore, this disclosure aims to provide a recording medium in which a bitstream is stored, which is received, decoded, and used to reconstruct an image by an image decoding apparatus according to this disclosure.

[0010] Those skilled in the art will understand that the technical objectives to be achieved in this disclosure are not limited to those mentioned above, and that other technical objectives not described herein will be clearly understood from the following description.

[0011] Technical solution

[0012] According to one aspect of this disclosure, an image decoding method performed by an image decoding apparatus is provided, the method comprising the steps of: determining a residual sample of a current block; and resetting the value of the residual sample based on whether a color space transformation has been applied. Here, the resetting of the residual sample value may be performed based on half the value of the chroma residual sample.

[0013] Furthermore, according to another aspect of this disclosure, an image decoding apparatus including a memory and at least one processor is provided, wherein the at least one processor is configured to: determine a residual sample of a current block, and reset the value of the residual sample based on whether a color space transformation has been applied. Here, the processor may be configured to reset the value of the residual sample based on a half-value of the chroma residual sample value.

[0014] Furthermore, according to another aspect of this disclosure, an image encoding method performed by an image encoding apparatus is provided, the method comprising the steps of: determining a residual sample of a current block; and resetting the value of the residual sample based on whether a color space transformation has been applied. Here, the resetting of the residual sample value can be performed based on half the value of the chroma residual sample.

[0015] In addition, according to another aspect of this disclosure, a method for transmitting a bit stream generated by the image encoding apparatus or image encoding method of this disclosure is provided.

[0016] In addition, according to another aspect of this disclosure, a computer-readable recording medium is provided in which a bitstream generated by the image encoding method or image encoding apparatus of this disclosure is stored.

[0017] The features described above in this brief overview are merely exemplary aspects of the following detailed description of this disclosure and do not limit the scope of this disclosure.

[0018] Beneficial effects

[0019] According to this disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.

[0020] In addition, according to this disclosure, an image encoding / decoding method and apparatus capable of improving encoding / decoding efficiency by performing lossless color transformation can be provided.

[0021] Additionally, according to this disclosure, a method for transmitting a bitstream generated by an image encoding method or apparatus according to this disclosure can be provided.

[0022] Additionally, according to this disclosure, a recording medium in which a bitstream generated by an image encoding method or apparatus according to this disclosure is stored can be provided.

[0023] Additionally, according to this disclosure, a recording medium may be provided in which a bitstream of an image is received, decoded, and used to reconstruct an image by an image decoding apparatus according to this disclosure is stored.

[0024] Those skilled in the art will understand that the effects achievable through this disclosure are not limited to those specifically described above, and that other advantages of this disclosure will become clearer from the following description. Attached Figure Description

[0025] Figure 1 This is a view schematically illustrating a video encoding system to which embodiments of this disclosure are applicable.

[0026] Figure 2 This is a schematic view illustrating an image encoding apparatus to which embodiments of the present disclosure are applicable.

[0027] Figure 3 This is a schematic view illustrating an image decoding apparatus to which embodiments of the present disclosure are applicable.

[0028] Figure 4 This is a view showing the segmentation structure of an image according to an embodiment.

[0029] Figure 5 This is a view illustrating an implementation of the block segmentation type based on a multi-type tree structure.

[0030] Figure 6 This is a view illustrating the signaling mechanism for block segmentation information of a quadtree structure with nested multi-type trees according to this disclosure.

[0031] Figure 7 This is a view showing an implementation where the CTU is divided into multiple CUs.

[0032] Figure 8 This is a view showing a neighboring reference sample according to an embodiment.

[0033] Figure 9 and Figure 10 This is a view illustrating intra-frame prediction according to an implementation method.

[0034] Figure 11 This is a view illustrating a coding method using inter-frame prediction according to an embodiment.

[0035] Figure 12 This is a view illustrating a decoding method using inter-frame prediction according to an embodiment.

[0036] Figure 13 This is a view showing a block diagram of CABAC according to an implementation method for encoding a syntax element.

[0037] Figures 14 to 17 This is a view illustrating entropy encoding and decoding according to an implementation method.

[0038] Figure 18 and Figure 19 This is a view illustrating an example of an image decoding and encoding process according to an implementation method.

[0039] Figure 20 This is a view showing the hierarchical structure of the encoded image according to an embodiment.

[0040] Figure 21 This is a view illustrating an implementation of the decoding process using ACT.

[0041] Figure 22 This is a view illustrating an implementation of a sequence parameter set syntax table in which the grammar elements associated with the ACT are signaled.

[0042] Figures 23 to 29 It is a view that continuously illustrates an implementation of a syntax table in which the encoding basis for signaling grammatical elements related to ACT is shown.

[0043] Figure 30 This is a view showing the encoding tree syntax according to the implementation method.

[0044] Figure 31 This is a view illustrating the encoding method of residual samples of BDPCM according to an embodiment.

[0045] Figure 32 This is a view showing a modified quantization residual block generated by performing BDPCM according to an embodiment.

[0046] Figure 33 This is a flowchart illustrating the process of encoding the current block by applying BDPCM in an image encoding apparatus according to an embodiment.

[0047] Figure 34 This is a flowchart illustrating the process of reconstructing the current block by applying BDPCM in an image decoding device according to an embodiment.

[0048] Figures 35 to 37 This is a schematic view illustrating the syntax used to signal information about BDPCM.

[0049] Figures 38 to 53 This is a view showing a syntax table for signaling ACT syntax elements according to various embodiments of the present disclosure.

[0050] Figure 54 and Figure 55 This is a view illustrating an image decoding method according to an embodiment.

[0051] Figure 56 and Figure 57 This is a view illustrating an image encoding method according to an embodiment.

[0052] Figure 58 This is a view illustrating a content streaming system to which embodiments of the present disclosure apply. Detailed Implementation

[0053] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings to facilitate implementation by those skilled in the art. However, this disclosure can be implemented in various different forms and is not limited to the embodiments described herein.

[0054] In describing this disclosure, detailed descriptions of relevant known functions or constructions will be omitted if they unnecessarily obscure the scope of this disclosure. In the accompanying drawings, portions irrelevant to the description of this disclosure are omitted, and similar reference numerals are assigned to similar portions.

[0055] In this disclosure, when a component is "connected," "coupled," or "linked" to another component, it may include not only direct connections but also indirect connections where intermediate components exist. Furthermore, when a component "comprises" or "has" other components, unless otherwise stated, it means that other components may be included, not excluded.

[0056] In this disclosure, the terms first, second, etc., are used only for the purpose of distinguishing one component from other components and do not limit the order or importance of the components, unless otherwise stated. Accordingly, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.

[0057] In this disclosure, the components are distinguished from each other to clearly describe each feature, but this does not mean that the components must be separate. That is, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed and implemented across multiple hardware or software units. Therefore, unless otherwise specified, implementations of these integrated or distributed components are included within the scope of this disclosure.

[0058] In this disclosure, the components described in the various embodiments are not necessarily essential components, and some components may be optional. Therefore, embodiments consisting of a subset of the components described in the embodiments are also included within the scope of this disclosure. Furthermore, embodiments that include other components besides those described in the various embodiments are also included within the scope of this disclosure.

[0059] This disclosure relates to the encoding and decoding of images. Unless redefined in this disclosure, the terms used herein may have the general meaning commonly used in the art to which this disclosure pertains.

[0060] In this disclosure, "video" can mean a collection of images over time. "Picture" generally means a basis representing an image within a specific time period, and slices / tiles are the coding basis that constitutes part of a picture during encoding. A picture can consist of one or more slices / tiles. Additionally, slices / tiles can include one or more coding tree units (CTUs). A picture can consist of one or more slices / tiles. A picture can consist of one or more groups of tiles. A group of tiles can include one or more tiles. A brick can refer to a quadrilateral region within a CTU row of a tile in a picture. A tile can include one or more bricks. A brick can refer to a quadrilateral region within a CTU row of a tile. A tile can be divided into multiple bricks, and each brick can include one or more CTU rows belonging to the tile. Tiles that are not divided into multiple bricks can also be treated as bricks.

[0061] In this disclosure, "pixel" or "pellet" can refer to the smallest unit that constitutes a picture (or image). Furthermore, "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a pixel value, or it can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.

[0062] In this disclosure, "unit" can refer to a basic unit of image processing. A unit may include a specific region of an image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., Cb, Cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "sample array," "block," or "region." Typically, an M×N block may include a set (or array) of samples (sample array) or transform coefficients in M ​​columns and N rows.

[0063] In this disclosure, "current block" can mean one of "current coding block," "current coding unit," "coding target block," "decoding target block," or "processing target block." When performing prediction, "current block" can mean "current prediction block" or "prediction target block." When performing transform (inverse transform) / quantization (dequantization), "current block" can mean "current transform block" or "transform target block." When performing filtering, "current block" can mean "filter target block."

[0064] Furthermore, in this disclosure, unless explicitly stated as a chroma block, "current block" may mean "the luminance block of the current block". "The chroma block of the current block" can be expressed by including an explicit description of a chroma block such as "chroma block" or "current chroma block".

[0065] In this disclosure, the terms “ / ” and “,” should be interpreted as indicating “and / or”. For example, the expressions “A / B” and “A, B” can mean “A and / or B”. In addition, “A / B / C” and “A, B, C” can mean “at least one of A, B and / or C”.

[0066] In this disclosure, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" can include 1) "A only", 2) "B only" and / or 3) both "A and B". In other words, in this disclosure, the term "or" should be interpreted as indicating "additionally or alternatively".

[0067] Overview of video coding systems

[0068] Figure 1 This is a view showing a video encoding system according to this disclosure.

[0069] The video encoding system according to the embodiment may include a source device 10 and a receiving device 20. The source device 10 may transmit the encoded video and / or image information or data to the receiving device 20 in the form of a file or stream via a digital storage medium or network.

[0070] The source device 10 according to an embodiment may include a video source generator 11, an encoding device 12, and a transmitter 13. The receiving device 20 according to an embodiment may include a receiver 21, a decoding device 22, and a renderer 23. The encoding device 12 may be referred to as a video / image encoding device, and the decoding device 22 may be referred to as a video / image decoding device. The transmitter 13 may be included in the encoding device 12. The receiver 21 may be included in the decoding device 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.

[0071] The video source generator 11 can acquire video / images through processes that capture, synthesize, or generate video / images. The video source generator 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device may include, for example, a computer, tablet, and smartphone, and can generate video / images (electronically). For example, virtual video / images can be generated by a computer, etc. In this case, the video / image capture process can be replaced by a process that generates related data.

[0072] Encoding device 12 can encode input video / images. Encoding device 12 can perform a series of processes such as prediction, transformation, and quantization for compression and encoding efficiency. Encoding device 12 can output encoded data (encoded video / image information) in the form of a bitstream.

[0073] Transmitter 13 can transmit encoded video / image information or data, output in bitstream form, to receiver 21 of receiving device 20 in the form of a file or stream via digital storage medium or network. Digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. Transmitter 13 may include elements for generating media files according to a predetermined file format and may include elements for transmission via broadcast / communication networks. Receiver 21 can extract / receive bitstreams from storage medium or network and transmit the bitstreams to decoding device 22.

[0074] Decoding device 22 can decode video / images by performing a series of processes such as inverse quantization, inverse transform, and prediction, which correspond to the operations of encoding device 12.

[0075] Renderer 23 can render decoded video / images. The rendered video / images can be displayed on a monitor.

[0076] Overview of Image Encoding Devices

[0077] Figure 2 This is a schematic view illustrating an image encoding device to which embodiments of this disclosure may be applied.

[0078] like Figure 2 As shown, the image encoding device 100 may include an image segmenter 110, a subtractor 115, a transformer 120, a quantizer 130, an inverse quantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-frame prediction unit 180, an intra-frame prediction unit 185, and an entropy encoder 190. The inter-frame prediction unit 180 and the intra-frame prediction unit 185 may be collectively referred to as "prediction units". The transformer 120, quantizer 130, inverse quantizer 140, and inverse transformer 150 may be included in a residual processor. The residual processor may also include a subtractor 115.

[0079] In some implementations, all or at least some of the components configuring the image encoding device 100 may be configured by a single hardware component (e.g., an encoder or a processor). Furthermore, the memory 170 may include a decoded image buffer (DPB) and may be configured by a digital storage medium.

[0080] Image segmenter 110 can segment an input image (or picture or frame) input to image encoding device 100 into one or more processing units. For example, a processing unit may be called an encoding unit (CU). Encoding units can be obtained by recursively segmenting encoding tree units (CTUs) or maximum encoding units (LCUs) according to a quadtree / binary tree / tritree (QT / BT / TT) structure. For example, an encoding unit can be segmented into multiple encoding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the segmentation of encoding units, a quadtree structure can be applied first, followed by a binary tree structure and / or a ternary tree structure. The encoding process according to this disclosure can be performed based on the final encoding unit that is no longer segmented. The maximum encoding unit can be used as the final encoding unit, or a deeper encoding unit obtained by segmenting the maximum encoding unit can be used as the final encoding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes, which will be described later. As another example, the processing unit of the encoding process may be a prediction unit (PU) or a transformation unit (TU). Prediction units and transform units can be partitioned or segmented from the final coding unit. Prediction units can be sample prediction units, and transform units can be units used to derive transform coefficients and / or units used to derive residual signals from transform coefficients.

[0081] The prediction unit (inter-frame prediction unit 180 or intra-frame prediction unit 185) can perform prediction on the block to be processed (the current block) and generate a prediction block that includes prediction samples of the current block. The prediction unit can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. The prediction unit can generate various information related to the prediction of the current block and send the generated information to the entropy encoder 190. The information about the prediction can be encoded in the entropy encoder 190 and output as a bitstream.

[0082] Intra-prediction unit 185 can predict the current block by referencing samples in the current image. Depending on the intra-prediction mode and / or intra-prediction technique, the reference samples may be located among the neighbors of the current block or may be placed separately. Intra-prediction modes may include multiple non-directional modes and multiple directional modes. Non-directional modes may include, for example, DC modes and planar modes. Depending on the level of detail in the prediction direction, directional modes may include, for example, 33 or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. Intra-prediction unit 185 can determine the prediction mode to be applied to the current block by using prediction modes applied to neighboring blocks.

[0083] The inter-frame prediction unit 180 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a juxtaposed reference block, a juxtaposed CU (colCU), etc. The reference image including the temporally neighboring block may be referred to as a juxtaposed image (colPic). For example, the inter-frame prediction unit 180 can configure a motion information candidate list based on neighboring blocks and generate information specifying which candidate to use to deduce the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame prediction unit 180 can use the motion information of neighboring blocks as the motion information of the current block. In skip mode, unlike merge mode, residual signals may not be transmitted. In motion vector prediction (MVP) mode, the motion vectors of neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be signaled by encoding motion vector differences and indicators of the motion vector predictors. The motion vector difference can refer to the difference between the motion vector of the current block and the motion vector predictor.

[0084] The prediction unit can generate a prediction signal based on various prediction methods and techniques described below. For example, the prediction unit can apply not only intra-frame prediction or inter-frame prediction, but also both intra-frame prediction and inter-frame prediction simultaneously to predict the current block. A prediction method that simultaneously applies both intra-frame prediction and inter-frame prediction to predict the current block can be called Combined Inter-Frame and Intra-Frame Prediction (CIIP). Furthermore, the prediction unit can perform Intra-Frame Block Copy (IBC) to predict the current block. Intra-Frame Block Copy can be used for content image / video coding in games, for example, Screen Content Coding (SCC). IBC is a method of predicting the current image using a previously reconstructed reference block in the current image at a predetermined distance from the current block. When IBC is applied, the position of the reference block in the current image can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC essentially performs prediction in the current image, but can be performed similarly to inter-frame prediction because the reference block is derived within the current image. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure.

[0085] The prediction signal generated by the prediction unit can be used to generate a reconstructed signal or a residual signal. Subtractor 115 can generate a residual signal (residual block or residual sample array) by subtracting the prediction signal (prediction block or prediction sample array) output from the prediction unit from the input image signal (original block or original sample array). The generated residual signal can be sent to converter 120.

[0086] Transformer 120 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform techniques may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented graphically. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. Furthermore, the transform processing can be applied to square pixel blocks of the same size or to blocks of variable size instead of square.

[0087] Quantizer 130 can quantize the transform coefficients and send them to entropy encoder 190. Entropy encoder 190 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 130 can rearrange the block-form quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and generate information about the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients.

[0088] The entropy encoder 190 can perform various encoding methods, such as exponential Columbus coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 190 can encode, either together or separately, the information required for video / image reconstruction other than the quantization transform coefficients (e.g., values ​​of syntax elements). The encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form at the Network Abstraction Layer (NAL) level. The video / image information may also include information about various parameter sets, such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). Furthermore, the video / image information may also include general constraint information. The signaled information, transmitted information, and / or syntax elements described in this disclosure can be encoded and included in the bitstream through the above encoding process.

[0089] The bitstream can be transmitted over a network or stored in a digital storage medium. The network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) for transmitting the signal output from the entropy encoder 190 and / or a storage unit (not shown) for storing the signal may be included as internal / external components of the image encoding device 100. Alternatively, a transmitter may be provided as a component of the entropy encoder 190.

[0090] The quantization transform coefficients output from quantizer 130 can be used to generate residual signals. For example, the residual signals (residual blocks or residual samples) can be reconstructed by applying dequantization and inverse transform to the quantization transform coefficients using dequantizer 140 and inverse transformer 150.

[0091] Adder 155 adds the reconstructed residual signal to the prediction signal output from inter-frame prediction unit 180 or intra-frame prediction unit 185 to generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). If the block to be processed has no residual, such as in the case of applying skip mode, the predicted block can be used as a reconstructed block. Adder 155 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, and can be used for inter-frame prediction of the next image by filtering as described below.

[0092] Filter 160 can improve the quality of subjective / objective images by applying filtering to the reconstructed signal. For example, filter 160 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 170, specifically in the DPB of memory 170. Various filtering methods can include, for example, deblocking filtering, sample adaptive shifting, adaptive loop filtering, bilateral filtering, etc. Filter 160 can generate various filtering-related information and transmit the generated information to entropy encoder 190, as described later in the description of each filtering method. The filtering-related information can be encoded by entropy encoder 190 and output as a bitstream.

[0093] The modified reconstructed image sent to memory 170 can be used as a reference image in inter-frame prediction unit 180. When inter-frame prediction is applied by image coding device 100, prediction mismatch between image coding device 100 and image decoding device can be avoided and coding efficiency can be improved.

[0094] The DPB of memory 170 can store modified reconstructed images for use as reference images in inter-frame prediction unit 180. Memory 170 can store motion information of blocks from which motion information in the current image is derived (or encoded) and / or motion information of already reconstructed blocks in the image. The stored motion information can be transmitted to inter-frame prediction unit 180 and used as motion information for spatially or temporally neighboring blocks. Memory 170 can store reconstructed samples of reconstructed blocks in the current image and can transmit the reconstructed samples to intra-frame prediction unit 185.

[0095] Overview of image decoding devices

[0096] Figure 3 This is a schematic view illustrating an image decoding device to which embodiments of the present disclosure may be applied.

[0097] like Figure 3 As shown, the image decoding device 200 may include an entropy decoder 210, an inverse quantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame prediction unit 260, and an intra-frame prediction unit 265. The inter-frame prediction unit 260 and the intra-frame prediction unit 265 may be collectively referred to as "prediction units". The inverse quantizer 220 and the inverse transformer 230 may be included in a residual processor.

[0098] According to an implementation, all or at least some of the components of the image decoding device 200 can be configured by hardware components (e.g., a decoder or a processor). Furthermore, the memory 250 may include a decoded image buffer (DPB) or may be configured by a digital storage medium.

[0099] The image decoding device 200, having received a bitstream including video / image information, can perform operations related to... Figure 2 The image is reconstructed by processing corresponding to the processing performed by the image encoding device 100. For example, the image decoding device 200 can perform decoding using a processing unit applied in the image encoding device. Therefore, the decoding processing unit can be, for example, an encoding unit. The encoding unit can be obtained by segmenting a coding tree unit or a maximum coding unit. The reconstructed image signal decoded and output by the image decoding device 200 can be reproduced by a reproduction device (not shown).

[0100] Image decoding device 200 can receive data in bitstream form from... Figure 2The received signal is output by the image encoding device. The entropy decoder 210 can decode the bitstream. For example, the entropy decoder 210 can parse the bitstream to derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets, such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). Furthermore, the video / image information may also include general constraint information. The image decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The information and / or syntax elements notified / received by signals described in this disclosure can be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 210 decodes the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, and outputs the values ​​of the syntax elements required for image reconstruction and the quantized values ​​of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine the context model using information about the target syntax element, decoding information of neighboring blocks and the target block, or information about symbols / bins decoded in the previous stage, perform arithmetic decoding on the bins based on the determined context model by predicting the occurrence probability of the bins, and generate symbols corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method can update the context model after determining the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin. The prediction-related information in the information decoded by the entropy decoder 210 can be provided to the prediction units (inter-frame prediction unit 260 and intra-frame prediction unit 265), and the residual value of entropy decoding performed in the entropy decoder 210, i.e., the quantization transform coefficients and related parameter information, can be input to the dequantizer 220. In addition, the filtering information in the information decoded by the entropy decoder 210 can be provided to the filter 240. Furthermore, the receiver (not shown) for receiving signals output from the image encoding device may be further configured as an internal / external element of the image decoding device 200, or the receiver may be a component of the entropy decoder 210.

[0101] Furthermore, the image decoding device according to this disclosure can be referred to as a video / image / picture decoding device. The image decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 210. The sample decoder may include at least one of a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame prediction unit 260, or an intra-frame prediction unit 265.

[0102] The dequantizer 220 can dequantize the quantized transform coefficients and output transform coefficients. The dequantizer 220 can rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scan order performed in the image encoding device. The dequantizer 220 can obtain the transform coefficients by performing dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information).

[0103] The inverse transformer 230 can perform inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).

[0104] The prediction unit can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. The prediction unit can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on the prediction information output from the entropy decoder 210, and can determine a specific intra-frame / inter-frame prediction mode (prediction technique).

[0105] Similar to that described in the prediction unit of the image coding device 100, the prediction unit can generate a prediction signal based on various prediction methods (techniques) described later.

[0106] Intra-prediction unit 265 can predict the current block by referring to samples in the current image. The description of intra-prediction unit 185 also applies to intra-prediction unit 265.

[0107] The inter-frame prediction unit 260 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. For example, the inter-frame prediction unit 260 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference image index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information specifying the inter-frame prediction mode of the current block.

[0108] Adder 235 generates a reconstruction signal (reconstructed image, reconstruction block, reconstruction sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including inter-frame prediction unit 260 and / or intra-frame prediction unit 265). If the block to be processed has no residual, for example when a skip mode is applied, the prediction block can be used as a reconstruction block. The description of adder 155 also applies to adder 235. Adder 235 can be referred to as a reconstructor or reconstruction block generator. The generated reconstruction signal can be used for intra-frame prediction of the next block to be processed in the current image, and can be used for inter-frame prediction of the next image by filtering as described below.

[0109] Filter 240 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 240 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 250, specifically in the DPB of memory 250. Various filtering methods may include, for example, deblocking filtering, adaptive sample shifting, adaptive loop filtering, bilateral filtering, etc.

[0110] The (modified) reconstructed image stored in the DPB of memory 250 can be used as a reference image in inter-frame prediction unit 260. Memory 250 can store motion information of blocks from which motion information in the current image is derived (or decoded) and / or motion information of already reconstructed blocks in the image. The stored motion information can be transmitted to inter-frame prediction unit 260 to be used as motion information for spatially or temporally neighboring blocks. Memory 250 can store reconstructed samples of reconstructed blocks in the current image and transmit the reconstructed samples to intra-frame prediction unit 265.

[0111] In this disclosure, the embodiments described in the filter 160, inter-frame prediction unit 180 and intra-frame prediction unit 185 of the image encoding device 100 can be equally or correspondingly applied to the filter 240, inter-frame prediction unit 260 and intra-frame prediction unit 265 of the image decoding device 200.

[0112] Overview of Image Segmentation

[0113] The video / image coding method according to this disclosure can be performed based on the image segmentation structure as follows. Specifically, the processes of prediction, residual processing (inverse transform, dequantization, etc.), syntax element encoding, and filtering, which will be described later, can be performed based on the CTU, CU (and / or TU, PU) derived from the image segmentation structure. The image can be segmented into block units, and the block segmentation process can be performed in the image segmenter 110 of the encoding device. Segmentation-related information can be encoded by the entropy encoder 190 and sent to the decoding device in the form of a bitstream. The entropy decoder 210 of the decoding device can deduce the block segmentation structure of the current image based on the segmentation-related information obtained from the bitstream, and based on this, a series of processes (e.g., prediction, residual processing, block / image reconstruction, in-loop filtering, etc.) can be performed for image decoding.

[0114] An image can be segmented into a sequence of coding tree units (CTUs). Figure 4 An example of an image being segmented into CTUs is shown. A CTU may correspond to a Coding Tree Block (CTB). Alternatively, a CTU may include a coded tree block for luma samples and two coded tree blocks for corresponding chroma samples. For example, for an image containing three sample arrays, a CTU may include an N×N block for luma samples and two corresponding blocks for chroma samples.

[0115] Overview of CTU segmentation

[0116] As described above, coding units can be obtained by recursively partitioning coding tree units (CTUs) or maximum coding units (LCUs) according to a quadtree / binary tree / ternary tree (QT / BT / TT) structure. For example, a CTU can be first partitioned into a quadtree structure. Subsequently, the leaf nodes of the quadtree structure can be further partitioned using multiple tree types.

[0117] The quadtree partitioning means that the current CU (or CTU) is equally divided into four. By partitioning according to the quadtree, the current CU can be divided into four CUs with the same width and height. When the current CU is no longer partitioned into a quadtree structure, the current CU corresponds to a leaf node of the quadtree structure. The CUs corresponding to the leaf nodes of the quadtree structure may not be further partitioned and can be used as the final encoding units described above. Alternatively, the CUs corresponding to the leaf nodes of the quadtree structure can be further partitioned using multiple types of tree structures.

[0118] Figure 5 This is a view showing the segmentation types of blocks based on multiple tree structures. Segmentation based on multiple tree structures can include two types of segmentation based on binary tree structures and two types of segmentation based on ternary tree structures.

[0119] The two types of partitioning based on the binary tree structure can be vertical binary partitioning (SPLIT_BT_VER) and horizontal binary partitioning (SPLIT_BT_HOR). Vertical binary partitioning (SPLIT_BT_VER) means that the current CU is equally divided into two in the vertical direction. For example... Figure 4 As shown, a vertical binary partition can generate two CUs with the same height as the current CU and a width half the width of the current CU. A horizontal binary partition (SPLIT_BT_HOR) means that the current CU is equally divided into two in the horizontal direction. Figure 5 As shown, by using horizontal binary partitioning, two CUs can be generated with a height that is half the height of the current CU and a width that is the same as the current CU.

[0120] The two types of partitioning based on the ternary number structure can include vertical ternary partitioning (SPLIT_TT_VER) and horizontal ternary partitioning (SPLIT_TT_HOR). In vertical ternary partitioning (SPLIT_TT_VER), the current CU is partitioned vertically in a 1:2:1 ratio. For example... Figure 5 As shown, a vertical truncation can generate two CUs with the same height as the current CU and a width one-quarter of the current CU's width, and one CU with the same height as the current CU and a width half of the current CU's width. In a horizontal truncation (SPLIT_TT_HOR), the current CU is divided horizontally in a 1:2:1 ratio. Figure 5 As shown, by dividing horizontally into three branches, two CUs with a height of 1 / 4 of the current CU's height and the same width as the current CU, and one CU with a height of half the current CU's height and the same width as the current CU, can be generated.

[0121] Figure 6 This is a view illustrating the signaling mechanism for block segmentation information of a quadtree structure with nested multi-type trees according to this disclosure.

[0122] Here, the CTU is treated as the root node of a quadtree and is initially split into a quadtree structure. A signal (e.g., `qt_split_flag`) indicates whether to perform a quadtree split for the current CU (a node in the quadtree (QT_node) or CTU). For example, when `qt_split_flag` has a first value (e.g., "1"), the current CU can be split into a quadtree. Alternatively, when `qt_split_flag` has a second value (e.g., "0"), the current CU is not split into a quadtree but becomes a leaf node (QT_leaf_node). Each quadtree leaf node can then be further split into a multi-type tree structure. That is, a leaf node of the quadtree can become a node in a multi-type tree (MTT_node). In the multi-type tree structure, a first flag (e.g., `Mtt_split_cu_flag`) is signaled to indicate whether the current node is further split. If the corresponding node is further split (e.g., if the first flag is 1), a second flag (e.g., `Mtt_split_cu_vertical_flag`) can be signaled to indicate the split direction. For example, if the second flag is 1, the split direction can be vertical, and if the second flag is 0, the split direction can be horizontal. Then, a third flag (e.g., `Mtt_split_cu_binary_flag`) can be signaled to indicate whether the split type is binary or ternary. For example, when the third flag is 1, the split type can be binary, and when the third flag is 0, the split type can be ternary. Nodes of the multi-type tree obtained through binary or ternary splits can be further split into multi-type tree structures. However, nodes of the multi-type tree cannot be split into quadtree structures. If the first flag is 0, the corresponding node of the multi-type tree is no longer split but becomes a leaf node (`MTT_leaf_node`) of the multi-type tree. The CU corresponding to the leaf node of the multi-type tree can be used as the final encoding unit mentioned above.

[0123] Based on `mtt_split_cu_vertical_flag` and `mtt_split_cu_binary_flag`, the multi-type tree partitioning mode (MttSplitMode) of the CU can be derived as shown in Table 1 below. In the following description, the multi-type tree partitioning mode may be referred to as the multi-tree partitioning type or partitioning type.

[0124] [Table 1]

[0125] MttSplitMode mtt_split_cu_vertical_flag mtt_split_cu_binary_flag SPLIT_TT_HOR 0 0 SPLIT_BT_HOR 0 1 SPLIT_TT_VER 1 0 SPLIT_BT_VER 1 1

[0126] Figure 7This is a view illustrating an example of partitioning a CTU into multiple CUs by applying a multi-type tree after applying a quadtree. Figure 7 In the diagram, bold border 710 represents quadtree partitioning, while remaining border 720 represents multi-type tree partitioning. A CU can correspond to a coded block (CB). In an implementation, a CU may include one coded block for luminance samples and two coded blocks for chrominance samples corresponding to the luminance samples. The size of the chrominance component (sample) CB or TB can be derived based on the component ratio of the color format (chrominance format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.) of the image / picture, based on the luminance component (sample) CB or TB size. In the case of a 4:4:4 color format, the chrominance component CB / TB size can be set to be equal to the luminance component CB / TB size. In the case of a 4:2:2 color format, the width of the chrominance component CB / TB can be set to half the width of the luminance component CB / TB, and the height of the chrominance component CB / TB can be set to the height of the luminance component CB / TB. In the 4:2:0 color format, the width of the chroma component CB / TB can be set to half the width of the luminance component CB / TB, and the height of the chroma component CB / TB can be set to half the height of the luminance component CB / TB.

[0127] In one implementation, when the size of the CTU is based on 128 luminance sample units, the size of the CU can be from 128×128 to 4×4, which is the same size as the CTU. In one implementation, in the case of a 4:2:0 color format (or chroma format), the chroma CB size can be from 64×64 to 2×2.

[0128] Furthermore, in implementations, the CU size and TU size can be the same. Alternatively, there can be multiple TUs within the CU region. The TU size typically represents the size of the luminance component (sample) transform block (TB).

[0129] The TU size can be derived based on the maximum permissible TB size, maxTbSize, which is a predetermined value. For example, when the CU size is greater than maxTbSize, multiple TUs (TBs) with maxTbSize can be derived from the CU, and transformations / inverse transformations can be performed on TUs (TBs) as units. For example, the maximum permissible luminance TB size can be 64×64 and the maximum permissible chrominance TB size can be 32×32. If the width or height of the CB segmented according to the tree structure is greater than the maximum transformation width or height, the CB can be automatically (or implicitly) segmented until the TB size limits in the horizontal and vertical directions are met.

[0130] Additionally, for example, when applying intra-frame prediction, the intra-frame prediction mode / type can be derived at the CU (or CB) level, and the neighbor reference sample derivation and prediction sample generation process can be performed at the TU (or TB) level. In this case, there can be one or more TUs (or TBs) in a CU (or CB) region, and multiple TUs or (TBs) can share the same intra-frame prediction mode / type.

[0131] Furthermore, for quadtree coding schemes with nested multi-type trees, the following parameters can be signaled from the encoding device to the decoding device as SPS syntax elements. For example, at least one of the following can be signaled: CTU size (representing the size of the quadtree root node), MinQTSize (representing the minimum allowed size of a quadtree leaf node), MaxBtSize (representing the maximum allowed size of a binary tree root node), MaxTtSize (representing the maximum allowed size of a ternary tree root node), MaxMttDepth (representing the maximum allowed depth of the multi-type tree partitioning starting from the quadtree leaf node), MinBtSize (representing the minimum allowed size of a binary tree leaf node), or MinTtSize (representing the minimum allowed size of a ternary tree leaf node).

[0132] As an implementation using the 4:2:0 chroma format, the CTU size can be set to 128×128 luma blocks and two corresponding 64×64 chroma blocks. In this case, MinOTSize can be set to 16×16, MaxBtSize to 128×128, MaxTtSzie to 64×64, MinBtSize and MinTtSize to 4×4, and MaxMttDepth to 4. Quadtree partitioning can be applied to the CTU to generate quadtree leaf nodes. Quadtree leaf nodes can be called leaf QT nodes. The size of quadtree leaf nodes can range from 16×16 (e.g., MinOTSize) to 128×128 (e.g., CTU size). If a leaf QT node is 128×128, it can be partitioned into a binary / ternary tree without additional partitioning. This is because, in this case, even if partitioned, it would exceed MaxBtsize and MaxTtszie (e.g., 64×64). In other cases, leaf QT nodes can be further segmented into multi-type trees. Therefore, a leaf QT node is the root node of the multi-type tree, and a leaf QT node can have a multi-type tree depth (mttDepth) of 0. If the multi-type tree depth reaches MaxMttdepth (e.g., 4), further segmentation can be disregarded. If the width of a multi-type tree node is equal to MinBtSize and less than or equal to 2xMinTtSize, further horizontal segmentation can be disregarded. If the height of a multi-type tree node is equal to MinBtSize and less than or equal to 2xMinTtSize, further vertical segmentation can be disregarded. When segmentation is disregarded, the encoding device can skip the signaling of segmentation information. In this case, the decoding device can deduce segmentation information with predetermined values.

[0133] Furthermore, a CTU can include a coded block of luma samples (hereinafter referred to as a "luma block") and two coded blocks of corresponding chroma samples (hereinafter referred to as "chroma blocks"). The above coding tree scheme can be applied equally or separately to the luma and chroma blocks of the current CU. Specifically, the luma and chroma blocks in a CTU can be partitioned into the same block tree structure, and in this case, the tree structure is represented as SINGLE_TREE. Alternatively, the luma and chroma blocks in a CTU can be partitioned into separate block tree structures, and in this case, the tree structure can be represented as DUAL_TREE. That is, when the CTU is partitioned into a dual-tree, the block tree structure for the luma block and the block tree structure for the chroma block can exist separately. In this case, the block tree structure for the luma block can be called DUAL_TREE_LUMA, and the block tree structure for the chroma components can be called DUAL_TREE_CHROMA. For P and B slice / piece groups, the luma and chroma blocks in a CTU can be restricted to having the same coding tree structure. However, for I-slice / patch groups, luma blocks and chroma blocks can have separate block tree structures. If separate block tree structures are applied, the luma CTB can be segmented into CUs based on a specific coding tree structure, and the chroma CTB can be segmented into chroma CUs based on another coding tree structure. That is, this means that CUs in I-slice / patch groups with separate block tree structures can include either a coded block of the luma component or coded blocks of the two chroma components, and CUs in P or B-slice / patch groups can include blocks of three color components (one luma component and two chroma components).

[0134] Although quadtree coding tree structures with nested multi-type trees have been described, the structures for splitting CUs are not limited to this. For example, BT and TT structures can be interpreted as concepts included in a multi-split tree (MPT) structure, and CUs can be interpreted as being split via QT and MPT structures. In the example of splitting CUs via QT and MPT structures, the splitting structure can be determined by signaling syntax elements (e.g., MPT_split_type) that include information about how many blocks the leaf nodes of the QT structure are split into, and syntax elements (e.g., MPT_split_mode) that include information about which direction (vertical or horizontal) the leaf nodes of the QT structure are split into.

[0135] In another example, the CU can be segmented in a manner different from the QT, BT, or TT structures. That is, unlike the QT structure which segments a lower-depth CU into 1 / 4 of a higher-depth CU, the BT structure which segments a lower-depth CU into 1 / 2 of a higher-depth CU, or the TT structure which segments a lower-depth CU into 1 / 4 or 1 / 2 of a higher-depth CU, in some cases the lower-depth CU can be segmented into 1 / 5, 1 / 3, 3 / 8, 3 / 5, 2 / 3, or 5 / 8 of the higher-depth CU, and the method of segmenting the CU is not limited to these.

[0136] Quadtree coded block structures with multiple tree types can provide highly flexible block partitioning structures. Due to the partitioning types supported in the multiple tree types, different partitioning patterns can potentially produce the same coded block structure in some cases. In encoding and decoding devices, the amount of data for partitioning information can be reduced by limiting the occurrence of such redundant partitioning patterns.

[0137] Furthermore, in the encoding and decoding of video / images according to this document, the image processing foundation can have a hierarchical structure. An image can be divided into one or more tiles, blocks, slices, and / or tile groups. A slice may include one or more blocks. A block may include one or more CTU rows within a tile. A slice may include blocks of an image, wherein the number of blocks is an integer. A tile group may include one or more tiles. A tile may include one or more CTUs. A CTU may be divided into one or more CUs. A tile may be a quadrilateral region within an image consisting of specific tile rows and specific tile columns composed of multiple CTUs. A tile group may include tiles raster scanned according to tiles within an image, wherein the number of tiles is an integer. A slice header may carry information / parameters applicable to the corresponding slice (blocks within a slice). When the encoding or decoding device has a multi-core processor, the encoding / decoding processes for tiles, slices, blocks, and / or tile groups can be executed in parallel.

[0138] In this disclosure, the names or concepts of slice or tile group can be used interchangeably. That is, the tile group header can be called the slice header. Here, a slice can have one of the slice types, including intra-frame (I) slices, prediction (P) slices, and bidirectional prediction (B) slices. For blocks within an I slice, inter-frame prediction is not used for prediction, and only intra-frame prediction can be used. Even in this case, the original sample values ​​can be encoded and signaled without prediction. For blocks within a P slice, either intra-frame prediction or inter-frame prediction can be used. When using inter-frame prediction, only unidirectional prediction can be used. Furthermore, for blocks within a B slice, either intra-frame prediction or inter-frame prediction can be used. When using inter-frame prediction, bidirectional prediction, as the maximum range, can be used.

[0139] Based on the characteristics of the video image (e.g., resolution) or considering coding efficiency or parallel processing, the encoding device can determine the tile / tile group, brick, slice, and the maximum and minimum coding unit size. Furthermore, information about this, or information used to derive this, can be included in the bitstream.

[0140] The decoding device can obtain information indicating whether a tile / tile group, brick, slice, or CTU within a tile of the current image has been divided into multiple coding units. The encoding and decoding devices only signal this information under specific conditions, thereby improving encoding efficiency.

[0141] A slice header (slice header syntax) can include information / parameters commonly applicable to a slice. An APS (APS syntax) or PPS (PPS syntax) can include information / parameters commonly applicable to one or more images. An SPS (SPS syntax) can include information / parameters commonly applicable to one or more sequences. A VPS (VPS syntax) can include information / parameters commonly applicable to multiple layers. A DPS (DPS syntax) can include information / parameters commonly applicable to the entire video. A DPS can include information / parameters related to the combination of encoded video sequences (CVS).

[0142] Additionally, information regarding the segmentation and configuration of tiles / tile groups / bricks / slices can be constructed using high-level syntax during the encoding stage and sent to the decoding device as a bitstream.

[0143] Overview of Intra-Frame Prediction

[0144] The intra-frame prediction performed by the aforementioned encoding and decoding devices will be described in detail below. Intra-frame prediction can refer to the prediction of a block based on reference samples in the image to which the current block belongs (hereinafter, the current image).

[0145] Reference Figure 8 This is described below. When intra-prediction is applied to the current block 801, neighboring reference samples to be used for intra-prediction of the current block 801 can be derived. The neighboring reference samples of the current block may include: a total of 2×nH samples including sample 811 adjacent to the left boundary of the current block of size nW×nH and sample 812 adjacent to the lower left side; a total of 2×nW samples including sample 821 adjacent to the upper boundary of the current block and sample 822 adjacent to the upper right side; and a sample 831 adjacent to the upper left side of the current block. Alternatively, the neighboring reference samples of the current block may include multiple column upper neighbor samples and multiple row left neighbor samples.

[0146] Additionally, the neighboring reference samples of the current block may include: a total of nH samples 841 adjacent to the right boundary of the current block of size nW×nH; a total of nW samples 851 adjacent to the bottom boundary of the current block; and a sample 842 adjacent to the lower right side of the current block.

[0147] However, some of the neighboring reference samples of the current block may not have been decoded or may be unavailable. In this case, the decoding device can construct neighboring reference samples to be used for prediction by replacing unavailable samples with available samples. Alternatively, neighboring reference samples to be used for prediction can be constructed by interpolation of available samples.

[0148] When deriving neighboring reference samples, (i) the predicted sample can be derived based on the average or interpolation of the neighboring reference samples of the current block, or (ii) the predicted sample can be derived based on a reference sample existing in a specific (prediction) direction relative to the predicted sample among the neighboring reference samples of the current block. Case (i) can be referred to as non-directional mode or non-angular mode, and case (ii) can be referred to as directional mode or angular mode. Additionally, a predicted sample can be generated based on the predicted sample of the current block among the neighboring reference samples by interpolating a second neighboring sample and a first neighboring sample located in a direction opposite to the prediction direction of the intra-prediction mode of the current block. This case can be referred to as Linear Interpolation Intra-Prediction (LIP). Furthermore, a chroma predicted sample can be generated based on a luminance sample using a linear model. This case can be referred to as LM mode. Additionally, a temporary predicted sample of the current block can be derived based on filtered neighboring reference samples, and the predicted sample of the current block can be derived by weighted summing of at least one reference sample derived according to the intra-prediction mode from the temporary predicted sample and the existing neighboring reference samples (i.e., unfiltered neighboring reference samples). The above situation can be called position-dependent intra-prediction (PDPC). Alternatively, the reference sample line with the highest prediction accuracy can be selected from multiple neighboring reference sample lines of the current block to derive the prediction sample using reference samples in the prediction direction of the corresponding line. In this case, intra-prediction coding can be performed by indicating (signaling) the reference sample line used to the decoding device. This situation can be called multi-reference line (MRL) intra-prediction or MRL-based intra-prediction. Furthermore, the current block can be divided into vertical or horizontal sub-partitions to perform intra-prediction based on the same intra-prediction mode, and neighboring reference samples can be derived and used on a per-sub-partition basis. That is, in this case, the intra-prediction mode of the current block is applied equivalently to the sub-partitions, and neighboring reference samples are derived and used on a per-sub-partition basis, thereby improving intra-prediction performance in some cases. This prediction method can be called intra-partition (ISP) or ISP-based intra-prediction. These intra-prediction methods can be called intra-prediction types, distinguished from intra-prediction modes (e.g., DC mode, planar mode, and directional mode). Intra-prediction types can be referred to by various terms such as intra-prediction schemes or additional intra-prediction modes. For example, an intra-prediction type (or additional intra-prediction mode) can include at least one selected from the group consisting of LIP, PDPC, MRL, and ISP mentioned above. A general intra-prediction method that excludes specific intra-prediction types such as LIP, PDPC, MRL, and ISP can be referred to as a normal intra-prediction type. A normal intra-prediction type can refer to a case where no specific intra-prediction type is applied, and prediction can be performed based on the intra-prediction modes mentioned above. Furthermore, post-filtering can be performed on the derived prediction samples when necessary.

[0149] Specifically, the intra-frame prediction process may include an intra-frame prediction mode / type determination step, a neighboring reference sample derivation step, and a prediction sample derivation step based on the intra-frame prediction mode / type. Additionally, if necessary, a post-filtering step may be performed on the derived prediction samples.

[0150] In addition to the intra-prediction types mentioned above, affine linear weighted intra-prediction (ALWIP) can also be used. ALWIP can be referred to as linear weighted intra-prediction (LWIP), matrix weighted intra-prediction, or matrix-based intra-prediction (MIP). When applying MIP to the current block, the predicted samples for the current block can be derived by i) using neighboring reference samples that have undergone an averaging process, ii) performing a matrix-vector multiplication process, and further iii) performing horizontal / vertical interpolation processes if necessary. The intra-prediction mode used for MIP can be different from the intra-prediction mode used in LIP, PDPC, MRL, ISP intra-prediction, or normal intra-prediction. The intra-prediction mode used for MIP can be referred to as MIP intra-prediction mode, MIP prediction mode, or MIP mode. For example, different matrices and offsets used in matrix-vector multiplication can be set according to the intra-prediction mode used for MIP. Here, the matrix can be referred to as the (MIP) weighted matrix, and the offset can be referred to as the (MIP) offset vector or (MIP) bias vector. The detailed MIP method will be described later.

[0151] The block reconstruction process based on intra-prediction and intra-prediction units in the coding apparatus may schematically include, for example, the following: Step S8910 may be performed by the intra-prediction unit 185 of the coding apparatus. Step S920 may be performed by a residual processor, which includes at least one selected from the group consisting of a subtractor 115, a transformer 120, a quantizer 130, an inverse quantizer 140, and an inverse transformer 150 of the coding apparatus. Specifically, step S920 may be performed by the subtractor 115 of the coding apparatus. In step S930, prediction information may be derived by the intra-prediction unit 185, and the prediction information may be encoded by the entropy encoder 190. In step S930, residual information may be derived by the residual processor, and the residual information may be encoded by the entropy encoder 190. The residual information is information about residual samples. The residual information may include information about the quantization transform coefficients of the residual samples. As described above, the residual samples can be derived into transform coefficients by the transformer 120 of the encoding device, and the transform coefficients can be derived into quantized transform coefficients by the quantizer 130. The information about the quantized transform coefficients can be encoded by the entropy encoder 190 through the residual encoding process.

[0152] In step S910, the encoding device may perform intra-prediction on the current block. The encoding device derives the intra-prediction mode / type of the current block, derives the neighboring reference samples of the current block, and generates prediction samples in the current block based on the intra-prediction mode / type and the neighboring reference samples. Here, the processes of determining the intra-prediction mode / type, deriving the neighboring reference samples, and generating the prediction samples may be performed simultaneously, or one of the processes may be performed before the other processes. For example, although not shown, the intra-prediction unit 185 of the encoding device may include an intra-prediction mode / type determination unit, a reference sample derivation unit, and a prediction sample derivation unit. The intra-prediction mode / type determination unit may determine the intra-prediction mode / type of the current block, the reference sample derivation unit may derive the neighboring reference samples of the current block, and the prediction sample derivation unit may derive the prediction samples of the current block. Furthermore, when performing the prediction sample filtering process, which will be described later, the intra-prediction unit 185 may also include a prediction sample filter. The encoding device may determine the mode / type applicable to the current block from among a variety of intra-prediction modes / types. The encoding device can compare the RD costs of intra-prediction modes / types and determine the best intra-prediction mode / type for the current block.

[0153] In addition, the encoding device can perform a prediction sample filtering process. Prediction sample filtering can also be called post-filtering. Through the prediction sample filtering process, some or all of the prediction samples can be filtered. In some cases, the prediction sample filtering process can be omitted.

[0154] In step S920, the encoding device can generate residual samples for the current block based on the (filtered) residual samples. The encoding device can compare the predicted samples in the original samples of the current block based on the phase and can derive the residual samples.

[0155] In step S930, the encoding device can encode image information including information about intra-frame prediction (prediction information) and residual information about residual samples. The prediction information may include intra-frame prediction mode information and intra-frame prediction type information. The encoding device can output the encoded image information as a bitstream. The output bitstream can be sent to the decoding device via a storage medium or network.

[0156] The residual information may include the residual coding syntax, which will be described later. The encoding device can derive the quantization transform coefficients by transforming / quantizing the residual samples. The residual information may include information about the quantization transform coefficients.

[0157] Furthermore, as described above, the encoding device can generate a reconstructed image (including reconstructed samples and reconstructed blocks). For this purpose, the encoding device can perform inverse quantization / inverse transform on the quantized transform coefficients and derive (modified) residual samples. The reason for performing inverse quantization / inverse transform after the transform / quantization of the residual samples is to derive residual samples identical to those derived by the decoding device as described above. The encoding device can generate a reconstructed block including reconstructed samples of the current block based on the predicted samples and the (modified) residual samples. Based on the reconstructed blocks, a reconstructed image of the current image can be generated. As described above, a loop filtering process can also be applied to the reconstructed image.

[0158] The video / image decoding process based on intra-frame prediction and intra-frame prediction units in a decoding device may schematically include, for example, the following: The decoding device may perform operations corresponding to those performed by the encoding device.

[0159] Steps S1010 to S1030 can be executed by the intra-frame prediction unit 265 of the decoding device. The prediction information from step S1010 and the residual information from step S1040 can be obtained from the bitstream by the entropy decoder 210 of the decoding device. A residual processor, including the dequantizer 220 or the inverse transformer 230 of the decoding device, or both, can derive the residual samples of the current block based on the residual information. Specifically, the dequantizer 220 of the residual processor can perform dequantization based on the quantization transform coefficients derived from the residual information and can derive the transform coefficients. The inverse transformer 230 of the residual processor can perform an inverse transform on the transform coefficients and can derive the residual samples of the current block. Step S1050 can be executed by the adder 235 or the reconstructor of the decoding device.

[0160] Specifically, in step S1010, the decoding device can deduce the intra-prediction mode / type of the current block based on the received prediction information (intra-prediction mode / type information). In step S1020, the decoding device can deduce the neighboring reference samples of the current block. In step S1030, the decoding device can generate prediction samples in the current block based on the intra-prediction mode / type and the neighboring reference samples. In this case, the decoding device can perform a prediction sample filtering process. Prediction sample filtering can be called post-filtering. Through the prediction sample filtering process, some or all of the prediction samples can be filtered. In some cases, the prediction sample filtering process can be omitted.

[0161] The decoding device can generate residual samples for the current block based on the received residual information. In step S1040, the decoding device can generate reconstructed samples for the current block based on the predicted samples and residual samples, and can deduce a reconstructed block including the reconstructed samples. Based on the reconstructed block, a reconstructed image of the current image can be generated. As described above, a loop filtering process can also be applied to the reconstructed image.

[0162] Here, although not shown, the intra-prediction unit 265 of the decoding device may include an intra-prediction mode / type determination unit, a reference sample derivation unit, and a prediction sample derivation unit. The intra-prediction mode / type determination unit can determine the intra-prediction mode / type of the current block based on the intra-prediction mode / type information obtained from the entropy decoder 210. The reference sample derivation unit can derive neighboring reference samples of the current block. The prediction sample derivation unit can derive the prediction samples of the current block. Furthermore, when performing the prediction sample filtering process described above, the intra-prediction unit 265 may also include a prediction sample filter.

[0163] Intra-luma_mpm_flag may include, for example, flag information indicating whether the most probable mode (MPM) or other modes are applied to the current block. When an MPM is applied to the current block, the prediction mode information may also include index information indicating one of the intra-luma_mpm_idx prediction mode candidates. Intra-luma_mpm_idx prediction mode candidates may be constructed as an MPM candidate list or an MPM list. Additionally, when no MPM is applied to the current block, the intra-luma_mpm_remainder prediction mode information may include information indicating one of the remaining intra-luma_mpm_remainder prediction modes besides the MPM candidates. The decoding device may determine the intra-luma_mpm_remainder prediction mode for the current block based on the intra-luma_mpm_remainder prediction mode information. A separate MPM list may be constructed for the MIP described above.

[0164] Furthermore, intra-prediction type information can be implemented in various forms. For example, intra-prediction type information may include intra-prediction type index information indicating one of the intra-prediction types. As another example, intra-prediction type information may include reference sample line information (e.g., intra_luma_ref_idx) indicating whether MRL is applied to the current block and which reference sample line is used when MRL is applied to the current block, ISP flag information (e.g., intra_subpartitions_mode_flag) indicating whether ISP is applied to the current block, ISP type information (e.g., intra_subpartitions_split_flag) indicating the segmentation type of the subpartition when ISP is applied, flag information indicating whether PDCP is applied, or flag information indicating whether LIP is applied. Additionally, intra-prediction type information may include a MIP flag indicating whether MIP is applied to the current block.

[0165] Intra-prediction mode information and / or intra-prediction type information can be encoded / decoded using the coding methods described in this document. For example, intra-prediction mode information and / or intra-prediction type information can be encoded / decoded using entropy coding (e.g., CABAC, CAVLC) based on truncated (Rice) binary code.

[0166] Overview of inter-frame prediction

[0167] The following will describe in reference Figure 2 and Figure 3 This section provides a detailed description of the inter-frame prediction method in the descriptions of encoding and decoding. In the case of a decoding apparatus, a video / image decoding method based on inter-frame prediction can be executed, and the inter-frame prediction unit in the decoding apparatus operates according to the following description. In the case of an encoding apparatus, a video / image encoding method based on inter-frame prediction can be executed, and the inter-frame prediction unit in the encoding apparatus operates according to the following description. Furthermore, the data encoded according to the following description can be stored in the form of a bitstream.

[0168] The prediction unit of the encoding / decoding apparatus can derive prediction samples by performing inter-frame prediction on a per-block basis. Inter-frame prediction can refer to a prediction derived by a method that depends on data elements (e.g., sample values ​​or motion information) of images other than the current image. When applying inter-frame prediction to the current block, the prediction block (prediction sample array) of the current block can be derived based on the reference block (reference sample array) specified by the motion vector on the reference image indicated by the reference image index. Here, in order to reduce the amount of motion information transmitted in inter-frame prediction mode, the motion information of the current block can be predicted on a per-block, sub-block, or sample basis based on the motion information correlation between neighboring blocks and the current block. Motion information can include motion vectors and reference image indices. Motion information can also include inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. When applying inter-frame prediction, neighboring blocks can include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block can be the same or different. The temporally neighboring block can be referred to as a juxtaposed reference block or a juxtaposed CU (colCU). Reference images, including those of temporally neighboring blocks, can be referred to as colpics. For example, a candidate list of motion information can be constructed based on the current block's neighboring blocks. Signals can be sent indicating which candidate to select (use) to derive the motion vector of the current block and / or the reference image index, along with flags or index information. Inter-frame prediction can be performed based on various prediction modes. For example, in skip and merge modes, the motion information of the current block can be the same as the motion information of the selected neighboring blocks. In skip mode, unlike merge mode, residual signals may not be sent. In motion information prediction (motion vector prediction (MVP)) mode, the motion vectors of the selected neighboring blocks can be used as motion vector predictors, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictors and the motion vector difference.

[0169] Depending on the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.), motion information can include L0 motion information and / or L1 motion information. Motion vectors in the L0 direction can be referred to as L0 motion vectors or MVL0, and motion vectors in the L1 direction can be referred to as L1 motion vectors or MVL1. Prediction based on L0 motion vectors can be called L0 prediction. Prediction based on L1 motion vectors can be called L1 prediction. Prediction based on both L0 and L1 motion vectors can be called bidirectional prediction. Here, L0 motion vectors can refer to motion vectors associated with a reference image list L0 (L0), and L1 motion vectors can refer to motion vectors associated with a reference image list L1 (L1). The reference image list L0 can include images that precede the current image in the output order as reference images. The reference image list L1 can include images that follow the current image in the output order. Previous images can be referred to as forward (reference) images, and subsequent images can be referred to as backward (reference) images. The reference image list L0 can also include images that follow the current image in the output order as reference images. In this scenario, within the reference image list L0, previous images can be indexed first, followed by subsequent images. The reference image list L1 can also include images that precede the current image in the output order. In this case, within the reference image list L1, subsequent images can be indexed first, followed by previous images. Here, the output order can correspond to the Image Order Count (POC) order.

[0170] The video / image coding process based on inter-frame prediction and inter-frame prediction units in the coding apparatus may schematically include, for example, the following: (Refer to...) Figure 11This is described below. In step S1110, the encoding device can perform inter-frame prediction on the current block. The encoding device can deduce the inter-frame prediction mode and motion information of the current block, and can generate prediction samples for the current block. Here, the processes of determining the intra-frame prediction mode, deduce motion information, and generate prediction samples can be performed simultaneously, or any one of these processes can be performed before the others. For example, the inter-frame prediction unit of the encoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit can determine the prediction mode of the current block. The motion information derivation unit can deduce the motion information of the current block. The prediction sample derivation unit can deduce the prediction samples for the current block. For example, the inter-frame prediction unit of the encoding device can search a predetermined region (search region) of a reference image for blocks similar to the current block through motion estimation, and can deduce a reference block whose difference from the current block is the smallest or less than or equal to a predetermined standard. Based on this, a reference image index indicating the reference image where the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding device can determine the pattern applicable to the current block from various prediction patterns. The encoding device can compare the RD costs of various prediction patterns and determine the best prediction pattern for the current block.

[0171] For example, when applying a skip mode or a merge mode to the current block, the encoding device can construct a merge candidate list, described later, and deduce the reference block among the reference blocks indicated by the merge candidates included in the merge candidate list that has the smallest difference from the current block or is less than or equal to a predetermined standard. In this case, a merge candidate associated with the deduced reference block can be selected, and merge index information indicating the selected merge candidate can be generated and signaled to the decoding device. The motion information of the current block can be deduced using the motion information of the selected merge candidate.

[0172] As another example, when applying the (A)MVP mode to the current block, the encoding device can construct an (A)MVP candidate list, described later, and use the motion vector of the selected MVP candidate from the MVP candidates included in the (A)MVP candidate list as the motion vector predictor (MVP) for the current block. In this case, for example, the motion vector of the reference block derived through the motion estimation described above can be used as the motion vector of the current block. Among the MVP candidates, the MVP candidate with the motion vector having the smallest difference from the motion vector of the current block can be the selected MVP candidate. The motion vector difference (MVD) can be derived as the difference obtained by subtracting the MVP from the motion vector of the current block. In this case, information about the MVD can be signaled to the decoding device. In addition, when applying the (A)MVP mode, the value of the reference image index can be constructed as reference image index information and can be signaled to the decoding device separately.

[0173] In step S1120, the encoding device can derive residual samples based on the predicted samples. The encoding device can compare the predicted samples of the current block with the original samples to derive the residual samples.

[0174] In step S1130, the encoding device can encode image information including prediction information and residual information. The encoding device can output the encoded image information in the form of a bitstream. Prediction information is information related to the prediction process and may include prediction mode information (e.g., skip flag, merge flag, or mode index) and information about motion information. Information about motion information may include candidate selection information (e.g., merge index, MVP flag, or MVP index) as information for deriving motion vectors. Additionally, information about motion information may include the aforementioned information about MVD and / or reference image index information. Furthermore, information about motion information may include information indicating whether L0 prediction, L1 prediction, or bidirectional prediction is applied. Residual information is information about residual samples. Residual information may include information about the quantization transform coefficients of the residual samples.

[0175] The output bitstream can be stored in (digital) storage media and sent to the decoding device, or it can be sent to the decoding device via a network.

[0176] Furthermore, as described above, the encoding device can generate a reconstructed image (including reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This is to derive a prediction result identical to the prediction result performed in the decoding device by the encoding device, thereby improving encoding efficiency. Therefore, the encoding device can store the reconstructed image (or reconstructed samples, reconstructed blocks) in memory and can use it as a reference image for inter-frame prediction. As described above, a loop filtering process can also be applied to the reconstructed image.

[0177] The video / image decoding process based on inter-frame prediction and inter-frame prediction units in the decoding device may schematically include, for example, the following:

[0178] The decoding device can perform operations corresponding to those performed by the encoding device. The decoding device can perform predictions on the current block based on received prediction information and can derive prediction samples.

[0179] Specifically, in step S1210, the decoding device can determine the prediction mode of the current block based on the received prediction information. The decoding device can determine which inter-frame prediction mode to apply to the current block based on the prediction mode information in the prediction information.

[0180] For example, based on the merge flag, it can be determined whether to apply a merge mode or determine the (A)MVP mode to the current block. Alternatively, one of a variety of inter-frame prediction mode candidates can be selected based on a mode index. Inter-frame prediction mode candidates may include skip mode, merge mode, and / or (A)MVP mode, or may include a variety of inter-frame prediction modes described later.

[0181] In step S1220, the decoding device can deduce the motion information of the current block based on the determined inter-frame prediction mode. For example, when a skip mode or merge mode is applied to the current block, the decoding device can construct a merge candidate list, described later, and select a merge candidate from the merge candidates included in the merge candidate list. The selection can be performed based on the selection information (merge index) described above. The motion information of the current block can be deduced using the motion information of the selected merge candidate. The motion information of the selected merge candidate can be used as the motion information of the current block.

[0182] As another example, when applying (A)MVP mode to the current block, the decoding device can construct an (A)MVP candidate list, described later, and use the motion vector of the selected MVP candidate from the MVP candidates included in the (A)MVP candidate list as the motion vector predictor (MVP) for the current block. Selection can be performed based on the aforementioned selection information (MVP flag or MVP index). In this case, the MVD of the current block can be derived based on information about the MVD, and the motion vector of the current block can be derived based on the MVP and MVD of the current block. Additionally, the reference image index of the current block can be derived based on reference image index information. The image indicated by the reference image index in the reference image list associated with the current block can be derived as the reference image referenced by the inter-frame prediction of the current block.

[0183] Furthermore, as will be described later, the motion information of the current block can be derived without constructing a candidate list. In this case, the motion information of the current block can be derived based on the process described in the prediction mode, which will be described later. In this case, the candidate list configuration described above can be omitted.

[0184] In step S1230, the decoding device can generate a predicted sample for the current block based on the motion information of the current block. In this case, a reference image can be derived based on the reference image index of the current block, and the predicted sample for the current block can be derived using samples of the reference block indicated by the motion vector of the current block on the reference image. In this case, as described later, in some situations, a predicted sample filtering process can be further performed on all or some of the predicted samples of the current block.

[0185] For example, the inter-frame prediction unit of the decoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit can determine the prediction mode of the current block based on the received prediction mode information. The motion information derivation unit can derive the motion information (motion vector and / or reference image index) of the current block based on the received motion information. The prediction sample derivation unit can derive the prediction samples of the current block.

[0186] In step S1240, the decoding device can generate residual samples for the current block based on the received residual information. In step S1250, the decoding device can generate reconstructed samples for the current block based on the predicted samples and residual samples, and can generate a reconstructed image based on this. Thereafter, as described above, a loop filtering process can be further applied to the reconstructed image.

[0187] As described above, the inter-frame prediction process may include the steps of determining an inter-frame prediction mode, deriving motion information based on the determined prediction mode, and performing prediction (generating prediction samples) based on the derived motion information. The inter-frame prediction process may be performed in the encoding and decoding apparatus described above.

[0188] Quantization / Inverse Quantization

[0189] As described above, the quantizer of the encoding device can derive the quantized transform coefficients by applying quantization to the transform coefficients. Similarly, the dequantizer of the encoding device or the dequantizer of the decoding device can derive the transform coefficients by applying dequantization to the quantized transform coefficients.

[0190] In video / still image encoding and decoding, the quantization ratio can be changed, and the compression ratio can be adjusted using the changed quantization ratio. From an implementation perspective, considering complexity, a quantization parameter (QP) can be used instead of the quantization ratio directly. For example, a quantization parameter with integer values ​​from 0 to 63 can be used, and each quantization parameter value can correspond to the actual quantization ratio. Additionally, the quantization parameter QP for the luma component (luma sample) can also be used. Y Quantization parameter QP of chromaticity components (chromaticity samples) C They can be set to be different from each other.

[0191] In quantization, the transform coefficients C can be input and divided by the quantization ratio Qstep, yielding the quantized transform coefficients C'. In this case, considering computational complexity, the quantization ratio can be multiplied by a scale to an integer value, and a shift operation can be performed using the corresponding scale value. The quantization scale can be derived based on the multiplication of the quantization ratio and the scale value. That is, the quantization scale can be derived from QP. The quantization scale can be applied to the transform coefficients C, and based on this, the quantized transform coefficients C' can be derived.

[0192] Dequantization is the inverse process of quantization. The quantized transform coefficients C' can be multiplied by the quantization ratio Qstep, and based on this, the reconstructed transform coefficients C'' can be obtained. In this case, the level scale can be derived from the quantization parameters. The level scale can be applied to the quantized transform coefficients C', and based on this, the reconstructed transform coefficients C'' can be derived. Due to losses in the transform and / or quantization processes, the reconstructed transform coefficients C'' may differ slightly from the original transform coefficients C. Therefore, the encoding device can perform dequantization in the same way as in the decoding device.

[0193] Furthermore, adaptive frequency-weighted quantization (IFQ) can be applied, where the quantization strength is adjusted according to the frequency. IFQ is a method that applies a frequency-dependent quantization strength. In IFQ, the frequency-dependent quantization strength can be applied using a predefined quantization scaling matrix. That is, the aforementioned quantization / dequantization processing can also be performed based on the quantization scaling matrix. For example, different quantization scaling matrices can be used depending on the size of the current block and / or whether the prediction mode applied to the current block to generate the residual signal of the current block is inter-frame prediction or intra-frame prediction. The quantization scaling matrix can be referred to as the quantization matrix or the scaling matrix. The quantization scaling matrix can be predefined. Additionally, for frequency-adaptive scaling, the frequency-specific quantization scaling information used for the quantization scaling matrix can be constructed / encoded by the encoding device and can be signaled to the decoding device. The frequency-specific quantization scaling information can be referred to as quantization scaling information. The frequency-specific quantization scaling information can include scaling list data. The (modified) quantization scaling matrix can be derived based on the scaling list data. Furthermore, the frequency-specific quantization scaling information can include presence flags indicating the presence of scaling list data. Alternatively, when the zoom list data is signaled at a higher level (e.g., SPS), it may further include information indicating whether the zoom list data has been modified at a lower level (e.g., PPS or tile group headers).

[0194] Transform / Inverse Transform

[0195] As described above, the encoding device can derive residual blocks (residual samples) based on blocks (predicted samples) predicted via intra / inter / IBC prediction, and can derive quantization transform coefficients by applying transform and quantization to the derived residual samples. Information about the quantization transform coefficients (residual information) included in the residual coding syntax can be encoded and output as a bitstream. The decoding device can obtain information about the quantization transform coefficients (residual information) from the bitstream and derive the quantization transform coefficients by performing decoding. The decoding device can derive residual samples by inverse quantization / inverse transform based on the quantization transform coefficients. As described above, quantization / inverse quantization or transform / inverse transform, or both, can be omitted. When transform / inverse transform is omitted, transform coefficients can be referred to as coefficients or residual coefficients, or, for consistency, can still be referred to as transform coefficients. A transform skip flag (e.g., transform_skip_flag) can be used to signal whether transform / inverse transform is omitted.

[0196] Transform / inverse transforms can be performed based on transform kernels. For example, a multiple transform selection (MTS) scheme for performing transform / inverse transforms can be applied. In this case, some of the multiple transform kernels in the set can be selected and applied to the current block. Transform kernels can be referred to by various terms such as transform matrix, transform type, etc. For example, a transform kernel set can refer to a combination of vertical transform kernels (vertical transform kernels) and horizontal transform kernels (horizontal transform kernels).

[0197] Transform / inverse transform can be performed on a per-CU or per-TU basis. That is, a transform / inverse transform can be applied to residual samples in a CU or a residual sample in a TU. The CU size and TU size can be the same, or multiple TUs can exist in a CU region. Furthermore, the CU size can typically refer to the size of the luma component (sample) CB. The TU size can typically refer to the size of the luma component (sample) TB. The chroma component (sample) CB or TB size can be derived based on the component ratio according to the color format (chroma format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.) based on the luma component (sample) CB or TB size. The TU size can be derived based on maxTbSize. For example, when the CU size is greater than maxTbSize, multiple TUs (TBs) with maxTbSize can be derived from the CU, and a transform / inverse transform can be performed on a per-TU (TB) basis. maxTbSize can be considered when determining whether to apply various intra-frame prediction types such as ISP. Information about maxTbSize can be predetermined. Alternatively, information about maxTbSize can be generated and encoded by the encoding device and signaled to the encoding device.

[0198] Entropy coding

[0199] As referenced above Figure 2 The described information suggests that some or all of the video / image information can be encoded by an entropy encoder with 190 entropy. (Reference) Figure 3 Some or all of the described video / image information can be entropy decoded by the entropy decoder 310. In this case, the video / image information can be encoded / decoded on a per-syntax element basis. In this document, the encoding / decoding of information can include encoding / decoding performed by the methods described in this paragraph.

[0200] Figure 13A block diagram of CABAC for encoding a syntax element is shown. In CABAC encoding, firstly, when the input signal is a non-binary syntax element, it is transformed into a binary value through binarization. When the input signal is already a binary value, the binarization is bypassed. Here, each binary number 0 or 1 that constitutes the binary value can be called a bin. For example, the binarized binary string (bin string) is 110, and each of 1, 1, and 0 is called a bin. The bin of a syntax element can refer to the value of the syntax element.

[0201] Binary bins can be input into either a regular encoding engine or a bypass encoding engine. A regular encoding engine assigns a context model with applied probability values ​​to the corresponding bin and encodes the bin based on the assigned context model. After encoding each bin, the regular encoding engine can update the bin's probability model. Bins encoded in this way can be called context-encoded bins. A bypass encoding engine can omit the processes for estimating the probabilities of the input bins and for updating the probability models applied to the bins after encoding. In the case of a bypass encoding engine, the encoding rate is improved by applying a uniform probability distribution (e.g., 50:50) to the input bins instead of assigning context. Bins encoded in this way can be called bypass bins. Context models can be assigned and updated for each bin to be context-encoded (regular encoding), and the context model can be indicated based on ctxidx or ctxInc. ctxidx can be derived based on ctxInc. Specifically, for example, the context index (ctxidx) indicating the context model for each of the bins used for regular encoding can be derived as the sum of the context index increment (ctxInc) and the context index offset (ctxIdxOffset). Here, ctxInc, which varies from bin to bin, can be derived. ctxIdxOffset can be represented by the lowest value of ctxIdx. The lowest value of ctxIdx can be called the initial value (initValue) of ctxIdx. ctxIdxOffset is a value typically used to distinguish the context model from other syntax elements, and the context model of a syntax element can be distinguished / derived based on ctxinc.

[0202] During entropy encoding, it is determined whether encoding is performed using a regular encoding engine or a bypass encoding engine, and the encoding path can be switched. Entropy decoding can perform the same processing as entropy encoding in reverse order.

[0203] For example, it can be like Figure 14 and Figure 15 The entropy encoding described above is performed in the central region. (Refer to...) Figure 14 and Figure 15 An encoding device (entropy encoder) can perform entropy coding on image / video information. This image / video information may include segmentation-related information, prediction-related information (e.g., inter-frame / intra-frame prediction classification information, intra-frame prediction mode information, and inter-frame prediction mode information), residual information, and loop filtering-related information, or it may include various associated syntax elements. Entropy coding can be performed on a per-syntax-element basis. Figure 14 S1410 to S1420 can be derived from the above. Figure 2 The entropy encoder 190 of the encoding device is executed.

[0204] In step S1410, the encoding device may perform binary conversion on the target syntax element. Here, the binary conversion may be based on various binary conversion methods such as truncated Rice binary conversion or fixed-length binary conversion, and the binary conversion method used for the target syntax element may be predefined. The binary conversion process may be performed by the binary conversion unit 191 in the entropy encoder 190.

[0205] In step S1420, the encoding device may perform entropy encoding on the target syntax element. The encoding device may perform regular encoding (context-based) or bypass encoding on the bin string of the target syntax element based on an entropy encoding scheme such as Context Adaptive Arithmetic Encoding (CABAC) or Context Adaptive Variable Length Encoding (CAVLC). The output may be included in the bitstream. The entropy encoding process may be executed by the entropy encoding processor 192 in the entropy encoder 190. As described above, the bitstream may be sent to the decoding device via a (digital) storage medium or network.

[0206] Reference Figure 16 and Figure 17 The decoding device (entropy decoder) can decode the encoded image / video information. The image / video information may include segmentation-related information, prediction-related information (e.g., inter-frame / intra-frame prediction classification information, intra-frame prediction mode information, and inter-frame prediction mode information), residual information, and loop filtering-related information, or may include various related syntax elements. Entropy coding can be performed on a per-syntax-element basis. Steps S1610 to S1620 can be performed as described above. Figure 3 The entropy decoder 210 of the decoding device is executed.

[0207] In step S1610, the decoding device may perform binaryization on the target syntax element. Here, binaryization may be based on various binaryization methods such as truncated Rice binaryization and fixed-length binaryization, and the binaryization method used for the target syntax element may be predefined. The decoding device may deduce available binaries (binary candidate strings) of available values ​​for the target syntax element through the binaryization process. The binaryization process may be performed by the binaryization unit 211 in the entropy decoder 210.

[0208] In step S1620, the decoding device can perform entropy decoding on the target syntax element. The decoding device sequentially decodes and parses each bin of the target syntax element from the input bits in the bitstream, and can compare the derived bin string with the available bin strings of the syntax element. When the derived bin string matches one of the available bin strings, the value corresponding to the bin string can be deduced as the value of the syntax element. Otherwise, the next bit in the bitstream is further parsed, and the above process is repeated. Through this process, the start or end bits of specific information (specific syntax elements) in the bitstream are not used, but variable-length bits are used to signal information. Accordingly, relatively fewer bits are assigned to low values, thereby improving overall encoding efficiency.

[0209] The decoding device can perform context-based or bypass-based decoding on each bin in the bin string of the bitstream based on an entropy coding scheme such as CABAC or CAVLC. The entropy decoding process can be executed by the entropy decoding processor 212 in the entropy decoder 210. The bitstream can include various types of information used for image / video decoding as described above. As mentioned above, the bitstream can be sent to the decoding device via a (digital) storage medium or a network.

[0210] In this document, a table including syntax elements (syntax table) can be used to represent signaling information from an encoding device to a decoding device. The order of syntax elements in the table used in this document may refer to the parsing order of syntax elements in the bitstream. The encoding device can construct and encode the syntax table so that the decoding device can parse the syntax elements in the parsing order. The decoding device can parse and decode the syntax elements of the syntax table from the bitstream in the parsing order, and thus obtain the values ​​of the syntax elements.

[0211] General image / video encoding process

[0212] In image / video coding, the images that make up an image / video can be encoded / decoded according to the decoding order in the sequence. The image order corresponding to the output order of the decoded images can be set to be different from the decoding order, and based on this, forward prediction and backward prediction can be performed when performing inter-frame prediction.

[0213] Figure 18 An example of an illustrative image decoding process applicable to the embodiments described in this document is shown. Figure 18 In the above reference, step S1810 can be obtained from the above reference. Figure 3 The entropy decoder 210 of the described decoding device is executed. Step S1820 can be executed by a prediction unit including an intra-frame prediction unit 265 and an inter-frame prediction unit 260. Step S1830 can be executed by a residual processor including an inverse quantizer 220 and an inverse transformer 230. Step S1840 can be executed by an adder 235. Step S1850 can be executed by a filter 240. Step S1810 may include the information decoding process described in this document. Step S1820 may include the inter-frame / intra-frame prediction process described in this document. Step S1830 may include the residual processing process described in this document. Step S1840 may include the block / picture reconstruction process described in this document. Step S1850 may include the loop filtering process described in this document.

[0214] refer to Figure 18 The image decoding process can be illustrated as shown in the reference above. Figure 3The process of obtaining image / video information from the bitstream in step S1810 (through decoding), the image reconstruction process in steps S1820 to S1840, and the loop filtering process for reconstructing the image in step S1850 are described. The image reconstruction process can be performed based on the prediction samples and residual samples obtained through inter-frame / intra-frame prediction in step S1820 as described in this document, and the residual processing in step S1830 (inverse quantization and inverse transform of quantization transform coefficients). For the reconstructed image generated by the image reconstruction process, a modified reconstructed image can be generated through a loop filtering process. The modified reconstructed image can be output as a decoded image and can be stored in the memory 250 of the decoding device or the decoded image buffer for later use as a reference image in the inter-frame prediction process during image decoding. In some cases, the loop filtering process can be omitted. In this case, the reconstructed image can be output as a decoded image and can be stored in the memory 250 of the decoding device or the decoded image buffer for later use as a reference image in the inter-frame prediction process during image decoding. As described above, the loop filtering process in step S1850 may include a deblocking filtering process, a sample adaptive offset (SAO) process, an adaptive loop filtering (ALF) process, and / or a bidirectional filter process. Some or all of these may be omitted. Alternatively, one or more of the deblocking filtering process, the sample adaptive offset (SAO) process, the adaptive loop filtering (ALF) process, and the bidirectional filter process may be applied sequentially, or all of them may be applied sequentially. For example, a deblocking filtering process may be applied to the reconstructed image, and then the SAO process may be performed. Alternatively, for example, a deblocking filtering process may be applied to the reconstructed image, and then the ALF process may be performed. This can be performed in the same manner as in the encoding device.

[0215] Figure 19 An example of an illustrative image encoding process applicable to the embodiments described in this document is shown. Figure 19 In step S1910, the above references can be included. Figure 2 The prediction unit of the intra-frame prediction unit 185 or inter-frame prediction unit 180 of the described coding apparatus is used for execution. Step S1920 can be executed by a residual processor including a transformer 120 and / or a quantizer 130. Step S1930 can be executed by an entropy encoder 190. Step S1910 may include the inter-frame / intra-frame prediction process described in this document. Step S1920 may include the residual processing process described in this document. Step S1930 may include the information encoding process described in this document.

[0216] Reference Figure 19 The image encoding process can be illustrated as shown in the references above. Figure 2The process described includes encoding information for image reconstruction (e.g., prediction information, residual information, and segmentation information) and outputting the information as a bitstream; generating a reconstructed image of the current image; and optionally applying loop filtering to the reconstructed image. The encoding device can derive (modified) residual samples from the quantization transform coefficients using dequantizer 140 and inverse transformer 150, and can generate a reconstructed image based on the prediction samples and (modified) residual samples as output by S1910. The generated reconstructed image can be the same as the reconstructed image generated by the decoding device described above. A loop filtering process can be performed on the reconstructed image to generate a modified reconstructed image. The modified reconstructed image can be stored in the decoded image buffer or memory 170. Similar to the case in the decoding device, the modified reconstructed image can later be used as a reference image in the inter-frame prediction process when encoding the image. As described above, in some cases, some or all of the loop filtering process can be omitted. When the loop filtering process is performed, the (loop) filtering-related information (parameters) can be encoded by entropy encoder 190 and output as a bitstream. Decoding devices can perform loop filtering based on filter-related information in the same way as encoding devices.

[0217] This loop filtering process reduces noise such as block artifacts (and ringing artifacts) generated during image / video encoding, and improves subjective / objective visual quality. Furthermore, since both the encoding and decoding devices perform the loop filtering process, they can derive the same prediction results, increasing the reliability of image encoding and reducing the amount of data sent for image encoding.

[0218] As described above, the image reconstruction process can be performed in both the decoding and encoding devices. Reconstructed blocks can be generated based on intra-frame prediction / inter-frame prediction on a block-by-block basis, and a reconstructed image including these blocks can be generated. When the current image / slice / tile group is an I-type image / slice / tile group, the blocks included in the current image / slice / tile group can be reconstructed based solely on intra-frame prediction. Furthermore, when the current image / slice / tile group is a P-type or B-type image / slice / tile group, the blocks included in the current image / slice / tile group can be reconstructed based on either intra-frame prediction or inter-frame prediction. In this case, inter-frame prediction can be applied to some blocks in the current image / slice / tile group, and intra-frame prediction can be applied to the remaining blocks. The color components of the image can include luma and chroma components. Unless explicitly limited herein, the methods and implementations presented herein can be applied to both luma and chroma components.

[0219] Examples of coding levels and structures

[0220] The encoded videos / images described in this document can be processed according to, for example, the encoding hierarchy and structure described later.

[0221] Figure 20 This is a view showing the hierarchical structure of an encoded image. An encoded image can be divided into the decoding process that manipulates the image and its own Video Coding Layer (VCL), a subsystem for sending and storing encoded information, and a Network Abstraction Layer (NAL) that exists between the VCL and the subsystems and is responsible for network adaptation functions.

[0222] In VCL, VCL data including compressed image data (slice data) can be generated. Alternatively, parameter sets or supplementary enhancement information (SEI) messages additionally required for image decoding processing can be generated, including information such as picture parameter sets (PPS), sequence parameter sets (SPS), and video parameter sets (VPS).

[0223] In the NAL, header information (NAL unit header) is added to the raw byte sequence payload (RBSP) generated in the VCL, enabling the generation of NAL units. Here, RBSP refers to the slice data, parameter set, and SEI message generated in the VCL. The NAL unit header may include NAL unit type information specified based on the RBSP data included in the NAL unit.

[0224] As shown in the figure, based on the RBSP generated in the VCL, NAL units can be divided into VCL NAL units and non-VCL NAL units. A VCL NAL unit can refer to a NAL unit that includes information about the image (slice data). A non-VCL NAL unit can refer to a NAL unit that includes information required for decoding the image (parameter set or SEI message).

[0225] Using header information appended according to the subsystem's data standard, VCL NAL units and non-VCL NAL units can be transmitted over a network. For example, NAL units can be transformed into data in the form of predetermined standards such as H.266 / VVC file format, Real-time Transport Protocol (RTP), and Transport Stream (TS), and the resulting data can be transmitted over various networks.

[0226] As mentioned above, regarding NAL cells, the NAL cell type can be specified based on the RBSP data structure included in the NAL cell, and information about the NAL cell type can be stored in the NAL cell header and signaled.

[0227] For example, a rough classification of VCL NAL unit types and non-VCL NAL unit types can be made based on whether the NAL unit includes information about the image (slice data). VCL NAL unit types can be classified according to the characteristics and type of the image included in the VCL NAL unit, while non-VCL NAL unit types can be classified according to the type of parameter set.

[0228] Below, as an example, the NAL cell types specified according to the types of the parameter set included in the non-VCL NAL cell type are listed.

[0229] -APS (Adaptive Parameter Set) NAL Unit: The type of NAL unit including APS.

[0230] -DPS (Decoding Parameter Set) NAL Unit: The type of NAL unit including DPS.

[0231] -VPS (Video Parameter Set) NAL Unit: Includes the type of NAL unit for the VPS.

[0232] -SPS (Sequence Parameter Set) NAL Unit: The type of NAL unit that includes SPS.

[0233] -PPS (Image Parameter Set) NAL Unit: The type of NAL unit including PPS.

[0234] The aforementioned NAL unit types have syntax information specific to the NAL unit type, and this syntax information can be stored in the NAL unit header and signaled. For example, the syntax information can be `nal_unit_type`, and the NAL unit type can be specified through the `nal_unit_type` value.

[0235] A slice header (slice header syntax) may include information / parameters commonly applicable to a slice. An APS (APS syntax) or PPS (PPS syntax) may include information / parameters commonly applicable to one or more slices or images. An SPS (SPS syntax) may include information / parameters commonly applicable to one or more sequences. A VPS (VPS syntax) may include information / parameters commonly applicable to multiple layers. A DPS (DPS syntax) may include information / parameters commonly applicable to the entire video. A DPS may include information / parameters related to the concatenation of encoded video sequences (CVS). In this document, the High-Level Syntax (HLS) may include at least one selected from the group consisting of APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, and slice header syntax.

[0236] In this document, the image / video information encoded by the encoding device and sent to the decoding device in the form of a bitstream may include segmentation-related information within the image, intra / inter-frame prediction information, residual information and loop filtering information, as well as information included in the slice header, information included in the APS, information included in the PPS, information included in the SPS and / or information included in the VPS.

[0237] Overview of Adaptive Color Transformation (ACT)

[0238] Adaptive Color Transform (ACT) is a color space transformation (conversion) technique used to eliminate unnecessary overlap between color components, and it is already used in HEVC Screen Content Extended Edition. It can also be applied to VVC.

[0239] In HEVC Screen Content Extension (HEVC SCC Extension), ACT is used to adaptively transform the prediction residual from the existing color space to the YCgCo color space. One of the two color spaces can be optionally selected by signaling an ACT flag for each transformation basis.

[0240] For example, a first value of the flag (e.g., 1) can indicate that the residuals of the transform basis are encoded in the original color space. A second value of the flag (e.g., 1) can indicate that the residuals of the transform basis are encoded in the YCgCo color space.

[0241] Figure 21 This is a view illustrating an implementation of the decoding process using ACT. Figure 21 In some implementations, motion compensation prediction may correspond to inter-frame prediction in this disclosure.

[0242] like Figure 21 As shown, the reconstructed image (or reconstructed block, reconstructed sample array, reconstructed sample, reconstructed signal) can be generated based on the predicted output value and the residual output value. Here, the residual output value can be the inverse transform output value. Here, the inverse transform can be the normal inverse transform. Here, the normal inverse transform can be the inverse transform based on MTS or the inverse low-frequency non-separable transform (LFNST).

[0243] Here, the predicted output value can be a predicted block, a predicted sample array, a predicted sample, or a predicted signal. The residual output value can be a residual block, a residual sample array, a residual sample, or a residual signal.

[0244] For example, in terms of the encoding device, ACT processing can be performed on the residual samples derived from the predicted samples. Alternatively, the output of the ACT processing can be used as input for normal transform processing. Here, normal transform processing can be an MTS-based transform or an LFNST-based transform.

[0245] Information (parameters) about (inverse)ACT can be generated and encoded by the encoding device and sent to the decoding device in the form of a bit stream.

[0246] The decoding device can obtain, parse, and decode (inverse)ACT related information (parameters), and can perform inverse ACT based on (inverse)ACT related information (parameters).

[0247] Based on the inverse ACT, (modified) residual samples (or residual blocks) can be derived. For example, (transform) coefficients can be derived by applying inverse quantization to the quantization (transform) coefficients. Additionally, residual samples can be derived by performing an inverse transform on the (transform) coefficients. Furthermore, (modified) residual samples can be obtained by applying the inverse ACT to the residual samples. Information (parameters) regarding the (inverse) ACT will be described in detail later.

[0248] In implementations, the kernel transform function used in HEVC can be used as the kernel transform function (transform kernel) for color space transformation. For example, matrices for forward and backward transformations, as shown in the following formula, can be used.

[0249] [Formula 1]

[0250]

[0251] [Equation 2]

[0252]

[0253] Here, C0, C1, and C2 can correspond to G, B, and R. Here, G represents the green component, B represents the blue component, and R represents the red component. Additionally, C0', C1', and C2' can correspond to Y, Cg, and Co. Here, Y represents lightness, Cg represents green chromaticity, and Co represents the orange chromaticity component.

[0254] Additionally, to compensate for the change in dynamic range of the residuals before and after color transformation, QP adjustment can be applied to the transformation residuals using (-5, -5, -3). Details of QP adjustment will be described later.

[0255] Furthermore, in the encoding and decoding processes according to the implementation method, when ACT is applicable, the following limitations may be applied.

[0256] - ACT is disabled in the case of dual-tree encoding / decoding. For example, ACT can be applied only to single-tree encoding / decoding.

[0257] - ACT can be disabled when ISP encoding and decoding are applied.

[0258] - ACT can be disabled for chroma blocks that apply BDPCM. ACT can only be enabled for luma blocks that apply BDPCM.

[0259] - CCLM can be disabled when it is possible to apply ACT.

[0260] Figure 22 This is a view illustrating an implementation of a sequence parameter set syntax table in which the grammar elements associated with the ACT are signaled.

[0261] Figures 23 to 29 It is a view that continuously illustrates an implementation of a syntax table in which the encoding basis for signaling grammatical elements related to ACT is shown.

[0262] like Figure 22 As shown, the ACT enabling flag, which indicates whether ACT is enabled during decoding, can be sps_act_enabled_flag 2210.

[0263] The first value of sps_act_enabled_flag (e.g., 0) can indicate that ACT is not used, and the flags cu_act_enabled_flag 2310, 2910 indicating whether ACT is applied in the coding base are not provided in the syntax for the coding base.

[0264] The second value of sps_act_enabled_flag (e.g., 1) can indicate that ACT can be used, and cu_act_enabled_flag can be provided in the syntax for the encoding base.

[0265] When sps_act_enabled_flag is not available in the bitstream, the value of sps_act_enabled_flag can be deduced to be the first value (e.g., 0).

[0266] In addition, such as Figure 23 As shown, the ACT flag, which indicates whether the residual of the current encoding basis is encoded in the YCgCo color space, can be cu_act_enabled_flag 2310, 2910.

[0267] A first value for cu_act_enabled_flag (e.g., 0) can indicate that the residual of the current encoding basis is encoded in the original color space. A second value for cu_act_enabled_flag (e.g., 1) can indicate that the residual of the current encoding basis is encoded in the YCgCo color space.

[0268] When the cu_act_enabled_flag is not provided in the bitstream, the flag can be deduced to a first value (e.g., 0). Here, the original color space can be the RGB color space.

[0269] The QP derivation method based on the transformation of ACT QP offset.

[0270] In the implementation, the quantization parameter derivation and Qp update processes in the scaling process for the transform coefficients can be performed as follows. For example, the quantization parameter derivation process can be performed using the following parameters.

[0271] - Luminosity coordinates (xCb, yCb), which indicate the relative coordinates of the top-left luminosity sample of the current coded block with respect to the top-left luminosity sample of the current image.

[0272] - The variable cbWidth indicates the width of the current coded block based on each luminance sample.

[0273] - The variable cbHeight indicates the height of the current coded block based on each luminance sample.

[0274] - The variable `treeType` indicates whether a single tree (SINGLE_TREE) or a dual tree is used to segment the current encoding tree nodes, and if a dual tree is used, whether it is a luma component dual tree (DAUL_TREE_LUMA) or a chroma component dual tree (DAUL_TREE_CHROMA).

[0275] In this process, the luminance quantization parameter Qp'Y, the chromaticity quantization parameters Qp'Cb, Qp'Cr, and Qp'CbCr can be derived.

[0276] The variable brightness position (xQg, yQg) indicates the position of the top-left brightness sample of the current quantization group corresponding to the top-left sample of the current image. Here, the horizontal position xQg and the vertical position yQg can be set to be equal to the values ​​of the variables CuQgTopLeftX and CuQgTopLeftY, respectively. The variables CuQgTopLeftX and CuQgTopLeftY can be defined as follows: Figure 30 The predefined values ​​in the encoding tree syntax shown.

[0277] Here, the current quantization group can be a quadrilateral region within a coding tree block, and they can share the same qP. Y_PRED The value. Its width and height can be equal to the width and height of the coding tree node to which the top-left luminance sample position is assigned in CuQgTopLeftX and CuQgTopLeftY.

[0278] When treeType is SINGLE_TREE or DUAL_TREE_LUMA, the predicted value of the luminance quantization parameter qP is... Y_PRED The following steps can be used to derive the method.

[0279] 1. The variable qP can be derived as follows: Y_PRED .

[0280] (Condition 1) qP is true if any one of the following conditions is true. Y_PRED The value can be set to match SliceQp Y Same value (here, SliceQp) Y The quantization parameter Qp indicates all slices in the image. Y The initial value, and this can be obtained from the bitstream). Alternatively, qP Y_PRED The value can be set to the luminance quantization parameter Qp based on the last luminance code of the preceding quantization group, depending on the decoding order. Y The value of .

[0281] -(Condition 1-1) When the current quantization group is the first quantization group in the slice

[0282] -(Condition 1-2) When the current quantization group is the first quantization group in the tile

[0283] - (Conditions 1-3) When the current quantization group is the first quantization group in the CTB row of the tile and a scheduled synchronization occurs (e.g., when entropy_coding_sync_enabled_flag has a value of 1).

[0284] 2. The variable qP can be derived as follows: Y_A The value of .

[0285] (Condition 2) qP is true when at least one of the following conditions is true. Y_A The value can be set to qP Y_PRED The value of qP. Alternatively, qP Y_A The value can be set to the luminance quantization parameter Qp, which is the coding basis of the luminance coding block covering the luminance sample position (xQg-1, yQg). Y The value of .

[0286] - (Condition 2-1) For a block identified by sample location (xCb, yCb), the block identified by sample location (xQg-1, yQg) is not a usable neighboring block.

[0287] - (Condition 2-2) When the CTB of the luminance coding block including the luminance sample position (xQg-1, yQg) is different from the CTB of the current luminance coding block including the luminance sample position (xCb, yCb), for example, when all of the following conditions are true,

[0288] The value of -(condition 2-2-1)(xQg-1)>>CtbLog2SizeY is different from the value of (xCb)>>CtbLog2SizeY.

[0289] The value of (condition 2-2-2)(yQg)>>CtbLog2SizeY is different from the value of (yCb)>>CtbLog2SizeY.

[0290] 3. The variable qP can be derived as follows: Y_B The value of .

[0291] (Condition 3) qP is true when at least one of the following conditions is true. Y_B The value can be set to qP Y_PRED The value of qP. Alternatively, qP Y_B The value can be set to the luminance quantization parameter Qp, which is the coding basis of the luminance coding block covering the luminance sample position (xQg, yQg-1). Y The value of .

[0292] -(Condition 3-1) For a block identified by sample location (xCb, yCb), if the block identified by sample location (xQg, yQg-1) is not a usable neighboring block,

[0293] - (Condition 3-2) When the CTB of the luminance coding block including the luminance sample position (xQg, yQg-1) is different from the CTB of the current luminance coding block including the luminance sample position (xCb, yCb), for example, when all of the following conditions are true,

[0294] The value of (condition 3-2-1)(xQg)>>CtbLog2SizeY is different from the value of (xCb)>>CtbLog2SizeY.

[0295] The value of -(condition 3-2-2)(yQg-1)>>CtbLog2SizeY is different from the value of (yCb)>>CtbLog2SizeY.

[0296] 4. The predicted value of the brightness quantization parameter qP can be derived as follows: Y_PRED .

[0297] qP is true when all of the following conditions are true. Y_PREDThe luminance quantization parameter Qp can be set as the encoding basis for a luminance coding block that covers the luminance sample location (xQg, yQg-1). Y .

[0298] -(Condition 3-1) For a block identified by sample location (xCb, yCb), the block identified by sample location (xQg, yQg-1) is a usable neighboring block.

[0299] - When the current quantization group is the first quantization group in the CTB row of the tile.

[0300] Furthermore, when all conditions are false, qP can be derived as shown in the following formula. Y_PRED .

[0301] [Formula 3]

[0302] qP Y_PRED =(qP Y_A +qP Y_B +1)>>1

[0303] The variable Qp can be derived as shown in the following formula. Y .

[0304] [Formula 4]

[0305] Qp Y =((qP) Y_PRED +CuQpDeltaVal+64+2*QpBdOffset)%(64+QpBdOffset))-QpBdOffset

[0306] Here, CuQpDeltaVal indicates the difference between the luma quantization parameter used for encoding and its predicted value. Its value can be obtained from the bitstream. QpBdOffset indicates the range offset of the luma and chroma quantization parameters. QpBdOffset can be preset to a predetermined constant or obtained from the bitstream. For example, QpBdOffset can be calculated by multiplying a predetermined constant (e.g., 6) by the value of a syntax element (e.g., sps_bitdepth) indicating the bit depth of the luma or chroma sample. The luma quantization parameter Qp' can be derived as shown in the following equation. Y .

[0307] [Formula 5]

[0308] Qp' Y =Qp Y +QpBdOffset

[0309] When the value of the variable ChromaArrayType, which indicates the type of the chroma array, is not the first value (e.g., 0) and treeType is SINGLE_TREE or DUAL_TREE_CHROMA, the following processing can be performed.

[0310] - When the value of treeType is DUAL_TREE_CHROMA, the variable Qp Y The value can be set to the luminance quantization parameter Qp based on the luminance encoding at the luminance sample location (xCb+cbWidth / 2, yCb+cbHeight / 2). Y The value is the same as the value.

[0311] - The variable qP can be derived as shown in the following formula. Cb qP Cr and qP CbCr .

[0312] [Formula 6]

[0313] qP Chroma =Clip3(-QpBdOffset,63,Qp Y )

[0314] qP Cb =ChromaQpTable[0][qP Chroma ]

[0315] qP Cr =ChromaQpTable[1][qP Chroma ]

[0316] qP CbCr =ChromaQpTable[2][qP Chroma ]

[0317] The colorimetric parameters Qp' of the Cb and Cr components can be derived as shown in the following formula. Cb and Qp' Cr and the colorimetric parameter Qp' of the joint Cb-Cr encoding CbCr .

[0318] [Formula 7]

[0319] Qp' Cb =Clip3(-QpBdOffset,63,qP) Cb +pps_cb_qp_offset+slice_cb_qp_offset+CuQpOffset Cb )+QpBdOffset

[0320] Qp'Cr =Clip3(-QpBdOffset,63,qP) Cr +pps_cr_qp_offset+slice_cr_qp_offset+CuQpOffset Cr )+QpBdOffset

[0321] Qp' CbCr =Clip3(-QpBdOffset,63,qP) CbCr +pps_joint_cbcr_qp_offset+slice_joint_cbcr_qp_offset+CuQpOffset CbCr )+QpBdOffset

[0322] In the above formula, pps_cb_qp_offset and pps_cr_qp_offset are used to derive Qp'. Cb and Qp' Cr The offset can be obtained from the bitstream of the image parameter set. `slice_cb_qp_offset` and `slice_cr_qp_offset` are used to derive `Qp'`. Cb and Qp' Cr The offset can be obtained from the bitstream in the slice header. CuQpOffset Cb and CuQpOffset Cr It is used to derive Qp' C b and Qp' Cr The offset can be obtained from the bitstream of the transform base.

[0323] Alternatively, for example, the following parameters can be used to perform dequantization of the transform coefficients.

[0324] - Brightness coordinates (xTbY, yTbY), which refer to the relative coordinates of the top-left sample of the current brightness transformation block with respect to the top-left brightness sample of the current image.

[0325] - The variable nTbW indicates the width of the transform block.

[0326] - The variable nTbH indicates the height of the transform block.

[0327] - The variable cIdx indicates the color component of the current block.

[0328] The output of this process can be an array d of scaling transformation coefficients. Here, the size of array d can be (nTbW)×(nTbH). Each element constituting it can be labeled d[x][y].

[0329] Therefore, the quantization parameter qP can be derived as follows. When cIdx has a value of 0, qP can be derived as shown in the following formula.

[0330] [Formula 8]

[0331] qP=Qp' Y

[0332] Alternatively, when TuCResMode[xTbY][yTbY] has a value of 2, it can be derived as shown in the following formula.

[0333] [Formula 9]

[0334] qP=Qp' C b C r

[0335] Alternatively, when cIdx has a value of 1, qP can be derived as shown in the following formula.

[0336] [Formula 10]

[0337] qP=Qp' Cb

[0338] Alternatively, when cIdx has a value of 2, qP can be derived as shown in the following formula.

[0339] [Equation 11]

[0340] qP=Qp' Cr

[0341] Subsequently, the quantization parameter qP can be updated as follows. Additionally, the variables rectNonTsFlag and bdShift can be derived as follows. For example, when transform_skip_flag[xTbY][yTbY][cIdx] has a value of 0 (e.g., when the transformation of the current transform block is not skipped), it can be derived as shown in the following equation. Here, the first value of transform_skip_flag (e.g., 0) can indicate whether the transformation is omitted, determined by another syntax element. The second value of transform_skip_flag (e.g., 1) can indicate that the transformation is omitted (e.g., skipped).

[0342] [Equation 12]

[0343] qP=qP-(cu_act_enabled_flag[xTbY][yTbY]?5:0)

[0344] rectNonTsFlag = 0

[0345] bdShift = 10

[0346] Alternatively, when transform_skip_flag[xTbY][yTbY][cIdx] has a value of 1 (e.g., skipping the transformation of the current transform block), it can be derived as shown in the following formula.

[0347] [Equation 13]

[0348] qP=Max(QpPrimeTsMin,qP)-(cu_act_enabled_flag[xTbY][yTbY]?5:0)

[0349] rectNonTsFlag=(((Log2(nTbW)+Log2(nTbH))&1)==1

[0350] bdShift=BitDepth+(rectNonTsFlag?1:0)+((Log2(nTbW)+Log2(nTbH)) / 2)-+pic_dep_quant_enabled_flag

[0351] Here, QpPrimeTsMin can indicate the minimum quantization parameter value allowed when the transform skip mode is applied. This can be determined as a predetermined constant, or it can be derived from the syntax elements of the bitstream associated with it.

[0352] Here, the suffixes Y, Cb, and Cr can represent the G, B, and R color components in the RGB color model or the Y, Cg, and Co color components in the YCgCo color model.

[0353] Overview of Block Differential Pulse Code Modulation (BDPCM)

[0354] The image encoding apparatus and image decoding apparatus according to the embodiments can perform differential encoding of the residual signal. For example, the image encoding apparatus can encode the residual signal by subtracting the prediction signal from the residual signal of the current block, and the image decoding apparatus can decode the residual signal by adding the residual signal of the current block to the prediction signal. The image encoding apparatus and image decoding apparatus according to the embodiments can perform differential encoding of the residual signal by applying the following BDPCM.

[0355] The BDPCM according to this disclosure can be performed in the quantization residual domain. The quantization residual domain may include the quantization residual signal (or quantization residual coefficients), and when BDPCM is applied, the transformation of the quantization residual signal can be skipped. For example, when BDPCM is applied, the transformation of the residual signal can be skipped, and quantization can be performed. Alternatively, the quantization residual domain may include quantization transform coefficients.

[0356] In implementations using BDPCM, the image coding device can derive the residual block of the current block predicted in the intra-frame prediction mode and quantize the residual block, thereby deriving the residual block. When performing differential coding mode on the residual signal for the current block, the image coding device can perform differential coding on the residual block to derive the modified residual block. Additionally, the image coding device can encode differential coding mode information for a specified residual signal and the modified residual block, thereby generating a bitstream.

[0357] More specifically, when BDPCM is applied to the current block, a prediction block (prediction block) including prediction samples of the current block can be generated through intra-frame prediction. In this case, the intra-frame prediction mode used to perform intra-frame prediction can be notified via a bitstream signal and can be derived based on the prediction direction of BDPCM described below. Furthermore, in this case, the intra-frame prediction mode can be determined as either a vertical prediction direction mode or a horizontal prediction direction mode. For example, when the prediction direction of BDPCM is horizontal, the intra-frame prediction mode can be determined as a horizontal prediction direction mode, and the prediction block of the current block can be generated through horizontal intra-frame prediction. Alternatively, when the prediction direction of BDPCM is vertical, the intra-frame prediction mode can be determined as a vertical prediction direction mode, and the prediction block of the current block can be generated through vertical intra-frame prediction. When horizontal intra-frame prediction is applied, the values ​​of the pixels adjacent to the left of the current block can be determined as the prediction sample values ​​of the samples included in the corresponding row of the current block. When applying vertical intra-frame prediction, the values ​​of pixels adjacent to the top of the current block can be determined as predicted sample values ​​for the samples included in the corresponding row of the current block. When applying BDPCM to the current block, the method for generating the predicted block for the current block can be performed equally in both the image encoding and image decoding devices.

[0358] When applying BDPCM to the current block, the image coding device can generate a residual block that includes residual samples of the current block by subtracting prediction samples from the current block. The image coding device can quantize the residual block and then encode the difference (or increment) between the quantized residual samples and the predictors of the quantized residual samples. The image decoding device can generate a quantized residual block of the current block by obtaining the quantized residual samples of the current block based on the predictors and the difference reconstructed from the bitstream. Subsequently, the image decoding device can dequantize the quantized residual block and add it to the prediction block, thereby reconstructing the current block.

[0359] Figure 31 This is a view illustrating the encoding method of residual samples of BDPCM according to this disclosure. Figure 31 The residual block can be generated by subtracting the prediction block from the current block in the image coding device. Figure 31The quantized residual block can be generated by quantizing the residual block using an image coding device. Figure 31 In the middle, r i,j Specifies the value of the residual sample at coordinate (i,j) in the current block. When the size of the current block is M×N, the value i can be from 0 to M-1, including the extreme values. Similarly, the value j can be from 0 to N-1, including the extreme values. For example, the residual can refer to the difference between the original block and the predicted block. For example, r i,j This can be derived by subtracting the predicted sample value from the original sample value at coordinate (i,j) in the current block. For example, r i,j This can be the prediction residual after performing horizontal or vertical intra-frame prediction using unfiltered samples from the top or left boundary. In horizontal intra-frame prediction, the values ​​of the left neighboring pixels are copied along the line passing through the prediction block. In vertical intra-frame prediction, the top neighboring row is copied to each row of the prediction block.

[0360] exist Figure 31 In, Q(r) i,j ) refers to the value of the quantized residual sample at coordinate (i,j) in the current block. For example, Q(r i,j ) can refer to r i,j The quantized value.

[0361] right Figure 31 The quantized residual samples are used to perform BDPCM predictions, and a modified quantized residual block R' of size M×N can be generated, which includes the modified quantized residual samples r'.

[0362] When the prediction direction of BDPCM is horizontal, the modified quantized residual sample value r' of the coordinates (i,j) in the current block can be calculated as shown in the following formula. i,j .

[0363] [Formula 14]

[0364]

[0365] As shown in Equation 14, when the prediction direction of BDPCM is horizontal, the value of the quantized residual sample Q(r) is... 0,j The value r' assigned as is to the coordinate (0,j) 0,j The values ​​of other coordinates (i,j) r' i,j The value of the quantized residual sample Q(r) can be derived as the coordinate (i,j). i,j The value of the quantized residual sample Q(r) with coordinates (i-1,j) i-1,j The difference between (i,j) and (j). That is, the value Q(r) as the quantized residual sample of coordinates (i,j). i,j The encoding is replaced by using the value Q(r) of the quantized residual sample at coordinate (i-1,j).i-1,j The difference calculated as the predicted value is derived into the modified quantized residual sample value r'i,j, and then the value r'i,j is encoded.

[0366] When the prediction direction of BDPCM is vertical, the modified quantized residual sample value (r') of the coordinates (i,j) in the current block can be calculated as shown in the following formula. i,j ).

[0367] [Formula 15]

[0368]

[0369] As shown in Equation 15, when the prediction direction of BDPCM is vertical, the value of the quantized residual sample Q(r) is... i,0 The value r' assigned to coordinate (i,0) as is i,0 The values ​​of other coordinates (i,j) r' i,j The value of the quantized residual sample Q(r) can be derived as the coordinate (i,j). i,j The value of the quantized residual sample Q(r) with coordinates (i,j-1) i,j-1 The difference between (i,j) and (j). That is, the value Q(r) as the quantized residual sample of coordinates (i,j). i,j The encoding is replaced by using the value Q(r) of the quantized residual sample at coordinate (i,j-1). i,j-1 The difference calculated as the predicted value is derived into the modified quantized residual sample value r'i,j, and then the value r'i,j is encoded.

[0370] As mentioned above, the process of modifying the current quantized residual sample value by using nearby quantized residual sample values ​​as predicted values ​​can be called BDPCM prediction.

[0371] Finally, the image encoding device can encode the modified quantization residual block, which includes the modified quantization residual samples, and can send the resulting block to the image decoding device. Here, as described above, the transformation of the modified quantization residual block is not performed.

[0372] Figure 32 This is a view showing the modified quantization residual block generated by performing BDPCM according to this disclosure.

[0373] exist Figure 32 In the diagram, the horizontal BDPCM shows the modified quantization residual block generated according to Equation 14 when the prediction direction of the BDPCM is horizontal. Similarly, the vertical BDPCM shows the modified quantization residual block generated according to Equation 15 when the prediction direction of the BDPCM is vertical.

[0374] Figure 33 This is a flowchart illustrating the process of encoding the current block by applying BDPCM in an image encoding device.

[0375] First, when the current block, which is the target block for encoding, is input in step S3310, prediction can be performed on the current block in step S3320 to generate a prediction block. The prediction block in step S3320 can be an intra-frame prediction block, and the intra-frame prediction mode can be determined as described above. Based on the prediction block generated in step S3320, a residual block of the current block can be generated in step S3330. For example, the image coding device can generate a residual block (the value of the residual sample) by subtracting the prediction block (the value of the predicted sample) from the current block (the value of the original sample). For example, by executing step S3330, a prediction block can be generated. Figure 31 The residual block generated in step S3330. Quantization can be performed on the residual block generated in step S3340 to generate a quantized residual block, and BDPCM prediction can be performed on the quantized residual block in step S3350. The quantized residual block generated as a result of performing step S3340 can be... Figure 31 The quantized residual block. As the result of the BDPCM prediction in step S3350, it can be generated according to the prediction direction. Figure 32 The modified quantization residual block. Since it has already referenced... Figures 31 to 32 The BDPCM prediction in step S3350 has been described, therefore a detailed description of it will be omitted. Subsequently, the image encoding device can encode the modified quantization residual block to generate a bitstream in step S3360. Here, the transformation of the modified quantization residual block can be skipped.

[0376] refer to Figures 31 to 33 The BDPCM operation in the described image encoding device can be reversed by the image decoding device.

[0377] Figure 34 This is a flowchart illustrating the process of reconstructing the current block by applying BDPCM in an image decoding device.

[0378] In step S3410, the image decoding device can obtain the information (image information) required to reconstruct the current block from the bitstream. The information required to reconstruct the current block may include prediction information about the current block (prediction information) and residual information about the current block (residual information). The image decoding device can perform prediction on the current block based on the information about the current block, and can generate a prediction block in step S3420. The prediction about the current block can be intra-frame prediction, and its detailed description is consistent with the above reference. Figure 33 The detailed descriptions are identical. Figure 34The diagram illustrates step S3420, which generates the prediction block of the current block, prior to steps S3430 to S3450, which generate the residual block of the current block. However, no limitations are imposed on this. The prediction block of the current block can be generated after the residual block of the current block is generated. Alternatively, both the residual block and the prediction block of the current block can be generated simultaneously.

[0379] In step S3430, the image decoding device can generate a residual block for the current block by parsing the residual information of the current block from the bitstream. The residual block generated in step S3430 can be... Figure 32 The modified quantization residual block is shown in the image.

[0380] The image decoding device can perform the following steps in step S3440: Figure 32 The modified quantized residual block is used to perform BDPCM prediction to generate Figure 31 The quantized residual block. The BDPCM prediction in step S3440 is used to obtain the quantized residual block. Figure 32 Modified quantization residual block generation Figure 31 The process of quantizing the residual block corresponds to the inverse processing of step S3350 performed by the image encoding device. For example, when the differential coding mode information (e.g., bdpcm_flag) obtained from the bitstream indicates a differential coding mode for performing differential coding of residual coefficients when applying BDPCM, the image decoding device performs differential coding on the residual block to derive the modified residual block. Using the residual coefficients to be modified and the predicted residual coefficients, the image decoding device can modify at least one of the residual coefficients in the residual block. The predicted residual coefficients can be determined based on the prediction direction indicated by the differential coding direction information (e.g., bdpcm_dir_flag) obtained from the bitstream. The differential coding direction information can indicate a vertical or horizontal direction. The image decoding device can assign the value obtained by adding the residual coefficient to be modified to the position of the residual coefficient to be modified. Here, the predicted residual coefficient can be a coefficient that immediately precedes and is adjacent to the residual coefficient to be modified according to the order of the prediction direction.

[0381] The BDPCM prediction in step S3440 performed by the image decoding device will be described in more detail below. The decoding device can calculate the quantized residual sample Q(r) by reversing the calculation performed by the encoding device. i,j For example, when the prediction direction of BDPCM is horizontal, the image decoding device can generate a quantization residual block from the modified quantization residual block using Equation 16.

[0382] [Formula 16]

[0383]

[0384] As defined in Equation 16, the value of the quantized residual sample at coordinate (i,j) can be calculated by summing the values ​​of the modified quantized residual samples from coordinate (0,j) to coordinate (i,j). i,j ).

[0385] Alternatively, Equation 17 can be used instead of Equation 16 to calculate the value Q(r) of the quantized residual sample at coordinate (i,j). i,j ).

[0386] [Equation 17]

[0387]

[0388] Equation 17 is the inverse of Equation 14. According to Equation 17, the value of the quantized residual sample at coordinate (0,j) is Q(r). 0,j The value r' is derived as the modified quantized residual sample value at coordinates (0,j). 0,j Q(r) for other coordinates (i,j) i,j The value r' is derived as the modified quantized residual sample value at coordinates (i,j). i,j The value of the quantized residual sample Q(r) with coordinates (i-1,j) i-1,j The sum of ) . That is, by using the value Q(r) of the quantized residual sample at coordinates (i-1,j). i-1,j The difference r'i,j is summed using the predicted value as the summation, thereby deriving the quantized residual sample value Q(r). 1,j ).

[0389] When the prediction direction of BDPCM is vertical, the image decoding device can generate a quantization residual block from the modified quantization residual block using Equation 18.

[0390] [Formula 18]

[0391]

[0392] As defined in Equation 18, the value of the quantized residual sample at coordinate (i,j) can be calculated by summing the values ​​of the modified quantized residual samples from coordinate (i,j) to coordinate (i,j). i,j ).

[0393] Alternatively, Equation 19 can be used instead of Equation 18 to calculate the value Q(r) of the quantized residual sample at coordinate (i,j). i,j ).

[0394] [Formula 19]

[0395]

[0396] Equation 19 is the inverse of Equation 15. According to Equation 19, the value of the quantized residual sample at coordinate (i,0) is Q(r). i,0 The value r' is derived as the value of the modified quantized residual sample at coordinate (i,0). i,0 Q(r) for other coordinates (i,j) i,j The value r' is derived as the modified quantized residual sample value at coordinates (i,j). i,j The value of the quantized residual sample Q(r) with coordinates (i,j-1) i,j-1 The sum of ) . That is, by using the value Q(r) of the quantized residual sample at coordinates (i,j-1). i,j-1 ) as the predicted value for the difference r' i,j Summing, we can derive the quantized residual sample value Q(r) from this summation. i,j ).

[0397] When a quantization residual block consisting of quantization residual samples is generated by executing step S3440 according to the above method, the image decoding device performs inverse quantization on the quantization residual block in step S3450 to generate the residual block of the current block. When BDPCM is applied, the transformation of the current block is skipped as described above. Therefore, the inverse transformation of the inverse quantization residual block can be skipped.

[0398] Subsequently, in step S3460, the image decoding device can reconstruct the current block based on the prediction block generated in step S3420 and the residual block generated in step S3450. For example, the image decoding device can reconstruct the current block (the value of the reconstructed sample) by adding the prediction block (the value of the predicted sample) to the residual block (the value of the residual sample). For example, this can be achieved by adding the dequantized sample Q... -1 (Q(r i,j The reconstructed sample value is generated by adding the predicted value within the block to the reconstructed sample value. The differential coding mode information indicating whether BDPCM is applied to the current block can be signaled via a bitstream. Additionally, when BDPCM is applied to the current block, the differential coding direction information indicating the prediction direction of BDPCM can be signaled via a bitstream. When BDPCM is not applied to the current block, the differential coding direction information may not be signaled.

[0399] Figures 35 to 37 This is a schematic view illustrating the syntax used to signal information about BDPCM.

[0400] Figure 35This is a view illustrating the syntax of the sequence parameter set according to an implementation for signaling BDPCM information. In this implementation, all SPS RBSPs included in at least one access unit (AU) having a value of 0 as a time ID (TemporalId), or provided externally, can be set to be used before being referenced in the decoding process. Additionally, SPS NAL units including SPS RBSPs can be set to have the same nuh_layer_id as the PPS NAL unit referencing the SPS NAL unit. In CVS, all SPS NAL units with a specific sps_seq_parameter_set_id value can be set to have the same content. Figure 35 The seq_parameter_set_rbsp() syntax exposes the aforementioned sps_transform_skip_enable_flag and the sps_bdpcm_enabled_flag, which will be discussed later.

[0401] The syntax element `sps_bdpcm_enabled_flag` indicates whether `intra_bdpcm_flag` is provided in the CU syntax for intra coding units. For example, a first value of `sps_bdpcm_enabled_flag` (e.g., 0) indicates that `intra_bdpcm_flag` is not provided in the CU syntax for intra coding units. A second value of `sps_bdpcm_enabled_flag` (e.g., 1) indicates that `intra_bdpcm_flag` is provided in the CU syntax for intra coding units. Furthermore, when `sps_bdpcm_enabled_flag` is not provided, its value can be set to the first value (e.g., 0).

[0402] Figure 36 This is a view illustrating an implementation of the syntax for signaling whether a constraint on BDPCM is applied. In this implementation, predetermined constraints in the encoding / decoding process can be signaled using the `general_constraint_info()` syntax. Figure 36The syntax allows signaling to the `no_bdpcm_constraint_flag` element, which indicates whether the value of `sps_bdpcm_enabled_flag` should be set to 0. For example, a first value for `no_bdpcm_constraint_flag` (e.g., 0) can indicate that no constraint is applied. When the value of `no_bdpcm_constraint_flag` is a second value (e.g., 1), the value of `sps_bdpcm_enabled_flag` can be forced to the first value (e.g., 0).

[0403] Figure 37 This is a view illustrating an implementation of the coding_unit() syntax for signaling information about BDPCM to the coding unit. (As shown) Figure 37 As shown, the syntax elements intra_bdpcm_flag and intra_bdpcm_dir_flag can be signaled using the coding_unit() syntax. The syntax element intra_bdpcm_flag can indicate whether BDPCM is applied to the current luminance coding block located at (x0, y0).

[0404] For example, a first value of `intra_bdpcm_flag` (e.g., 0) can indicate that BDPCM is not applied to the current luma coding block. A second value of `intra_bdpcm_flag` (e.g., 1) can indicate that BDPCM is applied to the current luma coding block. By indicating the application of BDPCM, `intra_bdpcm_flag` can indicate whether to skip the transform and whether to perform intra-frame luma prediction mode via `intra_bdpcm_dir_flag`, which will be described later.

[0405] Furthermore, for x = x0..x0+cbWidth-1 and y = y0..y0+cbHeight-1, the value of the above variable BdpcmFlag[x][y] can be set to the value of intra_bdpcm_flag.

[0406] The syntax element `intra_bdpcm_dir_flag` can indicate the prediction direction of BDPCM. For example, a first value of `intra_bdpcm_dir_flag` (e.g., 0) can indicate that the BDPCM prediction direction is horizontal. A second value of `intra_bdpcm_dir_flag` (e.g., 1) can indicate that the BDPCM prediction direction is vertical.

[0407] Furthermore, for x = x0..x0 + cbWidth-1 and y = y0 + cbHeight-1, the value of the variable BdpcmDir[x][y] can be set to the value of intra_bdpcm_dir_flag.

[0408] Intra-frame prediction of chroma blocks

[0409] When performing intra-frame prediction on the current block, prediction of both the luma component block (luminance block) and the chroma component block (chroma block) of the current block can be performed. In this case, the intra-frame prediction mode for the chroma block can be set separately from the intra-frame prediction mode for the luma block.

[0410] For example, the intra-chroma prediction mode of a chroma block can be indicated based on intra-chroma prediction mode information, and this information can be signaled using the `intra_chroma_pred_mode` syntax element. For instance, the intra-chroma prediction mode information can represent one of the following modes: planar mode, DC mode, vertical mode, horizontal mode, derivation mode (DM), and cross-component linear model (CCLM). Here, planar mode can specify intra-prediction mode #0, DC mode can specify intra-prediction mode #1, vertical mode can specify intra-prediction mode #26, and horizontal mode can specify intra-prediction mode #10. DM can also be called direct mode. CCLM can also be called linear model (LM). CLM modes can include any of the following: L_CCLM, T_CCLM, and LT_CCLM.

[0411] Furthermore, DM and CCLM are dependent intra-prediction modes used to predict chroma blocks using information from luma blocks. DM can refer to an intra-prediction mode that is the same as the intra-prediction mode for the luma component, but is applied as the intra-prediction mode for the chroma component. CCLM can refer to an intra-prediction mode in which, in the process of generating the prediction block for the chroma block, reconstructed samples of the luma block are subsampled, and samples derived by applying CCLM parameters α and β to the subsampled samples are used as prediction samples for the chroma block.

[0412] Overview of Cross-Component Linear Model (CCLM) Modes

[0413] As described above, the CCLM mode can be applied to chroma blocks. The CCLM mode is an intra-frame prediction mode that uses the correlation between luma blocks and their corresponding chroma blocks, and is performed by deriving a linear model based on neighboring samples of the luma blocks and the chroma blocks. Alternatively, predicted samples of the chroma blocks can be derived based on reconstructed samples of the luma blocks and the derived linear model.

[0414] Specifically, when applying the CCLM mode to the current chroma block, the parameters of the linear model can be derived based on the neighbor samples of the intra-prediction used for the current chroma block and the neighbor samples of the intra-prediction used for the current luma block. For example, the linear model of CCLM can be represented based on the following formula.

[0415] [Formula 20]

[0416] pred c (i, j) = α·rec L ′(i,j)+β

[0417] Here, pred c (i,j) can refer to the predicted sample of the coordinates (i,j) of the current chroma block in the current CU. L '(i,j) can refer to the reconstructed sample of the coordinates (i,j) of the current luma block in the CU. For example, rec L '(i,j) can refer to the reconstructed sample of the current luma block after downsampling. The linear model coefficients α and β can be signaled or derived from neighboring samples.

[0418] Joint encoding of residuals (joint CbCr)

[0419] In the encoding / decoding process according to the implementation, chroma residuals can be encoded / decoded together. This can be referred to as joint encoding of residuals or joint CbCr. Whether the joint encoding mode of CbCr is applied (enabled) can be signaled by the joint encoding mode signaling flag tu_joint_cbcr_residual_flag, which is signaled at the transform basis level. Alternatively, the selected encoding mode can be derived from the chroma CBF. The flag tu_joint_cbcr_residual_flag can be present when the value of at least one chroma CBF used for the transform basis is 1. The chroma QP offset value indicates the difference between the general chroma QP offset value signaled for the regular chroma residual encoding mode and the chroma QP offset value for the CbCr joint encoding mode. The chroma QP offset value can be signaled via PPS or slice hair signal. This QP offset value can be used to drive the chroma QP value of the block using the joint chroma residual encoding mode.

[0420] When Mode 2, which is the corresponding joint chroma coding mode in the table below, is enabled for the transform basis, its chroma QP offset can be added to the target luminance-derived chroma QP (the applied luminance-derived chroma QP) during the quantization and decoding of the transform basis.

[0421] For the other modes in the table below, such as modes 1 and 3, the chromaticity QP can be derived in the same way as for general Cb or Cr blocks. The processing of reconstructing the chromaticity residuals (resCb and resCr) from the transform blocks can be selected according to the table below. When this mode is enabled, a single joint chromaticity residual block (resJointC[x][y] in the table below) is signaled, and information such as tu_cbf_cb, tu_cbf_cr, and CSign, which is the sign value exposed in the slice header, can be taken into account to derive the residual block resCb of Cb and the residual block resCr of Cr.

[0422] In the encoding device, the joint chroma components can be derived as follows. Based on the joint encoding mode, resJointC{1,2} can be generated in the following order. In mode 2 (with a single residual reconstructing Cb = C, Cr = CSign * C), the joint residual can be determined according to the following formula.

[0423] [Equation 21]

[0424] resJointC[x][y]=(resCb[x][y]+CSign*resCr[x][y]) / 2.

[0425] Alternatively, in the case of Mode 1 (with a single residual reconstructed as Cb = C, Cr = (CSign * C) / 2), the joint residual can be determined according to the following formula.

[0426] [Equation 22]

[0427] resJointC[x][y]=(4*resCb[x][y]+2*CSign*resCr[x][y]) / 5.

[0428] Alternatively, in Mode 3 (with a single residual of reconstructed Cr = C, Cb = (CSign*C) / 2), the joint residual can be determined according to the following formula.

[0429] [Equation 23]

[0430] resJointC[x][y]=(4*resCr[x][y]+2*CSign*resCb[x][y]) / 5.

[0431] [Table 2]

[0432]

[0433] The above shows the reconstruction of the chroma residual. CSign refers to the sign value +1 or -1 specified in the slice header. resJointC[][] refers to the transmitted residual. In the table, mode refers to TuCResMode, which will be described later. The three joint chroma encoding modes in the table can be supported only for I slices. For P and B slices, only mode 2 can be supported. Therefore, for P and B slices, the syntax element tu_joint_cbcr_residual_flag can only be provided if the chroma cbf values ​​(e.g., tu_cbf_cb and tu_cbf_cr) are both 1. Furthermore, the transform depth can be removed in the context modeling of tu_cbf_luma and tu_cbf_cb.

[0434] Implementation Method 1: QP Update Method Using ACT Qp_offset

[0435] As described above, QP can be updated to apply ACT. However, this QP update has several problems. For example, when using the above method, it is impossible to set different ACT Qp offsets for each color component. Additionally, the derived qP value may be negative. Therefore, in the following implementation, a method for applying clipping to Qp values ​​derived from ACT QP offset values ​​based on color component values ​​is described.

[0436] In the implementation, the quantization parameter qP can be derived as follows.

[0437] First, when cIdx has a value of 0, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0438] [Equation 24]

[0439] qP=Qp' Y

[0440] ActQpOffset=5

[0441] Alternatively, when TuCResMode[xTbY][yTbY] has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0442] [Equation 25]

[0443] qP=Qp' C b C r

[0444] ActQpOffset=5

[0445] Alternatively, when cIdx has a value of 1, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0446] [Equation 26]

[0447] qP=Qp' Cb

[0448] ActQpOffset=3

[0449] The quantization parameter qP can be updated as follows.

[0450] When transform_skip_flag[xTbY][yTbY][cIdx] has a value of 0, qP can be derived as shown in the following formula.

[0451] [Equation 27]

[0452] qP=Max(0,qP-(cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0))

[0453] Alternatively, when transform_skip_flag[xTbY][yTbY][cIdx] has a value of 1, qP can be derived as shown in the following formula.

[0454] [Equation 28]

[0455] qP=Max(0,Max(QpPrimeTsMin,qP)-(cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0))

[0456] In another implementation, when transform_skip_flag[xTbY][yTbY][cIdx] has a value of 1, the value of QpPrimeTsMin can be used instead of 0 to limit qP as shown in the following formula.

[0457] [Equation 29]

[0458] qP=Max(QpPrimeTsMin,qP-(cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0)

[0459] In another embodiment, the quantization parameter qP can be derived as follows.

[0460] First, when cIdx has a value of 0, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0461] [Formula 30]

[0462] qP=Qp' Y

[0463] ActQpOffset=5

[0464] Alternatively, when TuCResMode[xTbY][yTbY] has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0465] [Equation 31]

[0466] qP=Qp' C b C r

[0467] ActQpOffset=5

[0468] Alternatively, when cIdx has a value of 1, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0469] [Equation 32]

[0470] qP=Qp' Cb

[0471] ActQpOffset=5

[0472] Alternatively, when cIdx has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following equation.

[0473] [Equation 33]

[0474] qP=Qp' Cr

[0475] ActQpOffset=3

[0476] The quantization parameter qP can be updated as follows.

[0477] When transform_skip_flag[xTbY][yTbY][cIdx] has a value of 0, qP can be derived as shown in the following formula.

[0478] [Formula 34]

[0479] qP=Max(0,qP-(cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0))

[0480] Alternatively, when transform_skip_flag[xTbY][yTbY][cIdx] has a value of 1, qP can be derived as shown in the following formula.

[0481] [Formula 35]

[0482] qP=Max(0,Max(QpPrimeTsMin,qP)-(cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0))

[0483] In another implementation, when transform_skip_flag[xTbY][yTbY][cIdx] has a value of 1, the value of QpPrimeTsMin can be used instead of 0 to limit qP as shown in the following formula.

[0484] [Formula 36]

[0485] qP=Max(QpPrimeTsMin,qP-(cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0))

[0486] In another embodiment, the quantization parameter qP can be derived as follows.

[0487] First, when cIdx has a value of 0, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0488] [Formula 37]

[0489] qP=Qp' Y

[0490] ActQpOffset = -5

[0491] Alternatively, when TuCResMode[xTbY][yTbY] has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0492] [Formula 38]

[0493] qP=Qp' CbCr

[0494] ActQpOffset = -5

[0495] Alternatively, when cIdx has a value of 1, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0496] [Formula 39]

[0497] qP=Qp' C b

[0498] ActQpOffset = -5

[0499] Alternatively, when cIdx has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following equation.

[0500] [Formula 40]

[0501] qP=Qp' Cr

[0502] ActQpOffset = -3

[0503] The quantization parameter qP can be updated as follows.

[0504] When transform_skip_flag[xTbY][yTbY][cIdx] has a value of 0, qP can be derived as shown in the following formula.

[0505] [Formula 41]

[0506] qP=Max(0,qP+(cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0))

[0507] Alternatively, when transform_skip_flag[xTbY][yTbY][cIdx] has a value of 1, qP can be derived as shown in the following formula.

[0508] [Equation 42]

[0509] qP=Max(0,Max(QpPrimeTsMin,qP)+(cu_act_enabled_flag[xTbY][yTbY]?

[0510] ActQpOffset:0))

[0511] In another implementation, when transform_skip_flag[xTbY][yTbY][cIdx] has a value of 1, the value of QpPrimeTsMin can be used instead of 0 to limit qP as shown in the following formula.

[0512] [Formula 43]

[0513] qP=Max(QpPrimeTsMin,qP+(cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0))

[0514] In the above description, Y, Cb, and Cr can represent three color components. For example, in the ACT transform, Y can correspond to C0. Cb can correspond to C1 or Cg. Additionally, Cr can correspond to C2 or Co.

[0515] Additionally, the ACTQpOffset values ​​of -5, -5, and -3 for the three color components can be replaced with other values ​​or other variables.

[0516] Implementation Method 2: Signaling for QP Offset Adjustment in ACT

[0517] In the above embodiments, the ACT QP offset adjustments are fixed at -5, -5, and -3 for the Y, Cg, and Co components, respectively. In this embodiment, to provide greater flexibility in adjusting the ACT QP offset, a method for signaling the ACT QP offset will be described. The ACT QP offset can be signaled as a parameter in the PPS.

[0518] In the implementation method, it can be based on Figure 38 The syntax table signals qp_offset. Its syntax elements are as follows.

[0519] The syntax element pps_act_qp_offsets_present_flag can indicate whether syntax elements related to ACTQP offsets exist in the PPS. For example, pps_act_qp_offsets_present_flag can indicate whether the syntax elements pps_act_y_qp_offset, pps_act_cb_qp_offset, and pps_act_cr_qp_offset, which will be described later, are signaled as PPS.

[0520] For example, the first value of pps_act_qp_offsets_present_flag (e.g., 0) can indicate that pps_act_y_qp_offset, pps_act_cb_qp_offset, and pps_act_cr_qp_offset do not signal via the PPS syntax table.

[0521] The second value of pps_act_qp_offsets_present_flag (e.g., 1) can instruct pps_act_y_qp_offset, pps_act_cb_qp_offset, and pps_act_cr_qp_offset to signal via the PPS syntax table.

[0522] When pps_act_qp_offsets_present_flag is not provided from the bitstream, pps_act_qp_offsets_present_flag can be deduced to have a first value (e.g., 0). For example, when a flag indicating whether an ACT is applied (e.g., sps_act_enabled_flag, which signals in SPS) has a first value (e.g., 0) indicating that an ACT is not applied, pps_act_qp_offsets_present_flag can be forced to have a first value (e.g., 0).

[0523] When the value of the syntax element `cu_act_enabled_flag` is a second value indicating that ACT will be applied to the current encoding base (e.g., 1), the syntax elements `pps_act_y_qp_offset_plus5`, `pps_act_cb_qp_offset_plus5s`, and `pps_act_cr_qp_offset_plus3` can be used to determine the offsets applied to the quantization parameter values ​​`qP` for the luminance, Cb, and Cr components, respectively. When values ​​for `pps_act_y_qp_offset_plus5`, `pps_act_cb_qp_offset_plus5`, and `pps_act_cr_qp_offset_plus3` are not present in the bitstream, each value can be set to 0.

[0524] Based on the syntax elements, the value of the variable PpsActQpOffsetY can be determined as pps_act_y_qp_offset_plus5-5. The value of the variable PpsActQpOffsetCb can be determined as pps_act_cb_qp_offset_plus5-5. Additionally, the value of the variable PpsActQpOffsetCr can be determined as pps_act_cb_qp_offset_plus3-3.

[0525] Here, ACT is not an orthogonal transformation, so 5, 5, and 3 can be applied as constant offset values ​​to be subtracted. In the implementation, for bitstream matching, the values ​​of PpsActQpOffsetY, PpsActQpOffsetCb, and PpsActQpOffsetCr can have values ​​ranging from -12 to 12. Additionally, according to the implementation, the Qp offset values, besides 5, 5, and 3, can be replaced with other constant values ​​and used.

[0526] In another implementation, a more flexible ACT_QP offset can be used to adjust the QP. An example where the ACT QP offset is signaled in the bitstream is described in the following implementation. Therefore, the ACT QP offset can have a wider offset range. Consequently, the QP updated using the ACT QP offset is more likely to exceed the available range, thus necessitating the application of upper and lower limits to the updated QP (more detailed implementations will be described later in implementations 6 and 7).

[0527] The variables PpsActQpOffsetY, PpsActQpOffsetCb, PpsActQpOffsetCr, and PpsActQpOffsetCbCr, which indicate the ACT QP offset, can be values ​​derived using the ACT QP offset signaled via the bitstream or preset constants. For bitstream consistency, PpsActQpOffsetY, PpsActQpOffsetCb, PpsActQpOffsetCr, and PpsActQpOffsetCbCr can have values ​​ranging from -12 to +12.

[0528] When signaling the value of the QP offset without using a fixed value and the value has a range from -12 to 12, it is necessary to limit the upper limit of the derived QP value in addition to limiting the lower limit of the derived QP value to avoid QP having a negative value.

[0529] To prevent qP from having negative values, the minimum value of qP can be forced to 0. Alternatively, the minimum value of qP can be set to a value determined by the signaling syntax element. For example, to signal the minimum value of qP when applying a transform skip mode, the syntax element QpPrimeTsMin, which indicates the value of qP applied when applying the transform skip mode, can be used. The maximum value of qP can be limited to the maximum available qP value determined by the signaling syntax element or the maximum available value of qP (e.g., 63).

[0530] Based on the above implementation, the quantization parameter qP can be derived as follows. First, when cIdx has a value of 0, qP and the ACT Qp offset can be derived as shown in the following formula.

[0531] [Formula 44]

[0532] qP=Qp' Y

[0533] ActQpOffset=PpsActQpOffsetY

[0534] Alternatively, when TuCResMode[xTbY][yTbY] has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0535] [Formula 45]

[0536] qP=Qp' CbCr

[0537] ActQpOffset=PpsActQpOffsetCbCr

[0538] Alternatively, when cIdx has a value of 1, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0539] [Formula 46]

[0540] qP=Qp' C b

[0541] ActQpOffset=PpsActQpOffsetCb

[0542] Alternatively, when cIdx has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following equation.

[0543] [Formula 47]

[0544] qP=Qp' Cr

[0545] ActQpOffset=PpsActQpOffsetCr

[0546] In the implementation, the quantization parameter qP can be updated as follows. When transform_skip_flag[xTbY][yTbY][cIdx] has a value of 0, qP can be derived as shown in the following formula.

[0547] [Formula 48]

[0548] qP=Clip3(0,63,qP-(cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0))

[0549] Alternatively, when transform_skip_flag[xTbY][yTbY][cIdx] has a value of 1, qP can be derived as shown in the following formula.

[0550] [Formula 49]

[0551] qP=Clip3(0,63,Max(QpPrimeTsMin,qP)-(cu_act_enabled_flag[xTbY][yTbY]?

[0552] ActQpOffset:0)

[0553] In another implementation, when transform_skip_flag[xTbY][yTbY][cIdx] has a value of 1, the minimum value of qP can be limited by replacing 0 with the value of QpPrimeTsMin as shown in the following formula.

[0554] [Formula 50]

[0555] The quantization parameter qP can be updated as follows.

[0556] When transform_skip_flag[xTbY][yTbY][cIdx] has a value of 0, qP can be derived as shown in the following formula.

[0557] qP=Clip3(0,63,qP-(cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0))

[0558] Alternatively, when transform_skip_flag[xTbY][yTbY][cIdx] has a value of 1, qP can be derived as shown in the following formula.

[0559] qP=Clip3(QpPrimeTsMin,63,qP-cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0)

[0560] In another implementation, the quantization parameter qP can be updated as follows.

[0561] When transform_skip_flag[xTbY][yTbY][cIdx] has a value of 0, qP can be derived as shown in the following formula.

[0562] [Formula 51]

[0563] qP=Clip3(0,63+QpBdOffset,qP+(cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0))

[0564] Alternatively, when transform_skip_flag[xTbY][yTbY][cIdx] has a value of 1, qP can be derived as shown in the following formula.

[0565] [Equation 52]

[0566] qP=Clip3(0,63+QpBdOffset,Max(QpPrimeTsMin,qP)+(cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0)

[0567] In another implementation, when transform_skip_flag[xTbY][yTbY][cIdx] has a value of 1, the minimum value of qP can be limited by replacing 0 with the value of QpPrimeTsMin as shown in the following formula.

[0568] [Formula 53]

[0569] The quantization parameter qP can be updated as follows.

[0570] When transform_skip_flag[xTbY][yTbY][cIdx] has a value of 0, qP can be derived as shown in the following formula.

[0571] qP=Clip3(0,63+QpBdOffset,qP+(cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0))

[0572] Alternatively, when transform_skip_flag[xTbY][yTbY][cIdx] has a value of 1, qP can be derived as shown in the following formula.

[0573] qP=Clip3(QpPrimeTsMin,63+QpBdOffset,qP+cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0)

[0574] Implementation Method 3: A method that allows ACT when performing chroma BDPCM

[0575] In this implementation, when BDPCM is applied to the luma component block, ACT can be applied to encode / decode the block. However, when BDPCM is applied to the chroma component block, ACT can be restricted from being applied to encode / decode the block.

[0576] Furthermore, even when BDPCM is applied to a chroma component block, ACT is also applied to that block, thereby improving coding efficiency. Figure 39 An implementation of a syntax configuration that applies ACT even when BDPCM is applied to chroma component blocks is shown. Figure 39 As shown, by removing the condition that the cu_act_enabled_flag indicating whether ACT is applied to the current encoding base is obtained, the BDCPM syntax elements for the chroma components can be obtained regardless of whether ACT is applied to the chroma component blocks, and BDCPM encoding can be performed accordingly.

[0577] Implementation Method 4: Applying the ACT method even when performing encoding / decoding with CCLM

[0578] Both CCLM and ACT aim to eliminate unwanted overlap between components. There is some overlap between CCLM and ACT, but even after applying all of them, it's impossible to completely eliminate the overlap between components. Therefore, applying CCLM and ACT together can further reduce the overlap between components.

[0579] The following implementation describes an implementation where CCLM and ACT are used together. During decoding, the decoding device may first apply CCLM and then apply ACT. When ACT is applied to both BDPCM and CCLM for the chroma components, the syntax table used to signal this can be as follows: Figure 40 The modifications shown are therefore, as illustrated. Figure 40 As shown in the syntax table, among the restrictions on the syntax elements related to signaling notifications with intra_bdpcm_chroma and cclm, the if(!cu_act_enabled_flag) syntax element that signals notifications based on whether ACT is applied can be removed from the syntax table.

[0580] Implementation Method 5: Applying a flexible ACTQp method including combined CbCr

[0581] When applying the ACT mode, the prediction residual can be transformed from a color space (e.g., GBR or YCbCr) to the YCgCo color space. Furthermore, the residual based on the transformation can be encoded in the YCgCo color space. As an implementation of the ACT kernel transform (transform kernel) for color space transformation, the following transform kernel, as described above, can be used.

[0582] [Formula 54]

[0583]

[0584] [Formula 55]

[0585]

[0586] As stated in the above formula, the transformations of C0', C1', and C2' (here, C0' = Y, C1' = Cg, C2' = Co) are not normalized. For example, the L2 norm does not have a value of 1. For example, the L2 norm for the transformations of each component can have a value of approximately 0.6 for C0' and C1', and a value of approximately 0.7 for C2'. Here, the L2 norm is obtained as the square root of the sum of the squares of the individual coefficients. For example, C0' can be calculated as 2 / 4 * C0 + 1 / 4 * C1 + 1 / 4 * C2. Therefore, the norm of C0' can be calculated as the square root of (2 / 4 * 2 / 4 + 1 / 4 * 1 / 4 + 1 / 4 * 1 / 4). Therefore, this can be calculated as the square root of 6 / 16, and can be calculated as having a value of approximately 0.6.

[0587] When normalization transformation is not applied, the dynamic range of each component is irregular. Furthermore, this degrades the coding performance of typical video compression systems.

[0588] To compensate for the dynamic range of the residual signal, QP offset values ​​are sent to compensate for changes in the dynamic range of each transform component, enabling QP adjustment. For example, this implementation can be applied to a general QP adjustment control method for combined CbCr and ACT transforms.

[0589] The individual color components are not encoded independently, but together, which may cause variations in the dynamic range between the individual color components due to the method described above for joint CbCr in Implementation 3.

[0590] In the encoding and decoding method according to the implementation, the ACT QP offset adjustment can be fixed at -5, which can also be applied to Y, Cg and Co.

[0591] In this implementation, to provide flexible Qp control for each component and the combined CbCr, different ACT Qp offsets can be used for Y, Cb, Cr, and / or the combined CbCr. The ACT Qp offset values ​​can be determined based on the component index and / or the combined CbCr and / or the combined CbCr pattern.

[0592] To indicate the ACT Qp offset, ppsActQpOffsetY, ppsActQpOffsetCb, and ppsActQpOffsetCr can be used. Additionally, ppsActQpOffsetCbCr can be used for the ACT QP offset in combined CbCr mode 2 with CBF, where all Cb and Cr components have non-zero values. These values ​​(e.g., ppsActQpOffsetY, ppsActQpOffsetCb, ppsActQpOffsetCr, and ppsActQpOffsetCbCr) can be predetermined or signaled via a bitstream. The ACT QP offset in combined CbCr mode can be set in another way or set to another value.

[0593] In implementation, ACT Qp offsets -5, -5, and -3 can be used for Y, Cb, and Cr, and ACT Qp offset -4 can be used for combined CbCr.

[0594] In another implementation, ACT Qp offsets -5, -4, and -3 can be used for Y, Cb, and Cr, and ACT Qp offset -3 can be used for the joint CbCr mode where the value of tu_cbf_cb is not 0.

[0595] In another implementation, the ACT QP offset of the joint CbCr mode 2 can have its own offset value. For another joint CbCr mode, the ACT QP offset can use the offset of the corresponding component. For example, the quantization parameter qP can be determined as follows. First, when cIdx has a value of 0, qP and the ACT Qp offset can be derived as shown in the following equation.

[0596] [Formula 56]

[0597] qP=Qp' Y

[0598] ActQpOffset=ppsActQpOffsetY

[0599] Alternatively, when TuCResMode[xTbY][yTbY] has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0600] [Formula 57]

[0601] qP=Qp' CbCr

[0602] ActQpOffset=ppsActQpOffsetCbCr

[0603] Alternatively, when cIdx has a value of 1, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0604] [Formula 58]

[0605] qP=Qp' Cb

[0606] ActQpOffset=ppsActQpOffsetCb

[0607] Alternatively, when cIdx has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following equation.

[0608] [Formula 59]

[0609] qP=Qp' Cr

[0610] ActQpOffset=ppsActQpOffsetCr

[0611] In this implementation, the quantization parameter qP can be updated as follows.

[0612] When transform_skip_flag[xTbY][yTbY][cIdx] has a value of 0, qP can be derived as shown in the following formula.

[0613] [Formula 60]

[0614] qP=Clip3(0,63+QpBdOffset,qP+(cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0))

[0615] Alternatively, when transform_skip_flag[xTbY][yTbY][cIdx] has a value of 1, qP can be derived as shown in the following formula.

[0616] [Formula 61]

[0617] qP=Clip3(QpPrimeTsMin,63+QpBdOffset,qP+cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0)

[0618] In another implementation, for the joint CbCr mode where tu_cbf_cb! = 0 (e.g., in modes 1 and 2), the offset of the joint CbCr can be determined using ppsActQpOffsetCb. Alternatively, for the joint CbCr mode where tu_cbf_cb == 0 (e.g., in mode 3), the offset of the joint CbCr can be determined using ppsActQpOffsetCr. For example, the above implementation can be modified and applied as follows.

[0619] The quantization parameter qP can be updated as follows. First, when cIdx has a value of 0, qP and the ACT Qp offset can be derived as shown in the following equation.

[0620] [Formula 62]

[0621] qP=Qp' Y

[0622] ActQpOffset=ppsActQpOffsetY

[0623] Alternatively, when TuCResMode[xTbY][yTbY] has a value of 2, qP can be derived as shown in the following formula.

[0624] [Formula 63]

[0625] qP=Qp' CbCr

[0626] Alternatively, when cIdx has a value of 1, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0627] [Formula 64]

[0628] qP=Qp' C b

[0629] ActQpOffset=ppsActQpOffsetCb

[0630] Alternatively, when cIdx has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following equation.

[0631] [Formula 65]

[0632] qP=Qp' Cr

[0633] ActQpOffset=ppsActQpOffsetCr

[0634] When cIdx does not have a value of 0 and TuCResMode[xTbY][yTbY] does not have a value of 0, the ACT Qp offset for the joint CbCr mode can be determined according to the following pseudocode.

[0635] [Formula 66]

[0636] if(TuCResMode[xTbY][yTbY]is euqal to 1or 2)

[0637] ActQpOffset=ppsActQpOffsetCb;

[0638] else

[0639] ActQpOffset=ppsActQpOffsetCr;

[0640] In the implementation, the quantization parameter qP can be updated as follows. When transform_skip_flag[xTbY][yTbY][cIdx] has a value of 0, qP can be derived as shown in the following formula.

[0641] [Formula 67]

[0642] qP=Clip3(0,63+QpBdOffset,qP+(cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0))

[0643] Alternatively, when transform_skip_flag[xTbY][yTbY][cIdx] has a value of 1, qP can be derived as shown in the following formula.

[0644] [Formula 68]

[0645] qP=Clip3(QpPrimeTsMin,63+QpBdOffset,qP+cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0)

[0646] In another implementation, regardless of the joint CbCr mode, ppsActQpOffsetY is used when the component index is Y, ppsActQpOffsetCb is used when the component index is Cb, and ppsActQpOffsetCr is used when the component index is Cr, thereby deriving qP. For example, the quantization parameter qP can be derived as follows.

[0647] First, when cIdx has a value of 0, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0648] [Formula 69]

[0649] qP=Qp' Y

[0650] ActQpOffset=ppsActQpOffsetY

[0651] Alternatively, when TuCResMode[xTbY][yTbY] has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0652] [Formula 70]

[0653] qP=Qp' C b C r

[0654] ActQpOffset=(cIdx==1)? ppsActQpOffsetCb:ppsActQpOffsetCr

[0655] Alternatively, when cIdx has a value of 1, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0656] [Formula 71]

[0657] qP=Qp' Cb

[0658] ActQpOffset=ppsActQpOffsetCb

[0659] Alternatively, when cIdx has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following equation.

[0660] [Equation 72]

[0661] qP=Qp' Cr

[0662] ActQpOffset=ppsActQpOffsetCr

[0663] The quantization parameter qP can be updated as follows.

[0664] When transform_skip_flag[xTbY][yTbY][cIdx] has a value of 0, qP can be derived as shown in the following formula.

[0665] [Formula 73]

[0666] qP=Clip3(0,63+QpBdOffset,qP+(cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0))

[0667] Alternatively, when transform_skip_flag[xTbY][yTbY][cIdx] has a value of 1, qP can be derived as shown in the following formula.

[0668] [Formula 74]

[0669] qP=Clip3(QpPrimeTsMin,63+QpBdOffset,qP+cu_act_enabled_flag[xTbY][yTbY]?ActQpOffset:0)

[0670] Implementation Method 6: A method for signaling the ACT Qp offset including combined CbCr

[0671] The following describes an example of signaling the ACT QP offset via a bitstream to provide greater flexibility. The ACT QP offset can be signaled via SPS, PPS, picture header, slice header, or other types of header sets. The combined CbCr ACT Qp offset can be signaled separately, or it can be derived from the ACT Qp offsets for Y, Cb, and Cr.

[0672] Without loss of generality, Figure 41 An example of the syntax table where the ACT Qp offset is signaled in PPS is shown. (See also...) Figure 41 In this implementation, for the combined CbCr, an ACT Qp offset can be signaled. This will be described in... Figure 41 The syntax elements indicated in the syntax table.

[0673] The syntax element pps_act_qp_offsets_present_flag can indicate whether a syntax element associated with an ACTQP offset exists in the PPS. For example, pps_act_qp_offsets_present_flag can indicate whether the syntax elements pps_act_y_qp_offset_plusX1, pps_act_cb_qp_offset_plusX2, pps_act_cr_qp_offset_plusX3, and pps_act_cbcr_qp_offset_plusX4, which will be described later, are signaled as part of the PPS.

[0674] For example, the first value of pps_act_qp_offsets_present_flag (e.g., 0) can instruct pps_act_y_qp_offset_plusX1, pps_act_cb_qp_offset_plusX2, pps_act_cr_qp_offset_plusX3, and pps_act_cbcr_qp_offset_plusX4 not to signal via the PPS syntax table.

[0675] The second value of pps_act_qp_offsets_present_flag (e.g., 1) can instruct pps_act_y_qp_offset, pps_act_cb_qp_offset, and pps_act_cr_qp_offset to signal via the PPS syntax table.

[0676] When pps_act_qp_offsets_present_flag is not provided from the bitstream, pps_act_qp_offsets_present_flag can be deduced to have a first value (e.g., 0). For example, when a flag indicating whether an ACT is applied (e.g., sps_act_enabled_flag, which signals in SPS) has a first value (e.g., 0) indicating that an ACT is not applied, pps_act_qp_offsets_present_flag can be forced to have a first value (e.g., 0).

[0677] When the value of the syntax element `cu_act_enabled_flag` is a second value (e.g., 1) indicating that ACT is applied to the current encoding base, the syntax elements `pps_act_y_qp_offset_plusX1`, `pps_act_cb_qp_offset_plusX2`, `pps_act_cr_qp_offset_plusX3`, and `pps_act_cbcr_qp_offset_plusX4` can be used to determine the offsets applied to the quantization parameter values ​​`qP` for the luminance, Cb, Cr components, and joint CbCr components, respectively. When `pps_act_y_qp_offset_plusX1`, `pps_act_cb_qp_offset_plusX2`, `pps_act_cr_qp_offset_plusX3`, and `pps_act_cbcr_qp_offset_plusX4` are not present in the bitstream, each value can be set to 0.

[0678] The values ​​of variables PpsActQpOffsetY, PpsActQpOffsetCb, PpsActQpOffsetCr, and PpsActQpOffsetCbCr can be determined according to the syntax elements as shown in the following formula.

[0679] [Formula 75]

[0680] PpsActQpOffsetY=pps_act_y_qp_offset_plusX1–X1

[0681] PpsActQpOffsetCb=pps_act_cb_qp_offset_plusX2-X2

[0682] PpsActQpOffsetCr=pps_act_cr_qp_offset_plusX3–X3

[0683] PpsActQpOffsetCbCr=pps_act_cbcr_qp_offset_plusX4–X4

[0684] Here, X1, X2, X3, and X4 can indicate predetermined constant values. These can be the same value, different values, or only some of them can have the same value.

[0685] In the implementation, for bitstream matching, the values ​​of PpsActQpOffsetY, PpsActQpOffsetCb, PpsActQpOffsetCr, and PpsActQpOffsetCbCr can be limited to values ​​ranging from -12 to 12.

[0686] Based on the determination of the variables, the quantization parameter qP can be determined as follows. First, when cIdx has a value of 0, qP and the ACT Qp offset can be derived as shown in the following formula.

[0687] [Formula 76]

[0688] qP=Qp' Y

[0689] ActQpOffset=PpsActQpOffsetY

[0690] Alternatively, when TuCResMode[xTbY][yTbY] has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0691] [Formula 77]

[0692] qP=Qp' C bC r

[0693] ActQpOffset=PpsActQpOffsetCbCr

[0694] Alternatively, when cIdx has a value of 1, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0695] [Formula 78]

[0696] qP=Qp' Cb

[0697] ActQpOffset=PpsActQpOffsetCb

[0698] Alternatively, when cIdx has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following equation.

[0699] [Formula 79]

[0700] qP=Qp' Cr

[0701] ActQpOffset=PpsActQpOffsetCr

[0702] In another implementation of signaling ACT Qp offsets, multiple ACT QP offsets can be signaled for different joint CbCr modes identified as mode A and mode B.

[0703] Joint CbCr pattern A can refer to joint CbCr patterns such as patterns 1 and 2 in Table 2 above, which have non-zero values ​​for tu_cbf_cb. Joint CbCr pattern B can refer to joint CbCr patterns such as pattern 3 in Table 2 above, which have a value of 0 for tu_cbf_cb. Figure 42 The corresponding syntax table is shown in [the document]. A description will follow in [the document / section]. Figure 42 The syntax elements indicated in the syntax table.

[0704] When the value of the syntax element cu_act_enabled_flag is a second value (e.g., 1) indicating that ACT is applied to the current encoding base, the syntax elements pps_act_y_qp_offset_plusX1, pps_act_cb_qp_offset_plusX2, pps_act_cr_qp_offset_plusX3, pps_act_cbcr_qp_offset_modeA_plusX4, and pps_act_cbcr_qp_offset_modeB_plusX5 can be used to determine the offsets applied to the quantization parameter values ​​qP for the luminance, Cb, Cr components, and joint CbCr components, respectively. When the values ​​of pps_act_y_qp_offset_plusX1, pps_act_cb_qp_offset_plusX2, pps_act_cr_qp_offset_plusX3, pps_act_cbcr_qp_offset_modeA_plusX4, and pps_act_cbcr_qp_offset_modeB_plusX5 are not present in the bitstream, each value can be set to 0.

[0705] Based on the syntax elements, the values ​​of variables PpsActQpOffsetY, PpsActQpOffsetCb, PpsActQpOffsetCr, PpsActQpOffsetCbCrModeA, and PpsActQpOffsetCbCrModeB can be determined as shown in the following formula.

[0706] [Formula 80]

[0707] PpsActQpOffsetY=pps_act_y_qp_offset_plusX1–X1

[0708] PpsActQpOffsetCb=pps_act_cb_qp_offset_plusX2-X2

[0709] PpsActQpOffsetCr=pps_act_cr_qp_offset_plusX3–X3

[0710] PpsActQpOffsetCbCrModeA=pps_act_cbcr_qp_offset_modeA_plusX4–X4

[0711] PpsActQpOffsetCbCrModeB=pps_act_cbcr_qp_offset_modeB_plusX5–X5

[0712] Here, X1, X2, X3, X4, and X5 can indicate predetermined constant values. These can be the same value, different values, or only some can have the same value. In the implementation, for bitstream matching, the values ​​of PpsActQpOffsetY, PpsActQpOffsetCb, PpsActQpOffsetCr, PpsActQpOffsetCbCrModeA, and PpsActQpOffsetCbCrModeB can be limited to values ​​ranging from -12 to 12.

[0713] Based on the determination of the variables, the quantization parameter qP can be determined as follows. First, when cIdx has a value of 0, qP and the ACT Qp offset can be derived as shown in the following formula.

[0714] [Formula 81]

[0715] qP=Qp' Y

[0716] ActQpOffset=PpsActQpOffsetY

[0717] Alternatively, when TuCResMode[xTbY][yTbY] has a value of 2, qP can be derived as shown in the following formula.

[0718] [Equation 82]

[0719] qP=Qp' C b C r

[0720] Alternatively, when cIdx has a value of 1, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0721] [Equation 83]

[0722] qP=Qp' C b

[0723] ActQpOffset=PpsActQpOffsetCb

[0724] Alternatively, when cIdx has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following equation.

[0725] [Formula 84]

[0726] qP=Qp' Cr

[0727] ActQpOffset=PpsActQpOffsetCr

[0728] Additionally, when cIdx does not have a value of 0 and TuCResMode[xTbY][yTbY] does not have a value of 0, the ACT Qp offset can be derived as shown in the following formula.

[0729] [Formula 85]

[0730] ActQpOffset=(tu_cbf_cb[xTbY][yTbY])? PpsActQpOffsetCbCrModeA:PpsActQpOffsetCbCrModeB

[0731] Furthermore, in another implementation, when TuCResMode[xTbY][yTbY] has a value of 2, ActQpOffset can be derived as shown in the following formula.

[0732] [Formula 86]

[0733] ActQpOffset=(tu_cbf_cb[xTbY][yTbY])?

[0734] (PPsQpOffsetCbCrModeA+slice_act_CbCr_qp_offset_ModeA):

[0735] (PPsQpOffsetCbCrModeB+slice_act_CbCr_qp_offset_ModeB)

[0736] In another implementation that signals the ACT Qp offset, such as in Figure 43 In the syntax table, only the ACT QP offsets for Y, Cb, and Cr can be signaled. The ACT QP offset for the joint CbCr can be derived from PpsActQpOffsetY, PpsActQpOffsetCb, and / or PpsActQpOffsetCr.

[0737] In one implementation, the ACT Qp offset for CbCr can be set to the value of PpsActQpOffsetCb. In another implementation, the ACT Qp offset for CbCr can be set to the same value as PpsActQpOffsetCb in the case of a joint CbCr pattern where tu_cbf_cb has a non-zero value, or it can be set to the same value as PpsActQpOffsetCr in the case of a joint CbCr pattern where tu_cbf_cb has a value of 0. Alternatively, it can be set in the opposite way.

[0738] Figure 43 This is a view illustrating another implementation of the syntax table in which the ACT Qp offset is signaled in the PPS. According to Figure 43 The grammatical elements are determined as follows, which determines the quantization parameter qP. First, when cIdx has a value of 0, qP and the ACT Qp offset can be derived as shown in the following equation.

[0739] [Formula 87]

[0740] qP=Qp' Y

[0741] ActQpOffset=PpsActQpOffsetY

[0742] Alternatively, when TuCResMode[xTbY][yTbY] has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0743] [Formula 88]

[0744] qP=Qp' C b C r

[0745] ActQpOffset=(cIdx==1)? PpsActQpOffsetCb:PpsActQpOffsetCr

[0746] In another embodiment, the value of ActQpOffset can be determined as follows.

[0747] [Formula 89]

[0748] ActQpOffset=(tu_cbf_cb[xTbY][yTbY])? PpsActQpOffsetCb:PpsActQpOffsetCr

[0749] Alternatively, when cIdx has a value of 1, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0750] [Formula 90]

[0751] qP=Qp' Cb

[0752] ActQpOffset=PpsActQpOffsetCb

[0753] Alternatively, when cIdx has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following equation.

[0754] [Formula 91]

[0755] qP=Qp' Cr

[0756] ActQpOffset=PpsActQpOffsetCr

[0757] Implementation Method 7: Signaling ACT Qp offset at multiple levels

[0758] In implementations, ACT QP offsets can be signaled at multiple levels. In addition to signaling ACT QP offsets at a level such as PPS as in Implementation 6 above, ACT QP offsets can be signaled at lower levels (e.g., slice headers, image headers, or other types of headers suitable for Qp control).

[0759] The following will describe two implementation methods. Figure 44 and Figure 45 This illustrates an example where the ACT QP offset is notified via sliced ​​head and image hair signals. In this way, the ACT QP offset can be signaled at multiple levels.

[0760] The following will describe Figure 44 and Figure 45 The syntax elements shown are: pps_slice_act_qp_offsets_present_flag, which can indicate whether the syntax elements slice_act_y_qp_offset, slice_act_cb_qp_offset, slice_act_cr_qp_offset, and slice_act_cbcr_qp_offset, which will be described later, exist in the slice header.

[0761] For example, the first value of pps_slice_act_qp_offsets_present_flag (e.g., 0) can indicate that slice_act_y_qp_offset, slice_act_cb_qp_offset, slice_act_cr_qp_offset, and slice_act_cbcr_qp_offset do not exist in the slice header.

[0762] For example, the second value of pps_slice_act_qp_offsets_present_flag (e.g., 1) can indicate the presence of slice_act_y_qp_offset, slice_act_cb_qp_offset, slice_act_cr_qp_offset, and slice_act_cbcr_qp_offset in the slice header.

[0763] The syntax elements slice_act_y_qp_offset, slice_act_cb_qp_offset, slice_act_cr_qp_offset, and slice_act_cbcr_qp_offset can indicate the offset of the quantization parameter value qP for the luminance, Cb, Cr, and joint CbCr components, respectively. The values ​​of slice_act_y_qp_offset, slice_act_cb_qp_offset, slice_act_cr_qp_offset, and slice_act_cbcr_qp_offset can be restricted to values ​​ranging from -12 to 12. Each value can be set to 0 when it is not present in the bitstream. The values ​​of PpsActQpOffsetY+slice_act_y_qp_offset, PpsActQpOffsetCb+slice_act_cb_qp_offset, PpsActQpOffsetCr+slice_act_cr_qp_offset, and PpsActQpOffsetCbCr+slice_act_cbcr_qp_offset can also be restricted to values ​​ranging from -12 to 12.

[0764] This can be applied to various implementations that signal at the PPS level for the ACT QP offset of the combined CbCr. For example, one QP offset can be signaled for the combined CbCr, multiple ACT Qp offsets can be signaled for different modes of the combined CbCr, or, without signaling the ACT Qp offset for the combined CbCr, a method can be applied that derives this by using the ACTQpOffsets for Y, Cb, and Cr and / or the mode of the combined CbCr when signaling via slice hair.

[0765] exist Figure 46 and Figure 47 Two modified implementations are shown in the figure. Figure 46 An implementation of which signals the ACT Qp offset in the slice header is shown. Figure 47 Another implementation is shown in which the ACT Qp offset is signaled in the slice header. Figure 47 In this implementation, only the ACT Qp offsets for Y, Cb, and Cr can be signaled, and the slice-level ACT Qp offset for the joint CbCr can be derived from slice_act_y_qp_offset, slice_act_cb_qp_offset, and / or slice_act_cr_qp_offset. This can be determined based on the mode type of the joint CbCr. In one implementation, the slice-level ACT Qp offset for CbCr can be set to the same value as slice_act_cb_qp_offset. In another implementation, in the case of a joint CbCr mode with a non-zero value tu_cbf_cb, the slice-level ACT Qp offset for the joint CbCr can be set to the same value as slice_act_cb_qp_offset. Additionally, in the case of a joint CbCr mode with a value of 0 for tu_cbf_cb, the slice-level ACT Qp offset for the joint CbCr can be set to the same value as slice_act_cr_qp_offset.

[0766] In another implementation, syntax elements can be signaled in the slice header or image header. To achieve this, encoding / decoding can be performed as follows.

[0767] - The flag pps_picture_slice_act_qp_offsets_present_flag, which indicates whether there is an ACT Qp offset in the image header or slice header, can be signaled in PPS.

[0768] - When ACT is applicable and the value of pps_picture_slice_act_qp_offsets_present_flag is the second value (e.g., 1), the flag pic_act_qp_offsets_present_flag indicating the presence of ACT Qp offsets in the image header is signaled in the image header. Here, the second value of pic_act_qp_offsets_present_flag (e.g., 1) can indicate that the ACT Qp offsets for all slices of the image corresponding to the image header are provided in the image header.

[0769] The first value of `-pic_act_qp_offsets_present_flag` (e.g., 0) can indicate that the ACT Qp offsets for all slices of the image corresponding to the image header are not provided in the image header. For example, when ACT applies and the value of `pps_picture_slice_act_qp_offsets_present_flag` is the second value (e.g., 1) and the value of `pic_act_qp_offsets_present_flag` is the first value (e.g., 0), the ACT Qp offsets for the slices can be provided in the slice header.

[0770] Figure 48 This is a view of the syntax table of the PPS in which the signal `pps_pic_slice_act_qp_offsets_present_flag` is sent. The syntax element `pps_pic_slice_act_qp_offsets_present_flag` can indicate whether the ACT Qp offset is provided in the picture header and / or slice header. For example, a first value (e.g., 0) of `pps_pic_slice_act_qp_offsets_present_flag` can indicate that the ACT Qp offset is not provided in the picture header and slice header. A second value (e.g., 1) of `pps_pic_slice_act_qp_offsets_present_flag` can indicate that the ACT Qp offset is provided in the picture header or slice header. When `pps_pic_slice_act_qp_offsets_present_flag` is not provided in the bitstream, the value of `pps_pic_slice_act_qp_offsets_present_flag` can be determined to be the first value (e.g., 0).

[0771] Figure 49This is a view showing the syntax table of the picture header used to signal the ACT Qp offset. The syntax element `pic_act_qp_offsets_present_flag` indicates whether the ACT Qp offset is provided in the picture header. A first value of `pic_act_qp_offsets_present_flag` (e.g., 0) indicates that the ACT Qp offset is not provided in the picture header but is provided in the slice header. A second value of `pic_act_qp_offsets_present_flag` (e.g., 1) indicates that the ACT Qp offset is provided in the picture header. When the value of `pic_act_qp_offsets_present_flag` is not provided in the bitstream, the value can be determined to be 0.

[0772] Figure 50 This is a view showing the syntax table of the slice header used to signal the ACT Qp offset. Figure 50 In the syntax table, the syntax elements slice_act_y_qp_offset, slice_act_cb_qp_offset, slice_act_cr_qp_offset, and slice_act_cbcr_qp_offset indicate the offset of the quantization parameter value qP for the luminance, Cb, and Cr components. The values ​​of slice_act_y_qp_offset, slice_act_cb_qp_offset, slice_act_cr_qp_offset, and slice_act_cbcr_qp_offset can have values ​​ranging from -12 to 12. Additionally, the values ​​of PpsActQpOffsetY+slice_act_y_qp_offset, PpsActQpOffsetCb+slice_act_cb_qp_offset, and PpsActQpOffsetCr+slice_act_cr_qp_offset can be restricted to a value range from -12 to 12.

[0773] Furthermore, if the values ​​of slice_act_y_qp_offset, slice_act_cb_qp_offset, slice_act_cr_qp_offset, and slice_act_cbcr_qp_offset are not provided in the bitstream, the values ​​of slice_act_y_qp_offset, slice_act_cb_qp_offset, and slice_act_cr_qp_offset can be determined to be 0 when the value of pps_pic_slice_act_qp_offsets_present_flag is the first value (e.g., 0). Alternatively, when the value of pps_pic_slice_act_qp_offsets_present_flag is the second value (e.g., 1), the values ​​of slice_act_y_qp_offset, slice_act_cb_qp_offset, and slice_act_cr_qp_offset can be determined to be the same as pps_act_y_qp_offset, pps_act_cb_qp_offset, and pps_act_cr_qp_offset, respectively.

[0774] Furthermore, when there is an ACT Qp offset in both the slice header and the image header, the final offset value used to derive the qP value can be determined by adding the offset value signaled in the PPS to the offset value signaled in the slice header or image header.

[0775] More specifically, in the implementation, the quantization parameter qP can be determined as follows. First, when cIdx has a value of 0, qP and the ACT Qp offset can be derived as shown in the following equation.

[0776] [Equation 92]

[0777] qP=Qp' Y

[0778] ActQpOffset=PPsQpOffsetY+slice_act_y_qp_offset

[0779] Alternatively, when TuCResMode[xTbY][yTbY] has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0780] [Formula 93]

[0781] qP=Qp' CbCr

[0782] ActQpOffset=PPsQpOffsetCbCr+slice_act_CbCr_qp_offset

[0783] Alternatively, when cIdx has a value of 1, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0784] [Formula 94]

[0785] qP=Qp' Cb

[0786] ActQpOffset=PpsActQpOffsetCb+slice_act_Cb_qp_offset

[0787] Alternatively, when cIdx has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following equation.

[0788] [Formula 95]

[0789] qP=Qp' Cr

[0790] ActQpOffset=PpsActQpOffsetCr+slice_act_Cr_qp_offset

[0791] In another implementation, when multiple ACT Qp offsets for the joint CbCr are signaled, the ActQpOffset for the joint CbCr can be determined as follows.

[0792] First, when cIdx has a value of 0, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0793] [Formula 96]

[0794] qP=Qp' Y

[0795] ActQpOffset=PPsQpOffsetY+slice_act_y_qp_offset

[0796] Alternatively, when TuCResMode[xTbY][yTbY] has a value of 2, qP can be derived as shown in the following formula.

[0797] [Formula 97]

[0798] qP=Qp' C b C r

[0799] Alternatively, when cIdx has a value of 1, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0800] [Formula 98]

[0801] qP=Qp' C b

[0802] ActQpOffset=PpsActQpOffsetCb+slice_act_Cb_qp_offset

[0803] Alternatively, when cIdx has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following equation.

[0804] [Formula 99]

[0805] qP=Qp' Cr

[0806] ActQpOffset=PpsActQpOffsetCr+slice_act_Cr_qp_offset

[0807] Additionally, when cIdx does not have a value of 0 and TuCResMode[xTbY][yTbY] does not have a value of 0, the ACT Qp offset can be derived as shown in the following formula.

[0808] [Formula 100]

[0809] ActQpOffset=(tu_cbf_cb[xTbY][yTbY])?

[0810] (PPsQpOffsetCbCrModeA+slice_act_CbCr_qp_offset_ModeA):

[0811] (PPsQpOffsetCbCrModeB+slice_act_CbCr_qp_offset_ModeB)

[0812] In another embodiment, when no ACT Qp offset for the combined CbCr is provided, the ActQpOffset and Qp of the Y, Cb, and / or Cr components are determined, and the ActQpOffset for the combined CbCr can be determined using the ACT Qp offsets of the Y, Cb, and / or Cr components as follows. For example, in the above embodiment, when TuCResMode[xTbY][yTbY] associated with Equation 97 has a value of 2, the calculation steps of qP can be changed and performed as follows.

[0813] "Alternatively, when TuCResMode[xTbY][yTbY] has a value of 2, the offsets qP and ACT Qp can be derived as shown in the following formula."

[0814] [Formula 101]

[0815] qP=Qp' CbCr

[0816] ActQpOffset=(cIdx==1])? (PPsQpOffsetCb+slice_act_Cb_qp_offset):(PPsQpOffsetCr+slice_act_Cr_qp_offset)”

[0817] In another embodiment, the value of ActQpOffset can be determined as follows.

[0818] [Equation 102]

[0819] ActQpOffset=(tu_cbf_cb[xTbY][yTbY])? (PPsQpOffsetCb+slice_act_Cb_qp_offset):(PPsQpOffsetCr+slice_act_Cr_qp_offset)

[0820] Implementation Method 8: Method for Signaling Multiple ACTQp Offset Sets

[0821] In this embodiment, a method for using a list of ACT Qp offsets will be described. For this purpose, the following processing can be performed.

[0822] a) Multiple sets of ACT Qp offsets can be signaled as a list within a parameter set (e.g., SPS or PPS). Each set in the list can include ACT Qp offsets for the Y, Cb, Cr, and joint CbCr components. For simplicity, the list of ACT Qp offsets can be signaled from the same parameter set as the list used to signal the chromaticity Qp offsets.

[0823] b) The number of sets of ACT Qp offsets in the list can be the same as the number of sets of chroma Qp offsets that are signaled in PPS.

[0824] c) As the ACT Qp offset used to derive qP for each coding basis, the ACT Qp offset can be a list of indices (e.g., cu_chroma_qp_offset_idx) that belong to the chroma Qp offset for the coding basis.

[0825] d) As an alternative implementation to b) and c), the following can be performed.

[0826] - It can signal the number of ACT Qp offset sets in the list. The number of ACT Qp offset sets in the list can be different from the number of chroma Qp offset sets.

[0827] - When ACT is applied, the index of the index indicating the ACT Qp offset used for encoding can be signaled.

[0828] Without departing from the above concept, we can use the syntax of signaling a list of ACT Qp offsets, such as... Figure 51 As shown in the diagram. For example, when cu_act_enabled_flag has a value of 1, pps_act_y_qp_offset, pps_act_cb_qp_offset, pps_act_cr_qp_offset, and pps_act_cbcr_qp_offset can be used to determine the offsets of the quantization parameter values ​​qP applied to the luminance, Cb, Cr components, and joint CbCr, respectively.

[0829] When there are no values ​​for pps_act_y_qp_offset, pps_act_cb_qp_offset, pps_act_cr_qp_offset, and pps_act_cbcr_qp_offset, each value can be deduced to be 0.

[0830] When the value of cu_act_enabled_flag is the second value (e.g., 1) and the value of cu_chroma_qp_offset_flag is the second value (e.g., 1), act_y_qp_offset_list[i], act_cb_qp_offset_list[i], act_cr_qp_offset_list[i], and act_cbcr_qp_offset_list[i] can be used to determine the offsets of the quantization parameter values ​​qP applied to the luminance, Cb, Cr components, and joint CbCr components, respectively. When no values ​​exist for act_y_qp_offset_list[i], act_cb_qp_offset_list[i], act_cr_qp_offset_list[i], and act_cbcr_qp_offset_list[i], each value can be deduced to be 0.

[0831] In this embodiment, the quantization parameter qP can be determined as follows. First, when cIdx has a value of 0, qP and the ACT Qp offset can be derived as shown in the following formula.

[0832] [Equation 103]

[0833] qP=Qp' Y

[0834] ActQpOffset = pps_act_y_qp_offset + (cu_chroma_qp_offset_flag) ? act_y_qp_offset_list[cu_chroma_qp_offset_idx]: 0 + slice_act_y_qp_offset Alternatively, when TuCResMode[xTbY][yTbY] has a value of 2, the qP and ACT Qp offsets can be derived as shown in the following formula.

[0835] [Equation 104]

[0836] qP=Qp' C b C r

[0837] ActQpOffset=pps_act_cbcr_qp_offset+(cu_chroma_qp_offset_flag)? act_cbcr_qp_offset_list[cu_chroma_qp_offset_idx]:0+slice_act_cbcr_qp_offset

[0838] Alternatively, when cIdx has a value of 1, the offsets qP and ACT Qp can be derived as shown in the following formula.

[0839] [Equation 105]

[0840] qP=Qp' C b

[0841] ActQpOffset = pps_act_cb_qp_offset + (cu_chroma_qp_offset_flag) ? act_cb_qp_offset_list[cu_chroma_qp_offset_idx]: 0 + slice_act_cb_qp_offset Alternatively, when cIdx has a value of 2, the qP and ACT Qp offsets can be derived as shown in the following formula.

[0842] [Equation 106]

[0843] qP=Qp' Cr

[0844] ActQpOffset=pps_act_cr_qp_offset+(cu_chroma_qp_offset_flag)? act_cr_qp_offset_list[cu_chroma_qp_offset_idx]:0slice_act_cr_qp_offset

[0845] Implementation Method 9: ACT Color Space Transformation Method Applied to Both Lossless and Lossy Encoding

[0846] The transformations between color spaces based on the matrices used for forward and backward transformations described above can be organized as follows.

[0847] [Table 3]

[0848]

[0849] Because some values ​​are lost during the processing of Co and Cg, this transformation cannot reconstruct the original state. For example, when sample values ​​in the RGB color space are converted to sample values ​​in the YCgCo color space and then converted back to sample values ​​in the RGB color space, the original sample values ​​are not fully reconstructed. Therefore, the transformation according to Table 3 cannot be used for lossless coding. It is necessary to improve the color space transformation algorithm so that even when applying lossless coding, there is no loss of sample values ​​after the color space transformation. Implementation methods 9 and 10 describe color space transformation algorithms that can be applied to both lossless and lossy coding.

[0850] In the following embodiments, a method for performing ACT using a color space transformation that can be restored to its original state is described, applicable to both lossless and lossy encoding. This restoreable color space transformation can be applied to the encoding and decoding methods described above. The ACT Qp offset can also be adjusted for the following color space transformation. The color space transformation according to the embodiments can be performed as shown in the following formula. For example, a forward conversion from the GBR color space to the YCgCo color space can be performed according to the following formula.

[0851] [Equation 107]

[0852] Co = RB;

[0853] t = B + (Co >> 1);

[0854] Cg = Gt;

[0855] Y = t + (Cg >> 1);

[0856] Alternatively, a backward conversion from the YCgCo color space to the GBR color space can be performed according to the following formula.

[0857] [Equation 108]

[0858] t = Y - (Cg >> 1)

[0859] G = Cg + t

[0860] B = t - (Co >> 1)

[0861] R = Co + B

[0862] The transformation between the YCgCo and RGB color spaces according to the above formula can be restored to the original state. That is, the color space transformation according to the formula supports perfect reconstruction. For example, even if a backward transformation is performed after a forward transformation, the sample values ​​are preserved. Therefore, the color space transformation according to the formula can be called a recoverable YCgCo-R color transformation. Here, R can stand for recoverable, which means achieving reconstruction to the original state. Compared with existing transformations, the YCgCo-R transformation can be provided by increasing the bit depth of Cg and Co by 1. If this condition is met, other types of recoverable transformations can be used like the transformations described above.

[0863] Because the transformation shown in the above formula has a different norm value than the transformations mentioned above, the ACT Qp offset for Y, Cg, and Co can be adjusted to compensate for the dynamic range changes caused by the color space transformation.

[0864] It has been described that when the above transformation is applied, the QCT Qp offset according to the embodiment can have values ​​(-5, -5, -5) for Y, Cg, and Co. However, when the recoverable transformation according to this embodiment is applied, other values ​​other than (-5, -5, -5) can be specified as the QCT Qp offset according to the embodiment. For example, the values ​​(-5, 1, 3) for Y, Cg, and Co can be used as the QCT Qp offset according to the embodiment.

[0865] In another embodiment, the ACT QP offset can be signaled via the bitstream described in embodiment 6 or 7 above.

[0866] For example, when the YCgCo-R transform described above is used with the ACT QP offset (-5,1,3), no coding loss is observed in lossy coding environments (e.g., QP 22, 27, 32, 37), as shown in the figure below. Furthermore, it is observed that applying ACT further improves coding performance by 5% when achieving lossless coding.

[0867] [Table 4]

[0868] sequence Y U V RGB, TGM 1080p 0.0% 0.2% 0.1% RGB, TGM 720p 0.2% -0.1% 0.1% RGB, animation -0.1% -0.1% 0.0% RGB, Mixed Content -0.1% 0.0% -0.1% RGB, content captured by the camera -0.3% 0.2% -0.3% Overall (RGB) 0.0% 0.0% 0.0%

[0869] The VVC specification, including the integrated ACT matrix, can be described in the table below.

[0870] [Table 5]

[0871]

[0872] For example, a residual sample array r of size (nTbW)×(nTbH) can be updated as follows: Y r Cb and r Cr [Equation 109]

[0873] tmp=r Y [x][y]–(r Cb [x][y]>>1)

[0874] r Y [x][y] = tmp + r Cb [x][y]

[0875] r Cb [x][y]=tmp–(r Cr [x][y]>>1)

[0876] r Cr [x][y]=r Cb [x][y]+r Cr [x][y]

[0877] Implementation Method 10: ACT Execution Method for Performing Multiple Color Transformations Based on Explicit Signaling

[0878] In this embodiment, at least one color transformation can be performed via ACT. Which color transformation is performed can be determined by flags signaled in the bitstream. These flags can be signaled at multiple levels, such as SPS, PPS, image headers, and slices, or at a recognizable granularity.

[0879] In this implementation, a predetermined flag can be signaled to indicate which ACT (Activity) to apply. For example, when the flag has a value of 1, an ACT based on a reversible color change can be applied. When the flag has a value of 0, an ACT based on a non-reversible color change can be applied.

[0880] In another implementation, a predetermined flag of the ACT can be signaled to indicate which color change to use. Figure 52 The text shows an example of the syntax for signaling notifications in SPS. It will be described... Figure 52The syntax element `sps_act_reversible_conversion` indicates whether to use a transformation formula that is irreversible to the original state. A first value of `sps_act_reversible_conversion` (e.g., 0) indicates that `ACT` uses a transformation formula that is irreversible to the original state. A second value of `sps_act_reversible_conversion` (e.g., 1) indicates that `ACT` uses a transformation formula that is reversible to the original state.

[0881] Therefore, the variable lossyCoding, which indicates whether lossy coding is performed, can be set as follows.

[0882] [Formula 110]

[0883] lossyCoding=(!sps_act_reversible_conversion)

[0884] By using the lossyCoding flag, the pseudocode for the decoding device to perform a backward conversion from YCgCo to GBR during the decoding process can be represented as follows.

[0885] [Formula 111]

[0886] If(sps_act_reversible_conversion==1)

[0887] {

[0888] / / YCgCo-R reversible conversion

[0889] t = Y - (Cg >> 1)

[0890] G = Cg + t

[0891] B = t - (Co >> 1)

[0892] R = Co + B

[0893] }

[0894] else{

[0895] t=Y–Cg

[0896] G = Y + Cg

[0897] B = t - Co

[0898] R = t + Co

[0899] }

[0900] Therefore, the VVC specification shown in Table 5 of Implementation 9 can be modified as shown in the table below.

[0901] [Table 6]

[0902]

[0903] Based on the table above, the residual update process using color space transformation can use the following parameters as input to this process: - Variable nTbW, which indicates the block width.

[0904] - The variable nTbH indicates the block height.

[0905] -by element r Y An array r of size (nTbW) × (nTbH) of luminance residual samples composed of [x] and [y]. Y

[0906] -by element r Cb An array r of size (nTbW) × (nTbH) of chromaticity residual samples composed of [x] and [y]. Cb

[0907] -by element r Cr An array r of size (nTbW) × (nTbH) of chromaticity residual samples composed of [x] and [y]. Cr ,

[0908] The output of this process is as follows.

[0909] - The updated array r of size (nTbW)×(nTbH) of the luminance residual samples Y ,

[0910] - The updated array r of size (nTbW) × (nTbH) of the chromaticity residual samples Cb ,

[0911] - The updated array r of size (nTbW) × (nTbH) of the chromaticity residual samples Cr ,

[0912] By performing this process, the residual sample array r of size (nTbW)×(nTbH) can be updated as follows: Y r Cb and r Cr .

[0913] First, when the value of sps_act_reversible_conversion is the second value (e.g., 1), the residual sample array r of size (nTbW) × (nTbH) can be updated as shown in the following formula. Y r Cb and r Cr.

[0914] [Equation 112]

[0915] tmp=r Y [x][y]-(r Cb [x][y]>>1))

[0916] r Y [x][y] = tmp + r Cb [x][y])

[0917] r Cb [x][y]=tmp-(r Cr [x][y]>>1))

[0918] r Cr [x][y]=r Cb [x][y]+r Cr [x][y]

[0919] Otherwise (e.g., when the value of sps_act_reversible_conversion is the first value (e.g., 0)), the residual sample array r of size (nTbW) × (nTbH) can be updated as shown in the following formula. Y r Cb and r Cr .

[0920] [Equation 113]

[0921] tmp=r Y [x][y]-r Cb [x][y])

[0922] r Y [x][y]=r Y [x][y]+r Cb [x][y]

[0923] r Cb [x][y] = tmp - r Cr [x][y]

[0924] r Cr [x][y] = tmp + r Cr [x][y]

[0925] The YCgCo backward transformation and the YCgCo-R backward transformation share some similarities. In a transformation that can be restored to the original state, this can be operated as a lossy backward transformation when Cg and Co are replaced with Cg' = Cg << 1 and Co' = Co << 1. The following equation illustrates its implementation.

[0926] [Equation 114]

[0927] t=Y–(Cg'>>1)=Y–Cg

[0928] G = Cg' + t = Y + Cg

[0929] B=t–(Co'>>1)=t–Co=Y–Cg-Co

[0930] R = Co' + B = t + Co = Y – Cg + Co

[0931] Therefore, in alternative implementations, only a transform that can be restored to the original state can be used, instead of maintaining both color transforms. In the case of lossy encoding, the Cg and Co components can be scaled by a factor of 1 / 2 during the operation of the encoding device and scaled twice during the operation of the decoding device. This makes it possible to use an integrated transform even when both lossy and lossless cases are supported. Additionally, there is the added advantage that the bit length remains unchanged even when lossy encoding is performed.

[0932] [Table 7]

[0933]

[0934] In the implementation method, it can be based on Figure 53 The syntax uses flags to indicate which ACT transformation to use (e.g., actShiftFlag). Figure 53 In the syntax table, the syntax element `sps_act_shift_flag` indicates whether the step of performing color component shifting is applied when applying ACT. For example, a first value of `sps_act_shift_flag` (e.g., 0) indicates that the step of performing color component shifting is not applied when applying ACT. A second value of `sps_act_shift_flag` (e.g., 1) indicates that the step of performing color component shifting is applied when applying ACT. The variable `actShiftFlag` can be set to the value of `sps_act_shift_flag`. Pseudocode for implementing a backward conversion from YCgCo to GBR in a decoding device can be written using `actShiftFlag` as follows.

[0935] [Table 8]

[0936]

[0937] Encoding and Decoding Methods

[0938] The following will describe an image encoding method performed by an image encoding apparatus and an image decoding method performed by an image decoding apparatus according to an embodiment.

[0939] First, the operation of the decoding device will be described. The image decoding device according to an embodiment may include a memory and a processor. The decoding device can perform decoding according to the operation of the processor. Figure 54 A decoding method of a decoding apparatus according to an embodiment is shown. For example... Figure 54 As shown, in step S5410, the decoding device according to the embodiment can determine the residual sample of the current block. Next, in step S5420, the decoding device can reset the value of the residual sample based on whether a color space transformation is applied. Here, the color space transformation can refer to the color space transformation described above. As mentioned above, when the value of the flag indicating whether a color space transformation is applied (e.g., cu_act_enabled_flag) indicates that a color space transformation has been applied (e.g., cu_act_enabled_flag == 1), the decoding device can determine that a color space transformation has been applied.

[0940] Reference Figure 55 The following will describe in more detail the determination of the residual samples of the current block by the decoding apparatus according to the embodiment. In step S5510, the decoding apparatus according to the embodiment can determine the prediction mode of the current block based on prediction information obtained from the bitstream. For example, the decoding apparatus can obtain a prediction mode flag (e.g., pred_mode_flag) from the bitstream that indicates whether the prediction mode of the current coding basis is inter-frame mode or intra-frame mode. When its value is a first value (e.g., pred_mode_flag == 0), the prediction mode of the current block is determined to be inter-frame mode. When the value is a second value (e.g., pred_mode_flag == 1), the prediction mode of the current block is determined to be intra-frame mode.

[0941] Next, in step S5520, the decoding device can obtain a flag indicating whether a color space transformation is applied to the residual samples of the current block based on the prediction mode of the current block. For example, the decoding device can obtain the cu_act_enabled_flag from the bitstream based on whether the prediction mode of the current block is an intra-prediction mode. For example, when the prediction mode of the current block is an intra-prediction mode, the decoding device can obtain the cu_act_enabled_flag from the bitstream. In an implementation, when the prediction mode of the current block is not an intra-prediction mode, the decoding device can obtain the cu_act_enabled_flag from the bitstream only if information about the transformation is obtained for the current coded block. For example, the decoding device can obtain the cu_act_enabled_flag from the bitstream even if the prediction mode of the current block is not an intra-prediction mode, only if the value of the flag (e.g., cu_coded_flag) indicating whether the syntax elements for the transform block are obtained from the bitstream based on the transform tree syntax (e.g., transform_tree()) or transform unit syntax (e.g., transform_unit()) is a predetermined value (e.g., 1). For example, when the syntax elements for the transform unit can be obtained from the bitstream and the prediction mode of the current encoding basis is inter-frame prediction mode or IBC mode, the decoding device can obtain cu_act_enabled_flag from the bitstream.

[0942] Next, as described in the previous embodiments, the decoding device may determine the quantization parameters of the current block based on the value of cu_act_enabled_flag in step S5530. Next, in step S5540, the decoding device may determine the transform coefficients of the current block based on the quantization parameters. Next, in step S5550, the decoding device may determine the residual samples based on the transform coefficients.

[0943] Here, the decoding device can determine the quantization parameters by applying a limiting factor to the quantization parameters so that the values ​​of the quantization parameters are within a predetermined range. Here, the lower limit of the predetermined range can be 0. Alternatively, the upper limit of the predetermined range can be determined based on a syntax element indicating the bit depth of the sample. For example, the upper limit of the predetermined range can be determined as 63 + QpBdOffset. QpBdOffset represents the offset of the range of the luma and chroma quantization parameters and can be preset as a predetermined constant or obtained from the bitstream. For example, QpBdOffset can be calculated by multiplying a predetermined constant (e.g., 6) by the value of a syntax element indicating the bit depth of the luma or chroma sample (e.g., sps_bitdepth).

[0944] Alternatively, the decoding device can determine the quantization parameters based on the color components of the current block, determine the quantization parameter offset based on the color components of the current block, and then reset the quantization parameters using the quantization parameter offset, thereby determining the quantization parameters. For example, the decoding device can add the quantization parameter offset (e.g., ActQpOffset) to the quantization parameter (e.g., qP), thereby resetting the quantization parameters.

[0945] The quantization parameter offset can be determined as follows: When a color space transformation is applied to the residual sample of the current block and the color component of the current block is a luminance component, the quantization parameter offset value can be determined to be -5. Alternatively, when a color space transformation is applied to the residual sample of the current block and the color component of the current block is a chromaticity (Cb) component, the quantization parameter offset value can be determined to be 1. Alternatively, when a color space transformation is applied to the residual sample of the current block and the color component of the current block is a chromaticity (Cr) component, the quantization parameter offset value can be determined to be 3.

[0946] The following will describe in detail step S5420, which involves resetting the value of the residual sample based on whether a color space transformation is applied. The reset of the residual sample value can be performed based on half the value of the chromaticity residual sample. For example, this can be performed as in Equation 109 of Embodiment 9 described above. For example, the half value of the chromaticity residual sample value can be derived by applying a right shift operation to the chromaticity residual sample value. For example, the reset of the residual sample value can be performed based on the half value of the chromaticity residual sample value and the luminance component residual sample value. For example, the reset of the residual sample value can be performed based on the sum of the half value of the chromaticity residual sample value and the luminance component residual sample value. In this embodiment, the reset of the residual sample value can be performed based on the sum of the half values ​​of the luminance component residual sample value and the half values ​​of the chromaticity Cb component residual sample value.

[0947] In one implementation, the luminance component residual sample value can be reset based on the sum of half the chrominance Cb component residual value and the luminance component residual sample value. Here, the half value of the chrominance Cb component residual sample, which is added to the luminance component residual sample value, can be determined based on the value obtained by applying a right shift operation to the chrominance Cb component residual sample value.

[0948] In this implementation, the chromaticity Cb component residual sample value can be reset by subtracting half the value of the chromaticity Cb component residual sample and half the value of the chromaticity Cr component residual sample from the luminance component residual sample value. Here, the half value of the chromaticity Cb component residual sample can be determined based on the value obtained by applying a right shift operation to the chromaticity Cb component residual sample value. Similarly, the half value of the chromaticity Cr component residual sample can be determined based on the value obtained by applying a right shift operation to the chromaticity Cr component residual sample value.

[0949] In one implementation, the chromaticity Cr component residual sample value can be reset by subtracting half the value of the chromaticity Cb component residual sample from the luminance component residual sample value and adding half the value of the chromaticity Cr component residual sample value. Here, the half value of the chromaticity Cb component residual sample can be determined based on the value obtained by applying a right shift operation to the chromaticity Cb component residual sample value. Similarly, the half value of the chromaticity Cr component residual sample can be determined by subtracting the value obtained by applying a right shift operation to the chromaticity Cr component residual sample value.

[0950] The operation of the encoding device will be described below. The image encoding device according to an embodiment may include a memory and a processor. The encoding device can perform encoding according to the operation of the processor in a manner corresponding to decoding by the decoding device. For example, as Figure 56 As shown, the encoding device can determine the residual sample of the current block in step S5610. Next, in step S5620, the encoding device can reset the value of the residual sample based on whether a color space transformation is applied. Here, the value of the residual sample can be reset based on half the value of the chroma residual sample to correspond to the decoding of the decoding device described above.

[0951] Reference Figure 57 The operation of the encoding device will be described in more detail below. In step S5710, the encoding device can determine the prediction mode of the current block. For example, the encoding device can determine the prediction mode of the current block as intra-frame, inter-frame, or IBC mode based on the coding rate. Next, in step S5720, the encoding device can generate a prediction block based on the determined prediction mode, and can generate a residual block based on the difference between the original image block and the prediction block, thereby determining the residual sample. Next, in step S5730, the encoding device can reset the value of the residual sample based on whether a color space transformation is applied. For example, the encoding device can determine whether a color space transformation is applied based on the coding rate. When a color space transformation is applied, the color space transformation of the sample can be performed using the transformation formulas (e.g., Equations 107 or 108) described in Embodiments 9 and 10. For example, in order for the decoding device to perform the inverse transformation, the encoding device can perform the color space transformation according to the formula corresponding to the inverse of Equation 109.

[0952] Next, in step S5740, the encoding device can determine the quantization parameters of the current block based on whether a color space transformation is applied. Corresponding to the operation of the decoding device described above, the encoding device can determine the quantization parameters based on the color components of the current block, determine the quantization parameter offset based on the color components of the current block, and then reset the quantization parameters using the quantization parameter offset, thereby determining the quantization parameters. For example, the encoding device can add the quantization parameter offset to the quantization parameters, thereby resetting the quantization parameters.

[0953] The quantization parameter offset can be determined as follows: When a color space transformation is applied to the residual sample of the current block and the color component of the current block is a luminance component, the quantization parameter offset value can be determined to be -5. Alternatively, when a color space transformation is applied to the residual sample of the current block and the color component of the current block is a chromaticity (Cb) component, the quantization parameter offset value can be determined to be 1. Alternatively, when a color space transformation is applied to the residual sample of the current block and the color component of the current block is a chromaticity (Cr) component, the quantization parameter offset value can be determined to be 3.

[0954] Additionally, the encoding device can determine the quantization parameters by applying a limiting factor to ensure that the values ​​of the quantization parameters fall within a predetermined range. Here, the lower limit of the predetermined range can be 0. Furthermore, the upper limit of the predetermined range can be set to a predetermined upper limit value. This upper limit value can be determined based on the bit depth of the sample. For example, the upper limit of the predetermined range can be determined as 63 + QpBdOffset. Here, QpBdOffset represents the offset of the luminance and chrominance quantization parameter ranges and can be preset as a predetermined constant or obtained from the bitstream. For example, QpBdOffset can be calculated by multiplying the predetermined constant by the value of the syntax element indicating the bit depth of the luminance or chrominance sample.

[0955] Next, in step S5750, the encoding device can determine the transform coefficients of the current block based on the quantization parameters. Next, in step S5760, the encoding device can encode information indicating whether to apply a color space transformation to the residual samples of the current block (e.g., cu_act_enabled_flag) based on the prediction mode of the current block.

[0956] Application and Implementation Methods

[0957] Although the exemplary methods of this disclosure described above are represented as a series of operations for clarity of description, they are not intended to limit the order in which the steps are performed, and these steps may be performed simultaneously or in different orders if necessary. To implement the method according to the invention, the described steps may further include other steps, including steps in addition to some steps, or may include additional steps in addition to some steps.

[0958] In this disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform an operation (step) that confirms the execution conditions or circumstances of the corresponding operation (step). For example, if it is described that a predetermined operation is performed when predetermined conditions are met, the image encoding device or the image decoding device can perform the predetermined operation after determining whether the predetermined conditions are met.

[0959] The various embodiments of this disclosure are not a list of all possible combinations and are intended to describe representative aspects of this disclosure; the matters described in the various embodiments may be applied independently or in combination of two or more.

[0960] Various embodiments of this disclosure can be implemented in hardware, firmware, software, or a combination thereof. When this disclosure is implemented in hardware, it can be implemented using application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, etc.

[0961] Furthermore, the image decoding and image encoding devices applying the embodiments of this disclosure can be included in multimedia broadcasting transmission and reception devices, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, video-on-demand (VoD) service providers, OTT (over-the-top video) devices, internet streaming service providers, three-dimensional (3D) video devices, video telephony devices, medical video devices, etc., and can be used to process video signals or data signals. For example, OTT video devices can include game consoles, Blu-ray players, internet access televisions, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.

[0962] Figure 58 This is a view illustrating a content streaming system to which embodiments of the present disclosure can be applied.

[0963] like Figure 58 As shown, the content streaming system using the embodiments of this disclosure may mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.

[0964] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream and then sends the bitstream to the streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server can be omitted.

[0965] The bitstream can be generated by an image encoding method or image encoding device applying the embodiments of this disclosure, and the streaming server can temporarily store the bitstream during the sending or receiving of the bitstream.

[0966] A streaming server sends multimedia data to a user device based on a user's request through a web server, and the web server acts as a medium for informing the user of the service. When a user requests a service from the web server, the web server can deliver it to the streaming server, and the streaming server can send the multimedia data to the user. In this scenario, the content streaming system may include a separate control server. In this case, the control server is used to control the commands / responses between devices in the content streaming system.

[0967] A streaming server can receive content from media storage devices and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined period of time.

[0968] Examples of user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays), digital televisions, desktop computers, digital signage, etc.

[0969] In a content streaming system, each server can operate as a distributed server, in which case the data received from each server can be distributed.

[0970] The scope of this disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) for enabling the operation of methods according to various embodiments to be executed on a device or computer, and non-transitory computer-readable media having such software or commands stored thereon and executable on a device or computer.

[0971] Industrial applicability

[0972] The embodiments disclosed herein can be used to encode or decode images.

Claims

1. An image decoding method performed by an image decoding device, the image decoding method comprising the following steps: Derive the quantization parameters for each of the luminance component block, chrominance Cb component block, and chrominance Cr component block; The transformation coefficients for each of the luminance component block, the chrominance Cb component block, and the chrominance Cr component block are derived based on the quantization parameters. The residual samples for each of the luminance component block, the chrominance Cb component block, and the chrominance Cr component block are determined based on the transformation coefficients. as well as The value of each residual sample is modified based on color space transformation information; The quantization parameters are derived based on the quantization parameter offset. The value of each residual sample is modified based on the half value of the residual sample for the chromaticity Cb component block.

2. The image decoding method according to claim 1, in, Based on the color space transformation information, for each of the luminance component block, the chromaticity Cb component block, and the chromaticity Cr component block, the value of the quantization parameter offset is derived to be a different value.

3. The image decoding method according to claim 1, in, Based on the color space transformation information, the value of the quantization parameter offset for the luminance component block is derived to be -5.

4. The image decoding method according to claim 1, in, Based on the color space transformation information, the value of the quantization parameter offset for the chromaticity Cb component block is derived to be 1.

5. The image decoding method according to claim 1, in, Based on the color space transformation information, the value of the quantization parameter offset for the chromaticity Cr component block is derived to be 3.

6. The image decoding method according to claim 1, in, The value of each residual sample is also modified based on the value of the residual sample for the luminance component block.

7. The image decoding method according to claim 1, in, The half-value of the residual sample for the chroma Cb component block is obtained based on a right shift operation of the value of the residual sample for the chroma Cb component block.

8. The image decoding method according to claim 1, in, The value of the residual sample for the luminance component block is modified based on the addition of the half value of the residual sample for the chrominance Cb component block.

9. The image decoding method according to claim 8, in, The half-value of the residual sample for the chroma Cb component block, which is added to the value of the residual sample for the luminance component block, is derived by subtracting the value of the residual sample for the chroma Cb component block obtained by a right shift operation from the value of the residual sample for the chroma Cb component block.

10. The image decoding method according to claim 1, in, The value of the residual sample for the chromaticity Cb component block is modified by subtracting both the half value of the residual sample for the chromaticity Cb component block and the half value of the residual sample for the chromaticity Cr component block from the value of the residual sample for the luminance component block.

11. The image decoding method according to claim 10, in, The half-value of the residual sample for the chroma Cb component block is derived based on a value obtained by a right shift operation on the value of the residual sample for the chroma Cb component block. The half-value of the residual sample for the chromaticity Cr component block is derived based on a value obtained by a right shift operation on the value of the residual sample for the chromaticity Cr component block.

12. The image decoding method according to claim 1, in, The value of the residual sample for the chromaticity Cr component block is modified by subtracting half the value of the residual sample for the chromaticity Cb component block from the value of the residual sample for the luminance component block and adding half the value of the residual sample for the chromaticity Cr component block.

13. The image decoding method according to claim 1, in, The residual sample r for the luminance component block Y The residual sample r for the chromaticity Cb component block Cb and the residual sample r for the chromaticity Cr component block Cr The modifications are based on the following: tmp=r Y [x][y]-(r Cb [x][y]>>1) r Y [x][y]=tmp+r Cb [x][y] r Cb [x][y]=tmp-(r Cr [x][y]>>1) r Cr [x][y]=r Cb [x][y]+r Cr [x][y]。 14. An image encoding method performed by an image encoding device, the image encoding method comprising the following steps: Derive the quantization parameters for each of the luminance component block, chrominance Cb component block, and chrominance Cr component block; The transformation coefficients for each of the luminance component block, the chrominance Cb component block, and the chrominance Cr component block are derived based on the quantization parameters. The residual samples for each of the luminance component block, the chrominance Cb component block, and the chrominance Cr component block are determined based on the transformation coefficients. as well as The value of each residual sample is modified based on color space transformation information; The quantization parameters are derived based on the quantization parameter offset. The value of each residual sample is modified based on the half-value of the residual sample for the chromaticity Cb component block.

15. A method for transmitting a bitstream generated by an image encoding method, the image encoding method comprising the following steps: Derive the quantization parameters for each of the luminance component block, chrominance Cb component block, and chrominance Cr component block; The transformation coefficients for each of the luminance component block, the chrominance Cb component block, and the chrominance Cr component block are derived based on the quantization parameters. The residual samples for each of the luminance component block, the chrominance Cb component block, and the chrominance Cr component block are determined based on the transformation coefficients. as well as The value of each residual sample is modified based on color space transformation information; The quantization parameters are derived based on the quantization parameter offset. The value of each residual sample is modified based on the half-value of the residual sample for the chromaticity Cb component block.

Citation Information

Patent Citations

  • Color residual prediction for video coding

    CN105723707A

  • Method and device for encoding / decoding images, and computer readable medium

    CN108712650A