Image encoding / decoding method and apparatus for selectively signaling filter availability information, and method for transmitting a bitstream.

The image encoding/decoding method addresses the high cost of high-resolution image transmission by selectively signaling filter-available information, enhancing encoding/decoding efficiency and reducing storage costs.

JP2026083357APending Publication Date: 2026-05-19LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2026-03-12
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality images leads to higher transmission and storage costs due to the increased amount of information, necessitating a more efficient image compression technique.

Method used

An image encoding/decoding method that selectively signals filter-available information, determining the number of tiles in a current picture and using a first flag to decide on filtering at tile boundaries, with a bitstream generation and transmission method.

Benefits of technology

Improves encoding/decoding efficiency by selectively signaling filter-available information, enabling efficient transmission and storage of high-resolution images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026083357000001_ABST
    Figure 2026083357000001_ABST
Patent Text Reader

Abstract

An image encoding / decoding method and apparatus are provided. [Solution] An image decoding method performed by the image decoding device according to the present disclosure may include the steps of: determining the number of tiles in the current picture based on the fact that the division of the current picture is not restricted; obtaining a first flag from the bitstream indicating whether filtering is available for the tile boundaries based on the fact that there are multiple tiles in the current picture; and determining whether or not to perform filtering on the boundaries of the tiles belonging to the current picture based on the value of the first flag.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly, to an image encoding / decoding method and apparatus for selectively signaling filter-available information, and a method for transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure.

Background Art

[0002] Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has been increasing in various fields. As the image data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to conventional image data. The increase in the amount of information or bits to be transmitted results in an increase in transmission costs and storage costs.

[0003] Therefore, there is a need for a highly efficient image compression technique for effectively transmitting, storing, and reproducing information of high-resolution, high-quality images.

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0005] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus for improving encoding / decoding efficiency by selectively signaling filter-available information.

[0006] Another object of the present disclosure is to provide a method for transmitting a bitstream generated by an image encoding method or apparatus according to the present disclosure.

[0007] Furthermore, this disclosure aims to provide a recording medium that stores a bitstream generated by the image encoding method or apparatus according to this disclosure.

[0008] Furthermore, this disclosure aims to provide a recording medium that stores a bitstream received by the image decoding device provided herein, decoded, and used for image restoration.

[0009] The technical problems that this disclosure seeks to solve are not limited to those described above, and other technical problems not mentioned above will be clearly understood by a person with ordinary skill in the art to which this disclosure pertains from the following description. [Means for solving the problem]

[0010] An image decoding method performed by an image decoding apparatus according to one aspect of the present disclosure may include the steps of: determining the number of tiles in a current picture based on the fact that the division of the current picture is not restricted; obtaining a first flag from a bitstream indicating whether filtering is available for the boundaries of the tiles based on the fact that there are multiple tiles in the current picture; and determining whether or not to perform filtering on the boundaries of the tiles belonging to the current picture based on the value of the first flag.

[0011] Furthermore, an image decoding apparatus according to one aspect of the present disclosure is an image decoding apparatus comprising a memory and at least one processor, wherein the at least one processor can determine the number of tiles in the current picture based on the fact that the division of the current picture is not restricted, obtain a first flag from the bitstream indicating whether filtering is available for the tile boundaries based on the fact that there are multiple tiles in the current picture, and determine whether or not to perform filtering on the boundaries of the tiles belonging to the current picture based on the value of the first flag.

[0012] Furthermore, an image encoding method performed by an image encoding apparatus according to one aspect of the present disclosure may include the steps of: determining the number of tiles in a current picture based on the fact that the division of the current picture is not restricted; determining the value of a first flag indicating whether filtering is available for the tile boundaries based on the fact that there are multiple tiles in the current picture; and generating a bitstream including the first flag.

[0013] Furthermore, a transmission method according to one aspect of the present disclosure can transmit a bitstream generated by an image encoding device or image encoding method of the present disclosure.

[0014] Furthermore, a computer-readable recording medium according to one aspect of the present disclosure can store a bitstream generated by an image encoding method or image encoding apparatus of the present disclosure.

[0015] The features described above, which are a brief summary of this disclosure, are merely illustrative examples of the detailed description of this disclosure described below and do not limit the scope of this disclosure. [Effects of the Invention]

[0016] According to this disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.

[0017] Furthermore, according to this disclosure, an image coding / decoding method and apparatus can be provided that can improve coding / decoding efficiency by selectively signaling filter-available information.

[0018] Furthermore, this disclosure provides a method for transmitting a bitstream generated by an image encoding method or apparatus according to this disclosure.

[0019] Furthermore, according to this disclosure, a recording medium storing a bitstream generated by the image encoding method or apparatus according to this disclosure can be provided.

[0020] In addition, according to the present disclosure, a recording medium storing a bitstream received by the image decoding device according to the present disclosure, decoded, and used for restoring an image can be provided.

[0021] The effects obtained by the present disclosure are not limited to the above-described effects, and other effects not described above will be clearly understood by those having ordinary knowledge in the technical field to which the present disclosure pertains from the following description.

Brief Description of Drawings

[0022] [Figure 1] It is a diagram schematically showing a video coding system to which an embodiment according to the present disclosure can be applied. [[ID=*]] [[ID=*]] [Figure 2] It is a diagram schematically showing an image encoding device to which an embodiment according to the present disclosure can be applied. [Figure 3] It is a diagram schematically showing an image decoding device to which an embodiment according to the present disclosure can be applied. [Figure 4] It is a diagram showing a division structure of an image according to an embodiment. [Figure 5] It is a diagram showing an embodiment of a division type of a block by a multi-type tree structure. [Figure 6] It is a diagram exemplifying a signaling mechanism of block division information in a quadtree with nested multi-type tree structure with a multi-type tree according to the present disclosure. [Figure 7] It is a diagram showing an embodiment in which a CTU is divided into multiple CUs. [Figure 8] It is a diagram showing peripheral reference samples according to an embodiment. [Figure 9] It is a diagram for explaining intra prediction according to an embodiment. [Figure 10] It is a diagram for explaining intra prediction according to an embodiment. [Figure 11] It is a diagram for explaining an encoding method using inter prediction according to an embodiment. [Figure 12]This figure illustrates a decoding method using interpretation according to one embodiment. [Figure 13] This block diagram shows a CABAC representation of one example for encoding a single syntax element. [Figure 14] This figure illustrates entropy coding and decoding according to one embodiment. [Figure 15] This figure illustrates entropy coding and decoding according to one embodiment. [Figure 16] This figure illustrates entropy coding and decoding according to one embodiment. [Figure 17] This figure illustrates entropy coding and decoding according to one embodiment. [Figure 18] This figure shows an example of a picture decoding and encoding procedure according to one embodiment. [Figure 19] This figure shows an example of a picture decoding and encoding procedure according to one embodiment. [Figure 20] This figure shows the hierarchical structure of an image coded according to one embodiment. [Figure 21] This figure shows an example of dividing a picture using tiles, slices, and subpictures. [Figure 22] This figure shows an example of dividing a picture using tiles, slices, and subpictures. [Figure 23] This figure shows an example of dividing a picture using tiles, slices, and subpictures. [Figure 24] This figure shows an example of dividing a picture using tiles, slices, and subpictures. [Figure 25] This figure shows individual examples of syntax for picture parameter sets. [Figure 26] This figure shows individual examples of syntax for picture parameter sets. [Figure 27] This figure shows individual examples of syntax for picture parameter sets. [Figure 28]This figure shows individual examples of syntax for picture parameter sets. [Figure 29] This figure shows one embodiment of a decoding method and an encoding method. [Figure 30] This figure shows one embodiment of a decoding method and an encoding method. [Figure 31] This figure illustrates a content streaming system to which the embodiments of this disclosure can be applied. [Modes for carrying out the invention]

[0023] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings, so that they can be easily implemented by a person with ordinary skill in the art to which the present disclosure pertains. However, the present disclosure can be implemented in a variety of different forms and is not limited to the embodiments described herein.

[0024] In describing embodiments of this disclosure, if it is determined that a specific description of a known configuration or function would obscure the gist of this disclosure, such detailed description will be omitted. In the drawings, parts unrelated to the description of this disclosure will be omitted, and similar parts will be denoted by the same reference numerals.

[0025] In this disclosure, when one component is described as being “connected,” “joined,” or “linked” to another component, this can include not only direct connections but also indirect connections where another component exists between them. Furthermore, when one component is described as “containing” or “having” another component, this means, unless otherwise stated to the contrary, that it may include another component rather than excluding it.

[0026] In this disclosure, terms such as "first," "second," etc., are used solely for the purpose of distinguishing one component from another, and do not limit the order or importance of the components unless otherwise specified. Therefore, within the scope of this disclosure, the first component of one embodiment may be called the second component in another embodiment, and similarly, the second component of one embodiment may be called the first component in another embodiment.

[0027] In this disclosure, components that are distinguished from each other are used to clearly describe their respective characteristics and do not necessarily mean that the components are separate. In other words, multiple components may be integrated to constitute a single hardware or software unit, or a single component may be distributed to constitute multiple hardware or software units. Therefore, such integrated or distributed embodiments are also included in the scope of this disclosure, without needing to be specifically mentioned.

[0028] In this disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, embodiments consisting of a subset of the components described in one embodiment are also included in the scope of this disclosure. Furthermore, embodiments that include additional components in addition to the components described in various embodiments are also included in the scope of this disclosure.

[0029] This disclosure relates to the encoding and decoding of images, and the terms used in this disclosure may have their ordinary meanings in the art to which this disclosure pertains, unless otherwise defined herein.

[0030] In this disclosure, "video" can mean a collection of images over time. "Picture" generally means a unit representing any one image at a particular time point in time, and a slice / tile is an encoding unit that constitutes part of a picture in encoding. A single picture can consist of one or more slices / tiles. A slice / tile can also contain one or more CTUs (coding tree units). The CTU can be divided into one or more CUs. A single picture can consist of one or more slices / tiles. A tile is a specific tile row and a specific tile column within a picture. A rectangular region within a Column (CTU) can consist of multiple CTUs. A tile row can be defined as a rectangular region of CTUs, having the same height as the picture and a width specified by syntax elements signaled from a bitstream portion such as the picture parameter set. A tile row can be defined as a rectangular region of CTUs, having the same width as the picture and a height specified by syntax elements signaled from a bitstream portion such as the picture parameter set. A tile scan is a predetermined sequential ordering method of CTUs that divide a picture. Here, the CTUs may be sequentially ordered within a tile according to a CTU raster scan, and the tiles within a picture may be sequentially ordered according to the raster scan order of the tiles in the picture. A slice can contain an integer number of complete tiles or an integer number of consecutive complete CTU rows within a single tile of a picture. A slice can exclusively be contained within a single NAL unit.

[0031] A picture can consist of one or more tile groups. A tile group can contain one or more tiles. A brick can represent a rectangular area of ​​a CTU row within a tile in a picture. A tile can contain one or more bricks. A brick can represent a rectangular area of ​​a CTU row within a tile. A tile can be divided into multiple bricks, and each brick can contain one or more CTU rows belonging to the tile. A tile that is not divided into multiple bricks can also be treated as a brick.

[0032] On the other hand, a single picture can be divided into two or more subpictures. A subpicture can be a rectangular region of one or more slices within the picture.

[0033] In this disclosure, “pixel” or “pel” may mean the smallest unit that constitutes a picture (or image). The term “sample” may also be used as a counterpart to pixel. A sample may generally represent a pixel or a pixel value, or it may represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.

[0034] In this disclosure, “unit” can refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information associated with that region. A unit may include one luma block and two chroma (e.g., Cb, Cr) blocks. The term “unit” may be used interchangeably with terms such as “sample array,” “block,” or “area,” as it may be used. Generally, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.

[0035] In this disclosure, “current block” can mean any one of the following: “current coding block,” “current coding unit,” “block to encode,” “block to decode,” or “block to process.” If prediction is performed, “current block” can mean “current prediction block” or “block to predict.” If transformation (inverse transformation) / quantization (inverse quantization) is performed, “current block” can mean “current transformation block” or “block to transform.” If filtering is performed, “current block” can mean “block to filter.”

[0036] Furthermore, in this disclosure, “current block” may mean “chroma block of the current block” unless there is an explicit mention of chroma block. “Chroma block of the current block” may be expressed explicitly as “chroma block” or “current chroma block,” including an explicit mention of chroma block.

[0037] In this disclosure, " / " and "," may be interpreted as "and / or." For example, "A / B" and "A, B" may be interpreted as "A and / or B." Also, "A / B / C" and "A, B, C" may mean "at least one of A, B and / or C."

[0038] In this disclosure, “or” may be interpreted as “and / or.” For example, “A or B” may mean 1) “A” only, 2) “B” only, or 3) “A and B.” Alternatively, in this disclosure, “or” may mean “additionally or alternatively.”

[0039] Overview of the video coding system

[0040] Figure 1 shows the video coding system according to this disclosure.

[0041] A video coding system according to one embodiment may include a source device 10 and a receiving device 20. The source device 10 can transmit encoded video and / or image information or data to the receiving device 20 via a digital storage medium or network in file or streaming format.

[0042] A source device 10 according to one embodiment may include a video source generation unit 11, an encoding device 12, and a transmission unit 13. A receiving device 20 according to one embodiment may include a receiving unit 21, a decoding device 22, and a rendering unit 23. The encoding device 12 may be called a video / image encoding device, and the decoding device 22 may be called a video / image decoding device. The transmission unit 13 may be included in the encoding device 12. The receiving unit 21 may be included in the decoding device 22. The rendering unit 23 may also include a display unit, which may be configured as a separate device or external component.

[0043] The video source generation unit 11 can acquire video / images through processes such as video / image capture, synthesis, or generation. The video source generation unit 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, or a video / image archive containing previously captured video / images. The video / image generation device may include, for example, a computer, tablet, and smartphone, and may generate video / images (electronically). For example, virtual video / images may be generated via a computer, in which case the video / image capture process may be replaced by a process in which the relevant data is generated.

[0044] The encoding device 12 can encode input video / images. The encoding device 12 can perform a series of steps, such as prediction, transformation, and quantization, for compression and encoding efficiency. The encoding device 12 can output the encoded data (encoded video / image information) in bitstream format.

[0045] The transmission unit 13 can transmit encoded video / image information or data, output in bitstream format, to the receiving unit 21 of the receiving device 20 via a digital storage medium or network in file or streaming format. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray®, HDD, and SSD. The transmission unit 13 may include elements for generating media files via a predetermined file format and elements for transmission via a broadcast / communication network. The receiving unit 21 can extract / receive the bitstream from the storage medium or network and transmit it to the decoding device 22.

[0046] The decoding device 22 can decode video / images by performing a series of steps such as inverse quantization, inverse transform, and prediction, corresponding to the operation of the encoding device 12.

[0047] The rendering unit 23 can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0048] Overview of Image Encoding Devices

[0049] Figure 2 is a schematic diagram showing an image encoding device to which the embodiments of this disclosure can be applied.

[0050] As shown in Figure 2, the image coding device 100 may include an image splitting unit 110, a subtraction unit 115, a transformation unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transformation unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter-prediction unit 180, an intra-prediction unit 185, and an entropy coding unit 190. The inter-prediction unit 180 and the intra-prediction unit 185 can together be called the "prediction unit". The transformation unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse transformation unit 150 may be included in a residual processing unit. The residual processing unit may further include a subtraction unit 115.

[0051] All or at least some of the multiple components constituting the image encoding device 100 can be implemented by a single hardware component (e.g., an encoder or processor) depending on the embodiment. Furthermore, the memory 170 may include a DPB (decoded picture buffer) and can be implemented by a digital storage medium.

[0052] The image splitting unit 110 can split an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. For example, the processing units may be called coding units (CUs). Coding units can be obtained by recursively splitting a coding tree unit (CTU) or the largest coding unit (LCU) using a QT / BT / TT (Quad-tree / binary-tree / ternary-tree) structure. For example, a single coding unit can be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure and / or a ternary-tree structure. For the splitting of coding units, a quad-tree structure may be applied first, followed by a binary-tree structure and / or a ternary-tree structure. Based on the final coding unit that cannot be further split, the coding procedure according to this disclosure can be performed. The largest coding unit can be used as the final coding unit, or a lower-depth coding unit obtained by dividing the largest coding unit can be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and / or restoration, as described later. As another example, the processing units of the coding procedure may be prediction units (PU) or transformation units (TU). The prediction unit and the transformation unit may be divided or partitioned from the final coding unit, respectively. The prediction unit may be a unit of sample prediction, and the transformation unit may be a unit that derives transformation coefficients and / or a unit that derives a residual signal from transformation coefficients.

[0053] The prediction unit (inter-prediction unit 180 or intra-prediction unit 185) can make predictions for the block to be processed (current block) and generate a predicted block that includes prediction samples for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block or on a CU basis. The prediction unit can generate various information regarding the prediction of the current block and transmit it to the entropy coding unit 190. The prediction information can be encoded by the entropy coding unit 190 and output in bitstream format.

[0054] The intra-prediction unit 185 can predict the current block by referring to a sample in the current picture. The referenced sample may be located in the vicinity (neighbor) or at a distance from the current block, according to the intra-prediction mode and / or intra-prediction technique. The intra-prediction mode may include multiple non-directional modes and multiple directional modes. The non-directional modes may include, for example, a DC mode and a Planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 185 may also determine the prediction mode to be applied to the current block using the prediction modes applied to the surrounding blocks.

[0055] The interprediction unit 180 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between the surrounding blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, the surrounding blocks may include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different from each other. The temporal neighboring block may be called a collocated reference block, colCU, etc. The reference picture containing the temporal neighboring block may be called a collocated picture (colPic). For example, the interpretation unit 180 can construct a motion information candidate list based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Interpretation can be performed based on various prediction modes; for example, in skip mode and merge mode, the interpretation unit 180 can use the motion information of surrounding blocks as the motion information of the current block. In skip mode, unlike merge mode, the residual signal may not be transmitted.In motion vector prediction (MVP) mode, the motion vector of the surrounding block is used as the motion vector predictor, and the motion vector of the current block can be signaled by encoding the motion vector difference and an indicator for the motion vector predictor. The motion vector difference can represent the difference between the motion vector of the current block and the motion vector predictor.

[0056] The prediction unit can generate a prediction signal based on various prediction methods and / or techniques described later. For example, the prediction unit can apply intra-prediction or inter-prediction to predict the current block, and can also apply intra-prediction and inter-prediction simultaneously. A prediction method that applies intra-prediction and inter-prediction simultaneously to predict the current block can be called CIIP (combined inter and intra prediction). The prediction unit can also perform intra-block copy (IBC) to predict the current block. Intra-block copy can be used for content image / video coding such as in games, for example, in SCC (screen content coding). IBC is a method of predicting the current block using a reference block that has already been restored in the current picture at a predetermined distance from the current block. When IBC is applied, the position of the reference block in the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but can be performed similarly to inter-prediction in that it derives the reference block within the current picture. In other words, IBC can use at least one of the interpretation techniques described in this disclosure.

[0057] The predicted signal generated by the prediction unit can be used to generate a reconstructed signal or a residual signal. The subtraction unit 115 can generate a residual signal (residual block, residual sample array) by subtracting the predicted signal output from the prediction unit (predicted block, predicted sample array) from the input image signal (original block, original sample array). The generated residual signal can be transmitted to the conversion unit 120.

[0058] The transformation unit 120 can generate transformation coefficients by applying transformation techniques to the residual signal. For example, the transformation techniques may include at least one of the following: DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented by this graph. CNT refers to a transformation obtained by generating a prediction signal using all previously reconstructed pixels. The transformation process can be applied to pixel blocks of the same size and are square, or to non-square, variable-sized blocks.

[0059] The quantization unit 130 can quantize the conversion coefficients and transmit them to the entropy coding unit 190. The entropy coding unit 190 can encode the quantized signal (information about the quantized conversion coefficients) and output it in bitstream format. The information about the quantized conversion coefficients can be called residual information. The quantization unit 130 can rearrange the block-form quantized conversion coefficients into a one-dimensional vector format based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector format of the quantized conversion coefficients.

[0060] The entropy coding unit 190 can perform various coding methods, such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). The entropy encoding unit 190 can encode, together or separately, information necessary for video / image restoration (e.g., syntax elements (values ​​of syntax elements), in addition to the quantized conversion coefficients). The encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream format in units of network abstraction layer (NAL) units. The video / image information may further include information about various parameter sets, such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). The video / image information may also further include general constraint information. The signaling information, transmitted information, and / or syntax elements referred to in this disclosure can be encoded via the encoding procedure described above and included in the bitstream.

[0061] The bitstream can be transmitted over a network or stored on a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmission unit (not shown) for transmitting the signal output from the entropy encoding unit 190 and / or a storage unit (not shown) for storing it may be provided as an internal / external element of the image encoding device 100, or the transmission unit may be provided as a component of the entropy encoding unit 190.

[0062] The quantized conversion coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, by applying inverse quantization and inverse transformation to the quantized conversion coefficients via the inverse quantization unit 140 and the inverse transformation unit 150, a residual signal (residual block or residual sample) can be reconstructed.

[0063] The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-prediction unit 180 or the intra-prediction unit 185. If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The adder 155 may be called the reconstruction unit or the reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, or, as described later, for inter-prediction of the next picture after filtering.

[0064] The filtering unit 160 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 170, specifically in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter. The filtering unit 160 can generate various filtering-related information, as will be described later in the explanation of each filtering method, and transmit it to the entropy coding unit 190. The filtering-related information can be encoded by the entropy coding unit 190 and output in bitstream format.

[0065] The corrected restored picture transmitted to memory 170 can be used as a reference picture in the interpretation unit 180. When interpretation is applied via this, the image encoding device 100 can avoid prediction mismatches between the image encoding device 100 and the image decoding device, and can also improve encoding efficiency.

[0066] The DPB in memory 170 can store the modified restored picture for use as a reference picture in the inter-prediction unit 180. Memory 170 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 180 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 170 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 185.

[0067] Overview of the image decoding device

[0068] Figure 3 is a schematic diagram showing an image decoding apparatus to which the embodiments of this disclosure can be applied.

[0069] As shown in Figure 3, the image decoding device 200 can be configured to include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an additive unit 235, a filtering unit 240, a memory 250, an inter-prediction unit 260, and an intra-prediction unit 265. The inter-prediction unit 260 and the intra-prediction unit 265 can together be called the "prediction unit". The inverse quantization unit 220 and the inverse transform unit 230 can be included in the residual processing unit.

[0070] All or at least some of the multiple components constituting the image decoding device 200 can be implemented by a single hardware component (e.g., a decoder or processor) according to the embodiment. Furthermore, the memory 170 may include a DPB and can be implemented by a digital storage medium.

[0071] An image decoding device 200, upon receiving a bitstream containing video / image information, can restore the image by executing a process corresponding to the process performed in the image encoding device 100 in Figure 1. For example, the image decoding device 200 can perform decoding using the processing unit applied in the image encoding device. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. The restored image signal decoded and output via the image decoding device 200 can then be reproduced via a playback device (not shown).

[0072] The image decoding device 200 can receive the signal output from the image encoding device 2 in bitstream format. The received signal can be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bitstream to derive information necessary for image restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as adaptive parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The image decoding device may further use the parameter set information and / or the general constraint information to decode the image. The signaling information, received information, and / or syntax elements referred to in this disclosure can be obtained from the bitstream by decoding via the decoding procedure. For example, the entropy decoding unit 210 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values ​​of syntax elements necessary for image reconstruction and the quantized values ​​of conversion coefficients related to the residual. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element from the bitstream, determines a context model using the syntax element information to be decoded, the decoding information of the surrounding blocks and the blocks to be decoded, or the symbol / bin information decoded in a previous step, predicts the probability of bin occurrence based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values ​​of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin.Of the information decoded by the entropy decoding unit 210, information related to prediction is provided to the prediction unit (inter-prediction unit 260 and intra-prediction unit 265), and the residual values ​​that have undergone entropy decoding in the entropy decoding unit 210, i.e., quantized conversion coefficients and related parameter information, can be input to the inverse quantization unit 220. In addition, of the information decoded by the entropy decoding unit 210, information related to filtering can be provided to the filtering unit 240. On the other hand, a receiving unit (not shown) that receives signals output from the image coding device may be further provided as an internal / external element of the image decoding device 200, or the receiving unit may be provided as a component of the entropy decoding unit 210.

[0073] On the other hand, the image decoding device according to this disclosure may be called a video / image / picture decoding device. The image decoding device may also include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoding unit 210, and the sample decoder may include at least one of an inverse quantization unit 220, an inverse transform unit 230, an adder unit 235, a filtering unit 240, a memory 250, an inter-prediction unit 260, and an intra-prediction unit 265.

[0074] The inverse quantization unit 220 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 220 can rearrange the quantized transformation coefficients in a two-dimensional block format. In this case, the rearrangement can be performed based on the coefficient scan order performed by the image encoding device. The inverse quantization unit 220 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) to obtain the transformation coefficients.

[0075] The inverse conversion unit 230 can inversely convert the conversion coefficients to obtain residual signals (residual blocks, residual sample arrays).

[0076] The prediction unit can make predictions for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 210, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block and can determine a specific intra / inter-prediction mode (prediction technique).

[0077] As described in the explanation of the prediction unit of the image coding device 100, the prediction unit can generate prediction signals based on various prediction methods (techniques) described later.

[0078] The intra-prediction unit 265 can predict the current block by referring to the samples in the current picture. The description of the intra-prediction unit 185 can also be applied to the intra-prediction unit 265.

[0079] The interprediction unit 260 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on a reference picture. In this case, to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in block, sub-block, or sample units based on the correlation of motion information between surrounding blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In interprediction, surrounding blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the interprediction unit 260 can construct a motion information candidate list based on surrounding blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction can be performed based on various prediction modes (techniques), and the prediction information may include information indicating the mode (technique) of interprediction for the current block.

[0080] The adder 235 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter-prediction unit 260 and / or intra-prediction unit 265). If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The description of the adder 155 can also be applied to the adder 235. The adder 235 can be called the reconstruction unit or reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, or, as described later, for inter-prediction of the next picture after filtering.

[0081] The filtering unit 240 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 240 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 250, specifically in the DPB of the memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.

[0082] The restored picture stored (modified) in the DPB of memory 250 can be used as a reference picture in the inter-prediction unit 260. Memory 250 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 260 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 250 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 265.

[0083] In this specification, the embodiments described for the filtering unit 160, inter-prediction unit 180, and intra-prediction unit 185 of the image coding device 100 can be applied similarly or in a corresponding manner to the filtering unit 240, inter-prediction unit 260, and intra-prediction unit 265 of the image decoding device 200, respectively.

[0084] Overview of image segmentation

[0085] The video / image coding method according to this disclosure can be performed based on the following image segmentation structure. Specifically, procedures such as prediction, residual processing (inverse transformation, inverse quantization, etc.), syntax element coding, and filtering, described later, can be performed based on CTU, CU (and / or TU, PU) derived from the image segmentation structure. The image can be segmented into blocks, and the block segmentation procedure can be performed in the image segmentation unit 110 of the encoding device described above. Segmentation-related information can be encoded in the entropy encoding unit 190 and transmitted to the decoding device in bitstream format. The entropy decoding unit 210 of the decoding device can derive the block segmentation structure of the current picture based on the segmentation-related information obtained from the bitstream, and perform a series of procedures for image decoding (e.g., prediction, residual processing, block / picture reconstruction, in-loop filtering, etc.) based on this.

[0086] A picture can be divided into a sequence of coding tree units (CTUs). Figure 4 shows an example of a picture being divided into CTUs. A CTU can correspond to coding tree blocks (CTBs). Alternatively, a CTU can contain two coding tree blocks: one for a luma sample and one for a corresponding chroma sample. For example, for a picture containing three sample arrays, the CTU can contain an N×N block of luma samples and two corresponding blocks for chroma samples. The maximum allowable size of a CTU for coding and prediction can differ from the maximum allowable size of a CTU for transformation. For example, the maximum allowable size of a luma block within a CTU can be 128×128, even if the maximum size of a luma transformation block is 64×64.

[0087] Overview of CTU division

[0088] As mentioned above, coding units can be obtained by recursively partitioning a coding tree unit (CTU) or a maximum coding unit (LCU) using a QT / BT / TT (Quad-tree / binary-tree / ternary-tree) structure. For example, a CTU can first be partitioned into a quadtree structure. Then, the leaf nodes of the quadtree structure can be further partitioned into a multi-type tree structure.

[0089] A quadtree partition means dividing the current CU (or CTU) into four equal parts. Through a quadtree partition, the current CU can be divided into four CUs of the same width and height. If the current CU is not further divided into a quadtree structure, it corresponds to a leaf node in the quadtree structure. A CU that corresponds to a leaf node in a quadtree structure is not further divided and can be used as the final coding unit as described above. Alternatively, a CU that corresponds to a leaf node in a quadtree structure can be further divided by a multi-type tree structure.

[0090] Figure 5 shows the types of block partitioning using a multi-type tree structure. Partitioning using a multi-type tree structure can include two partitions using a binary tree structure and two partitions using a ternary tree structure.

[0091] The two types of partitioning using a binary tree structure include vertical binary splitting (SPLIT_BT_VER) and horizontal binary splitting (SPLIT_BT_HOR). Vertical binary splitting (SPLIT_BT_VER) means splitting the current CU vertically into two equal parts. As shown in Figure 4, vertical binary splitting can generate two CUs that have the same height as the current CU and half the width of the current CU. Horizontal binary splitting (SPLIT_BT_HOR) means splitting the current CU horizontally into two equal parts. As shown in Figure 5, horizontal binary splitting can generate two CUs that have half the height of the current CU and the same width as the current CU.

[0092] Two types of partitioning using a ternary structure are vertical ternary splitting (SPLIT_TT_VER) and horizontal ternary splitting (SPLIT_TT_HOR). Vertical ternary splitting (SPLIT_TT_VER) divides the current CU vertically in a 1:2:1 ratio. As shown in Figure 5, vertical ternary splitting can produce two CUs with the same height as the current CU and a width of 1 / 4 of the current CU's width, and one CU with the same height as the current CU and a width of half the current CU's width. Horizontal ternary splitting (SPLIT_TT_HOR) divides the current CU horizontally in a 1:2:1 ratio. As shown in Figure 4, horizontal ternary splitting can produce two CUs with the same height as the current CU and a width of 1 / 4 of the current CU's width, and one CU with the same height as the current CU and a width of half the current CU's width.

[0093] Figure 6 illustrates the signaling mechanism for block partitioning information in a quadtree with nested multi-type tree structure according to this disclosure.

[0094] Here, the CTU is treated as the root node of the quadtree, and the CTU is the first node to be split into a quadtree structure. Information (e.g., qt_split_flag) indicating whether or not to split the quadtree can be signaled to the current CU (CTU or quadtree node (QT_node)). For example, if qt_split_flag is the first value (e.g., "1"), the current CU can be split into a quadtree. If qt_split_flag is the second value (e.g., "0"), the current CU will not be split into a quadtree and will become a leaf node (QT_leaf_node) of the quadtree. Each leaf node of the quadtree can subsequently be further split into a multitype tree structure. In other words, a leaf node of a quadtree can become a node (MTT_node) of a multitype tree. In a multi-type tree structure, a first flag (e.g., mtt_split_cu_flag) can be signaled to indicate whether the current node will be further split. If the node is to be further split (e.g., the first flag is 1), a second flag (e.g., mtt_split_cu_verticla_flag) can be signaled to indicate the splitting direction. For example, if the second flag is 1, the splitting direction is vertical, and if the second flag is 0, the splitting direction is horizontal. Subsequently, a third flag (e.g., mtt_split_cu_binary_flag) can be signaled to indicate whether the splitting type is binary or ternary. For example, if the third flag is 1, the splitting type is binary, and if the third flag is 0, the splitting type is ternary. Nodes in a multitype tree obtained by binary partitioning or ternary partitioning can be further partitioned into a multitype tree structure. However, nodes in a multitype tree cannot be partitioned into a quadtree structure.If the first flag is 0, the corresponding node in the multitype tree is not further subdivided and becomes a leaf node (MTT_leaf_node) of the multitype tree. A CU corresponding to a leaf node in the multitype tree can be used as the final coding unit as described above.

[0095] Based on the aforementioned mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the multi-type tree splitting mode (MttSplitMode) of the CU can be derived as shown in Table 1. In the following description, the multi-tree splitting mode may be abbreviated as the multi-tree splitting type or splitting type.

[0096] [Table 1]

[0097] Figure 7 shows an example where a CTU is divided into multiple CUs by applying a multitype tree after a quadtree. In Figure 7, the bold block edge 710 represents the quadtree division, and the remaining edge 720 represents the multitype tree division. A CU can correspond to a coding block CB. In one embodiment, a CU may include two coding blocks: a coding block for a luma sample and a coding block for a chroma sample corresponding to the luma sample. The chroma component (sample) CB or TB size can be derived based on the luma component (sample) CB or TB size according to the component ratio of the picture / image color format (chroma format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.). If the color format is 4:4:4, the chroma component CB / TB size can be set to be the same as the luma component CB / TB size. If the color format is 4:2:2, the width of the chroma component CB / TB can be set to half the width of the luma component CB / TB, and the height of the chroma component CB / TB can be set to the height of the luma component CB / TB. If the color format is 4:2:0, the width of the chroma component CB / TB can be set to half the width of the luma component CB / TB, and the height of the chroma component CB / TB can be set to half the height of the luma component CB / TB.

[0098] In one embodiment, when the size of the CTU is 128 based on the luma sample unit, the size of the CU can range from 128 × 128, which is the same size as the CTU, to 4 × 4. In one embodiment, when the color format is 4:2:0 (or chroma format), the chroma CB size can range from 64 × 64 to 2 × 2.

[0099] On the other hand, in one embodiment, the CU size and TU size can be the same. Alternatively, multiple TUs can exist within the CU region. The TU size generally refers to the Luma component (sample) TB (Transform Block) size.

[0100] The TU size can be derived based on a preset value, the maximum allowable TB size (maxTbSize). For example, if the CU size is larger than the maxTbSize, multiple TUs (TBs) with the maxTbSize can be derived from the CU, and conversion / inverse conversion can be performed in units of the TUs (TBs). For example, the maximum allowable lumen TB size may be 64×64, and the maximum allowable chromen TB size may be 32×32. If the width or height of a CB divided by the tree structure is larger than the maximum conversion width or height, the CB can be automatically (or implicitly) divided until the horizontal and vertical TB size limits are satisfied.

[0101] Furthermore, for example, when intra-prediction is applied, the intra-prediction mode / type is derived on a CU (or CB) basis, and the peripheral reference sample derivation and prediction sample generation procedures can be performed on a TU (or TB) basis. In this case, one or more TUs (or TBs) can exist within a single CU (or CB) region, and in this case, the multiple TUs (or TBs) can share the same intra-prediction mode / type.

[0102] On the other hand, for a quadtree coding tree scheme with multitype trees, the following parameters can be signaled from the encoder to the decoder as SPS syntax elements. For example, at least one of the following can be signaled: CTUsize, which indicates the size of the root node of the quadtree; MinQTSize, which indicates the minimum allowed size of the leaf nodes of the quadtree; MaxBTSize, which indicates the maximum allowed size of the root node of the binary tree; MaxTTSize, which indicates the maximum allowed size of the root node of the ternary tree; MaxMttDepth, which indicates the maximum allowed hierarchy depth of the multitype trees that are split from the leaf nodes of the quadtree; MinBtSize, which indicates the minimum allowed leaf node size of the binary tree; and MinTtSize, which indicates the minimum allowed leaf node size of the ternary tree.

[0103] In one embodiment using the 4:2:0 chroma format, the CTU size can be set to a 128x128 chroma block and two corresponding 64x64 chroma blocks. In this case, MinQTSize can be set to 16x16, MaxBtSize to 128x128, MaxTtSzie to 64x64, MinBtSize and MinTtSize to 4x4, and MaxMttDepth to 4. A quadtree partition can be applied to the CTU to generate leaf nodes of the quadtree. Leaf nodes of a quadtree can be called leaf QT nodes. Leaf nodes of a quadtree can range in size from 16x16 (e.g., the MinQTSize) to 128x128 (e.g., the CTU size). If a leaf QT node is 128x128, it may not be further partitioned into a binary / ternary tree. This is because even if partitioned in this case, it would exceed MaxBtsize and MaxTtszie (e.g., 64x64). Otherwise, a leaf QT node can be further partitioned into a multitype tree. Thus, a leaf QT node is the root node for a multitype tree, and a leaf QT node can have a multitype tree depth (mttDepth) value of 0. If the multitype tree depth reaches MaxMttdepth (e.g., 4), further additional partitioning may not be considered. If the width of a multitype tree node is the same as MinBtSize and equal to or less than 2xMinTtSize, further additional horizontal partitioning may not be considered. If the height of a multitype tree node is the same as MinBtSize and equal to or less than 2xMinTtSize, further additional vertical partitioning may not be considered. When partitioning is not considered in this way, the encoding device can omit signaling of partitioning information. In such cases, the decoding device can induce the partitioning information to a predetermined value.

[0104] On the other hand, a single CTU can include a coding block for a luma sample (hereinafter referred to as a "luma block") and two coding blocks for corresponding chroma samples (hereinafter referred to as "chroma blocks"). The coding tree scheme described above can be applied similarly to the luma blocks and chroma blocks of a CU, or it can be applied separately. Specifically, luma blocks and chroma blocks within a single CTU can be divided into the same block tree structure, in which case the tree structure can be represented as a single tree (SINGLE_TREE). Alternatively, luma blocks and chroma blocks within a single CTU can be divided into separate block tree structures, in which case the tree structure can be represented as a dual tree (DUAL_TREE). In other words, when a CTU is divided into a dual tree, the block tree structure for luma blocks and the block tree structure for chroma blocks can exist separately. In this case, the block tree structure for a luma block can be called a dual-tree luma (DUAL_TREE_LUMA), and the block tree structure for a chroma block can be called a dual-tree chroma (DUAL_TREE_CHROMA). For P and B slice / tile groups, luma blocks and chroma blocks within a single CTU can be restricted to having the same coding tree structure. However, for I slice / tile groups, luma blocks and chroma blocks can have separate block tree structures from each other. If separate block tree structures are applied, a luma CTB (Coding Tree Block) can be divided into CUs based on a specific coding tree structure, and a chroma CTB can be divided into chroma CUs based on a different coding tree structure. That is, a CU within an I slice / tile group to which a separate block tree structure is applied can consist of a coding block for a luma component or a coding block for two chroma components, while a CU in a P or B slice / tile group can consist of a block for three color components (a luma component and two chroma components).

[0105] In the above, a quadtree coding tree structure with a multitype tree was described, but the structure in which a CU is split is not limited to this. For example, BT structures and TT structures can be interpreted as concepts included in multiple partitioning tree (MPT) structures, and a CU can be interpreted as being split by QT structures and MPT structures. In one example in which a CU is split by QT and MPT structures, the split structure can be determined by signaling a syntax element (e.g., MPT_split_type) containing information about how the leaf nodes of the QT structure are split into several blocks, and a syntax element (e.g., MPT_split_mode) containing information about whether the leaf nodes of the QT structure are split in the vertical or horizontal direction.

[0106] In another example, the CU can be divided in a way different from the QT, BT, or TT structures. That is, unlike the QT structure which divides the lower-depth CU into quarters the size of the upper-depth CU, or the BT structure which divides the lower-depth CU into half the size of the upper-depth CU, or the TT structure which divides the lower-depth CU into quarters or half the size of the upper-depth CU, the lower-depth CU can, depending on the case, be divided into 1 / 5, 1 / 3, 3 / 8, 3 / 5, 2 / 3, or 5 / 8 the size of the upper-depth CU, and the way in which the CU is divided is not limited to this.

[0107] Thus, the quadtree coding block structure with the multitype tree can provide a highly flexible block partition structure. On the other hand, due to the partition types supported by the multitype tree, different partition patterns may, in some cases, lead to potentially identical coding block structures. By limiting the occurrence of such redundant partition patterns, the encoding and decoding devices can reduce the amount of data in the partition information.

[0108] Furthermore, in the video / image encoding and decoding according to this disclosure, the image processing units can have a hierarchical structure. A picture can be divided into one or more tiles, bricks, slices, and / or tile groups. A slice can contain one or more bricks. A brick can contain one or more CTU rows within a tile. A slice can contain an integer number of bricks in a picture. A tile group can contain one or more tiles. A tile can contain one or more CTUs. The CTU can be divided into one or more CUs. A tile may be a rectangular area in a picture consisting of a specific tile row and a specific tile column, each consisting of multiple CTUs. A tile group can contain an integer number of tiles obtained by a tile raster scan in a picture. A slice header can carry information / parameters applicable to the slice (blocks within the slice). If the encoding or decoding device has a multicore processor, the encoding / decoding procedures for the tiles, slices, bricks, and / or tile groups can be processed in parallel.

[0109] In this disclosure, the designations or concepts of slice and tile group may be used interchangeably. That is, a tile group header may be called a slice header. Here, a slice may have one of the slice types, including intra(I)slice, predictive(P)slice, and bi-predictive(B)slice. For blocks in an I slice, only intra-predictive prediction may be used for prediction, and inter-predictive prediction may not be used. Of course, even in this case, original sample values ​​can be coded and signaled without prediction. For blocks in a P slice, intra-predictive or inter-predictive prediction may be used, and if inter-predictive prediction is used, only uni-predictive prediction may be used. On the other hand, for blocks in a B slice, intra-predictive or inter-predictive prediction may be used, and if inter-predictive prediction is used, up to bi-predictive prediction may be used.

[0110] The encoding device can determine tile / tile group, brick, slice, and maximum and minimum coding unit sizes depending on the characteristics of the video image (e.g., resolution), or considering coding efficiency or parallel processing. Information related to this, or information that can guide this determination, may be included in the bitstream.

[0111] The decoding device can now obtain information indicating whether the picture's tiles / tile groups, bricks, slices, and CTUs within the tiles have been divided into numerous coding units. The encoding and decoding devices can also improve encoding efficiency by signaling this information only under specific conditions.

[0112] The slice header (slice header syntax) may include information / parameters that are commonly applicable to the slice. APS (APS syntax) or PPS (PPS syntax) may include information / parameters that are commonly applicable to one or more pictures. SPS (SPS syntax) may include information / parameters that are commonly applicable to one or more sequences. VPS (VPS syntax) may include information / parameters that are commonly applicable to multiple layers. DPS (DPS syntax) may include information / parameters that are commonly applicable to video in general. DPS may include information / parameters related to the joining of CVS (coded video sequence).

[0113] Furthermore, information regarding the division and configuration of tiles / tile groups / bricks / slices, for example, can be configured during the encoding stage via the higher-level syntax and transmitted to the decoding device in bitstream format.

[0114] Overview of Intra Prediction

[0115] The intra-prediction performed by the encoding and decoding devices described above will be explained in more detail below. Intra-prediction can be described as a prediction that generates a prediction sample for the current block based on a reference sample within the picture to which the current block belongs (hereinafter referred to as the current picture).

[0116] This will be explained with reference to Figure 8. When intraprediction is applied to block 801, the peripheral reference samples to be used for intraprediction of block 801 can be derived. The peripheral reference samples of the current block may include a total of 2 × nH samples, including sample 811 adjacent to the left boundary of the current block of size nW × nH and sample 812 adjacent to the bottom-left, a total of 2 × nW samples, including sample 821 adjacent to the top boundary of the current block and sample 822 adjacent to the top-right, and one sample 831 adjacent to the top-light of the current block. Alternatively, the peripheral reference samples of the current block may include upper peripheral samples in multiple columns and left peripheral samples in multiple rows.

[0117] Furthermore, the peripheral reference samples of the current block may also include a total of nH samples 841 adjacent to the right boundary of the current block of size nW × nH, a total of nW samples 851 adjacent to the bottom boundary of the current block, and one sample 842 adjacent to the bottom-right side of the current block.

[0118] However, some of the surrounding reference samples in the current block may not yet be decoded or available. In this case, the decoder can construct the surrounding reference samples to be used for prediction by substituting the unavailable samples with available samples, or by interpolating the available samples.

[0119] If peripheral reference samples are derived, (i) predicted samples can be derived based on the average or interpolation of the neighboring reference samples of the current block, or (ii) predicted samples can be derived based on reference samples that are located in a specific (predicted) direction relative to the predicted sample among the peripheral reference samples of the current block. Case (i) can be called a non-directional mode or non-angular mode, and case (ii) can be called a directional mode or angular mode. In addition, among the peripheral reference samples, the predicted sample can also be generated through interpolation between the first peripheral sample and a second peripheral sample located in the opposite direction to the prediction direction of the intra-prediction mode of the current block, with respect to the predicted sample of the current block. In the above case, it can be called linear interpolation intra-prediction (LIP). In addition, chroma predicted samples can be generated based on chroma samples using a linear model. In this case, it can be called LM mode. Alternatively, a temporary prediction sample for the current block can be derived based on filtered peripheral reference samples, and the prediction sample for the current block can be derived by performing a weighted sum of the temporary prediction sample and at least one reference sample derived according to the intra-prediction mode from the existing peripheral reference samples, i.e., unfiltered peripheral reference samples. In this case, it can be called PDPC (Position dependent intra prediction). In addition, intra-predictive coding can be performed by selecting the reference sample line with the highest prediction accuracy from among the peripheral multiple reference sample lines of the current block, deriving the prediction sample using the reference sample located in the prediction direction on that line, and then instructing (signaling) the decoder to use the reference sample line.In the above case, it can be called multi-reference line (MRL) intra prediction or MRL-based intra prediction. Furthermore, intra prediction is performed based on the same intra prediction mode for each vertical or horizontal subpartition of the current block, but peripheral reference samples can be derived and used on a subpartition-by-subpartition basis. That is, in this case, the intra prediction mode for the current block is applied identically to the subpartition, but by deriving and using peripheral reference samples on a subpartition-by-subpartition basis, intra prediction performance can be improved in some cases. Such prediction methods can be called intra subpartitions (ISP) or ISP-based intra prediction. Such intra prediction methods can be distinguished from intra prediction modes (e.g., DC mode, Planar mode, and directional mode) and referred to as intra prediction types. These intra prediction types can be referred to by various terms such as intra prediction techniques or additional intra prediction modes. For example, the intra prediction type (or additional intra prediction mode, etc.) may include at least one of the above-mentioned LIP, PDPC, MRL, and ISP. General intra-prediction methods, excluding specific intra-prediction types such as LIP, PDPC, MRL, and ISP, can be called normal intra-prediction types. Normal intra-prediction types refer to cases where the specific intra-prediction types described above do not apply, and predictions can be made based on the intra-prediction modes described above. Meanwhile, post-processing filtering can be performed on the derived prediction samples as needed.

[0120] Specifically, the intra-prediction procedure may include an intra-prediction mode / type determination step, a peripheral reference sample derivation step, and an intra-prediction mode / type-based prediction sample derivation step. Furthermore, a post-filtering step may be performed on the derived prediction samples as needed.

[0121] On the other hand, in addition to the intra-prediction types described above, ALWIP (affine linear weighted intra prediction) can also be used. ALWIP can also be called LWIP (linear weighted intra prediction) or MIP (matrix weighted intra prediction or matrix-based intra prediction). When MIP is applied to a current block, predicted samples for the current block can be derived by i) using averaging peripheral reference samples, ii) performing a matrix-vector-multiplication procedure, and iii) further performing horizontal / vertical interpolation procedures as needed. The intra-prediction mode used for MIP can be configured differently from the intra-prediction modes used for LIP, PDPC, MRL, ISP intra-prediction, or normal intra-prediction described above. The intra-prediction mode for MIP can be called the MIP intra-prediction mode, MIP prediction mode, or MIP mode. For example, the matrix and offset used in the matrix-vector-multiplication can be set differently depending on the intra-prediction mode for MIP. Here, the matrix can be called the (MIP) weight matrix, and the offset can be called the (MIP) offset vector or (MIP) bias vector. The specific MIP method will be described later.

[0122] A block reconstruction procedure based on intra-prediction and an intra-prediction unit within an encoding device may schematically include the following: S910 can be performed by the intra-prediction unit 185 of the encoding device, and S920 can be performed by a residual processing unit including at least one of the subtraction unit 115, transformation unit 120, quantization unit 130, inverse quantization unit 140, and inverse transformation unit 150 of the encoding device. Specifically, S920 can be performed by the subtraction unit 115 of the encoding device. In S930, prediction information can be derived by the intra-prediction unit 185 and encoded by the entropy encoding unit 190. In S930, residual information can be derived by the residual processing unit and encoded by the entropy encoding unit 190. The residual information is information about the residual sample. The residual information may include information about the quantized transformation coefficients for the residual sample. As described above, the residual sample is derived into a conversion coefficient via the conversion unit 120 of the encoding device, and the conversion coefficient can be derived as a quantized conversion coefficient via the quantization unit 130. Information regarding the quantized conversion coefficient can be encoded in the entropy encoding unit 190 via the residual coding procedure.

[0123] The encoding device can perform intra-prediction for the current block (S910). The encoding device can derive an intra-prediction mode / type for the current block, derive peripheral reference samples for the current block, and generate predicted samples within the current block based on the intra-prediction mode / type and the peripheral reference samples. Here, the procedures for determining the intra-prediction mode / type, deriving peripheral reference samples, and generating predicted samples may be performed simultaneously, or one of the procedures may be performed before the others. For example, although not shown, the intra-prediction unit 185 of the encoding device may include an intra-prediction mode / type determination unit, a reference sample derivation unit, and a predicted sample derivation unit, where the intra-prediction mode / type determination unit determines the intra-prediction mode / type for the current block, the reference sample derivation unit derives peripheral reference samples for the current block, and the predicted sample derivation unit derives predicted samples for the current block. On the other hand, if the predicted sample filtering procedure described later is performed, the intra-prediction unit 185 may further include a predicted sample filtering unit. The encoding device can determine which of a plurality of intra-prediction modes / types is to apply to the current block. The encoding device can compare the RD costs for the intra-prediction modes / types and determine the optimal intra-prediction mode / type for the current block.

[0124] On the other hand, the encoding device can also perform a predictive sample filtering procedure. This predictive sample filtering can be called post-filtering. The predictive sample filtering procedure can filter out some or all of the predictive samples. In some cases, the predictive sample filtering procedure can be omitted.

[0125] The encoding device can generate a residual sample for the current block based on the (filtered) predicted sample (S920). The encoding device can derive the residual sample by comparing the predicted sample with the original sample of the current block based on phase.

[0126] The encoding device can encode image information including information relating to the intra-prediction (prediction information) and residual information relating to the residual sample (S930). The prediction information may include the intra-prediction mode information and the intra-prediction type information. The encoding device can output the encoded image information in bitstream format. The output bitstream can be transmitted to a decoding device via a storage medium or network.

[0127] The residual information may include the residual coding syntax described later. The encoding device can transform / quantize the residual samples to derive quantized transformation coefficients. The residual information may include information relating to the quantized transformation coefficients.

[0128] On the other hand, as described above, the encoding device can generate a restored picture (including restored samples and restored blocks). To this end, the encoding device can decrypt the quantized conversion coefficients again to derive (corrected) residual samples. The reason for decrypting / quantizing the residual samples again is to derive the same residual samples as those derived by the decoding device, as described above. Based on the predicted samples and the (corrected) residual samples, the encoding device can generate a restored block containing restored samples for the current block. Based on the restored block, a restored picture for the current picture can be generated. As described above, in-loop filtering procedures and the like can be further applied to the restored picture.

[0129] The intra-prediction-based video / image decoding procedure and the intra-prediction unit within the decoding device may include, schematically, the following: The decoding device can perform operations corresponding to those performed by the encoding device.

[0130] S1010 to S1030 can be performed by the intra-prediction unit 265 of the decoding device, and the prediction information in S1010 and the residual information in S1040 can be obtained from the bitstream by the entropy decoding unit 210 of the decoding device. A residual processing unit including at least one of the inverse quantization unit 220 and the inverse transform unit 230 of the decoding device can derive a residual sample for the current block based on the residual information. Specifically, the inverse quantization unit 220 of the residual processing unit can derive a transformation coefficient by performing inverse quantization based on the quantized transformation coefficient derived based on the residual information, and the inverse transform unit 230 of the residual processing unit can derive a residual sample for the current block by performing an inverse transform on the transformation coefficient. S1050 can be performed by the addition unit 235 or the restoration unit of the decoding device.

[0131] Specifically, the decoding device can derive an intra-prediction mode / type for the current block based on the received prediction information (intra-prediction mode / type information) (S1010). The decoding device can derive a peripheral reference sample for the current block (S1020). The decoding device can generate a prediction sample within the current block based on the intra-prediction mode / type and the peripheral reference sample (S1030). In this case, the decoding device can perform a prediction sample filtering procedure. Prediction sample filtering can be called post-filtering. The prediction sample filtering procedure can filter out some or all of the prediction samples. In some cases, the prediction sample filtering procedure can be omitted.

[0132] The decoding device can generate a residual sample for the current block based on the received residual information. The decoding device can generate a restored sample for the current block based on the predicted sample and the residual sample, and derive a restored block containing the restored sample (S1040). A restored picture for the current picture can be generated based on the restored block. As described above, in-loop filtering procedures and the like can be further applied to the restored picture.

[0133] Here, the intra-prediction unit 265 of the decoding device may include, even if not shown, an intra-prediction mode / type determination unit, a reference sample derivation unit, and a prediction sample derivation unit. The intra-prediction mode / type determination unit determines the intra-prediction mode / type for the current block based on the intra-prediction mode / type information acquired by the entropy decoding unit 210, the reference sample derivation unit derives a peripheral reference sample of the current block, and the prediction sample derivation unit derives a prediction sample of the current block. On the other hand, if the prediction sample filtering procedure described above is performed, the intra-prediction unit 265 may further include a prediction sample filtering unit.

[0134] The intra-prediction mode information may include, for example, flag information (e.g., intra_luma_mpm_flag) indicating whether the MPM (most probable mode) or the remaining mode is applied to the current block. If the MPM is applied to the current block, the prediction mode information may further include index information (e.g., intra_luma_mpm_idx) pointing to one of the intra-prediction mode candidates (MPM candidates). The intra-prediction mode candidates (MPM candidates) may consist of an MPM candidate list or an MPM list. If the MPM is not applied to the current block, the intra-prediction mode information may further include remaining mode information (e.g., intra_luma_mpm_remainder) pointing to one of the remaining intra-prediction modes excluding the intra-prediction mode candidates (MPM candidates). The decoding device can determine the intra-prediction mode of the current block based on the intra-prediction mode information. A separate MPM list may be configured for the MIP described above.

[0135] Furthermore, the intra-prediction type information can be implemented in various forms. For example, the intra-prediction type information may include intra-prediction type index information indicating any one of the intra-prediction types. As another example, the intra-prediction type information may include reference sample line information (e.g., intra_luma_ref_idx) indicating whether the MRL is applied to the current block and, if so, which reference sample line is used; ISP flag information (e.g., intra_subpartitions_mode_flag) indicating whether the ISP is applied to the current block; ISP type information (e.g., intra_subpartitions_split_flag) indicating the split type of the subpartition if the ISP is applied; flag information indicating whether PDCP is applied; or flag information indicating whether LIP is applied. The intra-prediction type information may also include an MIP flag indicating whether MIP is applied to the current block.

[0136] The intra-prediction mode information and / or the intra-prediction type information can be encoded / decoded via the coding methods described herein. For example, the intra-prediction mode information and / or the intra-prediction type information can be encoded / decoded via entropy coding (e.g., CABAC, CAVLC) coding based on truncated (rice) binary code.

[0137] Overview of Interpretation

[0138] The following describes the detailed techniques of the interpretation method in the description of encoding and decoding with reference to Figures 2 and 3. In the case of a decoding device, the video / image decoding method based on interpretation and the interpretation unit within the decoding device can operate according to the following description. In the case of an encoding device, the video / image encoding method based on interpretation and the interpretation unit within the encoding device can operate according to the following description. In addition, the data encoded according to the following description can be stored in bitstream format.

[0139] The prediction unit of the encoding / image decoding device can perform interpretation on a block-by-block basis to derive predicted samples. Interpretation can indicate predictions derived in a manner dependent on data elements of pictures other than the current picture (e.g., sample values ​​or motion information). When interpretation is applied to the current block, a predicted block (predicted sample array) for the current block can be derived based on the reference block (reference sample array) identified by the motion vector on the reference picture pointed to by the reference picture index. At this time, in order to reduce the amount of motion information transmitted in interpretation mode, the motion information of the current block can be predicted on a block, subblock, or sample basis based on the correlation of motion information between the surrounding blocks and the current block. The motion information may include the motion vector and the reference picture index. The motion information may further include interpretation type information (L0 prediction, L1 prediction, Bi prediction, etc.). When interpretation is applied, the surrounding blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture containing the aforementioned reference block and the reference picture containing the aforementioned time-peripheral block may be the same or different. The time-peripheral block may be referred to by names such as collocated reference block (colCU), and the reference picture containing the time-peripheral block may be referred to by name as collocated picture (colPic). For example, a list of motion information candidates can be constructed based on the surrounding blocks of the current block, and a flag or index information indicating which candidate is selected (used) can be signaled to derive the motion vector and / or reference picture index of the current block. Interpretation can be performed based on various prediction modes; for example, in skip mode and merge mode, the motion information of the current block may be the same as the motion information of the selected surrounding block.In skip mode, unlike merge mode, the residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected surrounding block can be used as the motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.

[0140] The motion information may include L0 motion information and / or L1 motion information based on the interpretation type (L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction may be called an L0 motion vector or MVL0, and a motion vector in the L1 direction may be called an L1 motion vector or MVL1. A prediction based on the L0 motion vector may be called an L0 prediction, a prediction based on the L1 motion vector may be called an L1 prediction, and a prediction based on both the L0 motion vector and the L1 motion vector may be called a bi (Bi) prediction. Here, the L0 motion vector may represent a motion vector associated with the reference picture list L0 (L0), and the L1 motion vector may represent a motion vector associated with the reference picture list L1 (L1). The reference picture list L0 may include pictures earlier in the output order than the current picture as reference pictures, and the reference picture list L1 may include pictures later in the output order than the current picture. The aforementioned earlier pictures can be called forward (reference) pictures, and the aforementioned later pictures can be called backward (reference) pictures. The reference picture list L0 may further include pictures that are later in the output order than the current picture as reference pictures. In this case, the earlier pictures may be indexed first in the reference picture list L0, and the later pictures may be indexed next. The reference picture list L1 may further include pictures that are earlier in the output order than the current picture as reference pictures. In this case, the later pictures may be indexed first in the reference picture list L1, and the earlier pictures may be indexed next. Here, the output order can correspond to the POC (picture order count) order.

[0141] A video / image coding procedure based on interpretation and an interpretation unit within an coding device may include, in general terms, the following, as illustrated with reference to Figure 11. The coding device performs interpretation for the current block (S1110). The coding device can derive the interpretation mode and motion information of the current block and generate a prediction sample for the current block. Here, the interpretation mode determination, motion information derivation, and prediction sample generation procedures may be performed simultaneously, or one procedure may be performed before the others. For example, the interpretation unit of the coding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit, where the prediction mode determination unit determines the prediction mode for the current block, the motion information derivation unit derives the motion information of the current block, and the prediction sample derivation unit derives a prediction sample for the current block. For example, the interpretation unit of the coding device can search for blocks similar to the current block within a certain area (search area) of a reference picture via motion estimation and derive a reference block whose difference from the current block is the minimum or below a certain standard. Based on this, a reference picture index pointing to the reference picture where the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding device can determine which of the various prediction modes is applied to the current block. The encoding device can compare the RD costs for the various prediction modes and determine the optimal prediction mode for the current block.

[0142] For example, when a skip mode or merge mode is applied to the current block, the encoding device can configure a merge candidate list, as described later, and derive a reference block from among the reference blocks pointed to by the merge candidates included in the merge candidate list whose difference from the current block is the minimum or below a certain standard. In this case, a merge candidate associated with the derived reference block is selected, and merge index information pointing to the selected merge candidate is generated and signaled to the decoding device. The movement information of the current block can be derived using the movement information of the selected merge candidate.

[0143] As another example, when the (A)MVP mode is applied to the current block, the encoding device can configure the (A)MVP candidate list described later, and use the motion vector of the selected mvp candidate from among the mvp (motion vector predictor) candidates included in the (A)MVP candidate list as the mvp of the current block. In this case, for example, the motion vector pointing to the reference block derived by the motion estimation described above can be used as the motion vector of the current block, and the mvp candidate with the smallest difference from the motion vector of the current block among the mvp candidates can become the selected mvp candidate. The MVD (motion vector difference), which is the difference obtained by subtracting the mvp from the motion vector of the current block, can be derived. In this case, information regarding the MVD can be signaled to the decoding device. Also, when the (A)MVP mode is applied, the value of the reference picture index can be configured with reference picture index information and separately signaled to the decoding device.

[0144] The encoding device can derive a residual sample based on the predicted sample (S1120). The encoding device can derive the residual sample by comparing the original sample of the current block with the predicted sample.

[0145] The encoding device encodes image information including prediction information and residual information (S1130). The encoding device can output the encoded image information in bitstream format. The prediction information is information related to the prediction procedure and may include prediction mode information (e.g., skip flag, merge flag, or mode index) and motion information. The motion information may include candidate selection information (e.g., merge index, mvp flag, or mvp index) which is information for deriving a motion vector. The motion information may also include the above-mentioned MVD information and / or reference picture index information. Furthermore, the motion information may include information indicating whether L0 prediction, L1 prediction, or bi prediction is applied. The residual information is information about the residual sample. The residual information may include information about the quantized transformation coefficients for the residual sample.

[0146] The output bitstream can be stored on a (digital) storage medium and transmitted to a decoding device, or it can be transmitted to a decoding device via a network.

[0147] On the other hand, as described above, the encoding device can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference sample and the residual sample. This is because the encoding device can derive the same prediction results as the decoding device, thereby improving coding efficiency. Therefore, the encoding device can store the reconstructed picture (or reconstructed sample, reconstructed block) in memory and use it as a picture for interpretation. As described above, in-loop filtering procedures and the like can be further applied to the reconstructed picture.

[0148] The video / image decoding procedure and the interpretation unit within the decoding device based on interpretation prediction may, in general, include, for example, the following:

[0149] The decoding device can perform operations corresponding to those performed by the encoding device. The decoding device can make predictions for the current block based on the received prediction information and derive prediction samples.

[0150] Specifically, the decoding device can determine the prediction mode for the current block based on the received prediction information (S1210). Based on the prediction mode information in the prediction information, the decoding device can determine which inter-prediction mode is applied to the current block.

[0151] For example, based on the merge flag, it can be determined whether the merge mode is applied to the current block or whether the (A)MVP mode is determined. Alternatively, based on the mode index, one of several inter-prediction mode candidates can be selected. The inter-prediction mode candidates may include skip mode, merge mode and / or (A)MVP mode, or may include various inter-prediction modes as described later.

[0152] The decoding device derives motion information for the current block based on the determined interprediction mode (S1220). For example, if a skip mode or merge mode is applied to the current block, the decoding device can configure a merge candidate list, as described later, and select one of the merge candidates included in the merge candidate list. This selection can be made based on the selection information (merge index) described above. The motion information for the current block can be derived using the motion information for the selected merge candidate. The motion information for the selected merge candidate can be used as the motion information for the current block.

[0153] As another example, when the (A)MVP mode is applied to the current block, the decoding device can configure the (A)MVP candidate list described later, and use the motion vector of the mvp candidate selected from the mvp (motion vector predictor) candidates included in the (A)MVP candidate list as the mvp of the current block. The selection can be made based on the selection information (mvp flag or mvp index) described above. In this case, the MVD of the current block can be derived based on the information regarding the MVD, and the motion vector of the current block can be derived based on the mvp of the current block and the MVD. Furthermore, the reference picture index of the current block can be derived based on the reference picture index information. The picture pointed to by the reference picture index in the related reference picture list for the current block can be derived as the reference picture referenced for interpretation of the current block.

[0154] On the other hand, as will be described later, the movement information of the current block can be derived without constructing a candidate list, in which case the movement information of the current block can be derived according to the procedure disclosed in the prediction mode described later. In this case, the above-mentioned candidate list configuration can be omitted.

[0155] The decoding device can generate predicted samples for the current block based on the motion information of the current block (S1230). In this case, the reference picture can be derived based on the reference picture index of the current block, and the predicted samples for the current block can be derived using the sample of the reference block pointed to by the motion vector of the current block on the reference picture. In this case, as will be described later, a prediction sample filtering procedure may be performed on all or some of the predicted samples for the current block, depending on the circumstances.

[0156] For example, the interpretation unit of the decoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. Based on the prediction mode information received from the prediction mode determination unit, the prediction mode for the current block is determined. Based on the motion information information received from the motion information derivation unit, motion information (such as motion vectors and / or reference picture indices) for the current block is derived. The prediction sample derivation unit can then derive prediction samples for the current block.

[0157] The decoding device generates a residual sample for the current block based on the received residual information (S1240). The decoding device can generate a reconstructed sample for the current block based on the predicted sample and the residual sample, and generate a reconstructed picture based on this (S1250). As described above, further procedures such as in-loop filtering can be applied to the reconstructed picture thereafter.

[0158] As described above, the interpretation procedure may include an interpretation mode determination step, a motion information derivation step based on the determined prediction mode, and a prediction execution (prediction sample generation) step based on the derived motion information. The interpretation procedure can be performed in the encoding and decoding devices as described above.

[0159] Quantization / Dequantization

[0160] As described above, the quantization unit of the encoding device can derive quantized conversion coefficients by applying quantization to the conversion coefficients, and the inverse quantization unit of the encoding device or the inverse quantization unit of the decoding device can derive conversion coefficients by applying inverse quantization to the quantized conversion coefficients.

[0161] In the encoding and decoding of moving / still images, the quantization rate can be changed, and the compression rate can be adjusted using the changed quantization rate. From an implementation standpoint, instead of directly using the quantization rate, a quantization parameter (QP) can be used, taking complexity into consideration. For example, a quantization parameter with integer values ​​from 0 to 63 can be used, and each quantization parameter value can correspond to an actual quantization rate. Furthermore, the quantization parameter QP for the luma component (luma sample) Y And the quantization parameter QP for the chromatic component (chromatic sample) C It can be configured differently.

[0162] The quantization process takes a transformation coefficient C as input and divides it by the quantization rate (Qstep) to obtain the quantized transformation coefficient C'. In this case, considering the computational complexity, the quantization rate can be multiplied by a scale to obtain an integer form, and a shift operation can be performed only for values ​​corresponding to the scale value. The quantization scale can be derived based on the product of the quantization rate and the scale value. That is, the quantization scale can be derived according to QP. The quantized transformation coefficient C' can also be derived by applying the quantization scale to the transformation coefficient C.

[0163] The inverse quantization process is the reverse process of the quantization process. By multiplying the quantized transformation coefficient C' by the quantization rate Qstep, the restored transformation coefficient C'' can be obtained. In this case, a level scale can be derived according to the quantization parameters, and the restored transformation coefficient C' can be derived by applying the level scale to the quantized transformation coefficient C'. The restored transformation coefficient C'' may differ somewhat from the initial transformation coefficient C due to losses in the transformation and / or quantization process. Therefore, inverse quantization can be performed in an encoding device as well as a decoding device.

[0164] On the other hand, adaptive frequency weighting quantization (AQU) is a technique that adjusts the quantization intensity according to frequency. This AQU is a method of applying different quantization intensities for different frequencies. This AQU can be applied using a predefined quantization scaling matrix to apply different quantization intensities for each frequency. In other words, the quantization / dequantization process described above can be performed based more on this quantization scaling matrix. For example, different quantization scaling matrices can be used depending on whether the prediction mode applied to the current block is inter-prediction or intra-prediction to generate the current block size and / or the current block's residual signal. This quantization scaling matrix can be called a quantization matrix or scaling matrix. This quantization scaling matrix can be predefined. Furthermore, for frequency adaptive scaling, frequency-specific quantization scale information for the quantization scaling matrix can be configured / encoded in an encoding device and signaled to a decoding device. This frequency-specific quantization scale information can be called quantization scaling information. The frequency-specific quantization scale information may include scaling list data. Based on the scaling list data, the (modified) quantization scaling matrix can be derived. The frequency-specific quantization scale information may also include present flag information indicating the presence or absence of the scaling list data. Alternatively, if the scaling list data is signaled at a higher level (e.g., SPS), it may further include information indicating whether the scaling list data is modified at a lower level (e.g., PPS or tile group header).

[0165] Conversion / Inverse Conversion

[0166] As described above, the encoding device can derive residual blocks (residual samples) based on predicted blocks (predicted samples) via intra / inter / IBC prediction, and can apply transformation and quantization to the derived residual samples to derive quantized transformation coefficients. Information on the quantized transformation coefficients (residual information) can be encoded in the residual coding syntax and then output in bitstream format. The decoding device can obtain the information on the quantized transformation coefficients (residual information) from the bitstream, decode it, and derive quantized transformation coefficients. The decoding device can derive residual samples via inverse quantization / inverse transformation based on the quantized transformation coefficients. As described above, at least one of the quantization / inverse quantization and / or transformation / inverse transformation is optional. If the transformation / inverse transformation is omitted, the transformation coefficients may also be called coefficients or residual coefficients, or may still be called transformation coefficients for consistency of representation. Whether the transformation / inverse transformation is omitted can be signaled based on a transformation skip flag (e.g., transform_skip_flag).

[0167] The aforementioned transformations / inverse transformations can be performed based on transformation kernels. For example, a multiple transform selection (MTS) scheme can be applied to perform transformations / inverse transformations. In this case, a portion of a set of multiple transformation kernels is selected and applied to the current block. Transformation kernels can be referred to by various terms, such as transformation matrices or transformation types. For example, a set of transformation kernels can represent a combination of vertical transformation kernels and horizontal transformation kernels.

[0168] The conversion / inverse conversion can be performed on a CU or TU basis. That is, the conversion / inverse conversion can be applied to residual samples within a CU or residual samples within a TU. The CU size and TU size may be the same, or multiple TUs may exist within the CU region. On the other hand, the CU size can generally represent the luminous component (sample) CB size. The TU size can generally represent the luminous component (sample) TB size. The chroma component (sample) CB or TB size can be derived based on the luminous component (sample) CB or TB size according to the component ratio of the color format (chroma format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.). The TU size can be derived based on maxTbSize. For example, if the CU size is larger than maxTbSize, multiple TUs (TBs) of maxTbSize can be derived from the CU, and the conversion / inverse conversion can be performed on a TU (TB) basis. maxTbSize can be considered when deciding whether to apply various intra-prediction types such as ISP. The information for maxTbSize may be predetermined, or it may be generated and encoded by an encoding device and signaled to the encoding device.

[0169] Entropy coding

[0170] As previously explained with reference to Figure 2, some or all of the video / image information can be entropically encoded by the entropy encoding unit 190, and some or all of the video / image information, as explained with reference to Figure 3, can be entropically decoded by the entropy decoding unit 310. In this case, the video / image information can be encoded / decoded on a syntax element basis. In this disclosure, the encoding / decoding of information may include encoding / decoding by the methods described in this paragraph.

[0171] Figure 13 shows a CABAC block diagram for encoding a single syntax element. The CABAC encoding process first converts the input signal to a binary value via binarization if the input signal is a syntax element rather than a binary value. If the input signal is already a binary value, the binarization step can be bypassed. Here, each binary digit 0 or 1 that makes up the binary value can be called a bin. For example, if the binary string after binarization (binstring) is 110, then 1, 1, and 0 can each be called a bin. The bins for a syntax element can represent the value of that syntax element.

[0172] Binarized bins can be input to either a regular coding engine or a bypass coding engine. The regular coding engine can assign a context model that reflects the probability values ​​to each bin and encode the bin based on the assigned context model. After coding each bin, the regular coding engine can update the probability model for that bin. These coded bins can be called context-coded bins. The bypass coding engine can omit the steps of estimating probabilities for the input bins and updating the probability model applied to the bins after coding. In the case of the bypass coding engine, coding speed can be improved by applying a uniform probability distribution (e.g., 50:50) to the input bins instead of assigning a context. These coded bins can be called bypass bins. Context models can be assigned and updated for each context-coded (regularly coded) bin, and context models can be indicated based on ctxidx or ctxInc. ctxidx can be derived based on ctxInc. Specifically, for example, a context index (ctxidx) that points to the context model for each of the normally coded bins can be derived as the sum of a context index increment (ctxInc) and a context index offset (ctxIdxOffset). Here, ctxInc can be derived differently for each bin. ctxIdxOffset can be represented by the lowest value of ctxIdx. The lowest value of ctxIdx can be called the initial value (initValue) of ctxIdx.The aforementioned ctxIdxOffset is a value generally used to distinguish it from the context model for other syntax elements, and the context model for a single syntax element can be distinguished / derived based on ctxinc.

[0173] The entropy coding procedure allows for the determination of whether to perform coding via a regular coding engine or a bypass coding engine, and the coding path can be switched accordingly. Entropy decoding can perform the same process as entropy coding in reverse order.

[0174] The entropy coding described above can be performed, for example, as shown in Figures 14 and 15. Referring to Figures 14 and 15, the encoding device (entropy encoding unit) can perform an entropy coding procedure for image / video information. The image / video information may include partitioning-related information, prediction-related information (e.g., inter / intra prediction division information, intra prediction mode information, inter prediction mode information, etc.), residual information, in-loop filtering-related information, etc., or may include various syntax elements related thereto. The entropy coding can be performed on a syntax element-by-syntax basis. Steps S1410 to S1420 in Figure 14 can be performed by the entropy encoding unit 190 of the encoding device shown in Figure 2 described above.

[0175] The encoding device can perform binarization on the target syntax element (S1410). Here, the binarization can be based on various binarization methods such as the Truncated Rice binarization process and the Fixed-length binarization process, and the binarization method for the target syntax element can be predefined. The binarization procedure can be performed by the binarization unit 191 within the entropy encoding unit 190.

[0176] The encoding device can perform entropy coding on the target syntax element (S1420). Based on entropy coding techniques such as CABAC (context-adaptive arithmetic coding) or CAVLC (context-adaptive variable length coding), the encoding device can perform normal coding-based (context-based) or bypass coding-based coding on the binstring of the target syntax element, and the output can be included in the bitstream. The entropy coding procedure can be performed by the entropy coding processing unit 192 within the entropy coding unit 190. As described above, the bitstream can be transmitted to the decoding device via a (digital) storage medium or a network.

[0177] Referring to Figures 16 and 17, the decoding device (entropy decoding unit) can decode the encoded image / video information. The image / video information may include partitioning-related information, prediction-related information (e.g., inter / intra prediction partitioning information, intra prediction mode information, inter prediction mode information, etc.), residual information, in-loop filtering-related information, etc., or may include various syntax elements related thereto. The entropy coding can be performed on a syntax element basis. Steps S1610 to S1620 can be performed by the entropy decoding unit 210 of the decoding device shown in Figure 3 above.

[0178] The decoding device can perform binarization on the target syntax element (S1610). Here, the binarization can be based on various binarization methods such as the Truncated Rice binarization process and the Fixed-length binarization process, and the binarization method for the target syntax element can be predefined. Through the binarization procedure, the decoding device can derive available binstrings (binstring candidates) for the available values ​​of the target syntax element. The binarization procedure can be performed by the binarization unit 211 within the entropy decoding unit 210.

[0179] The decoding device can perform entropy decoding on the target syntax element (S1620). The decoding device can sequentially decode and parse each bin for the target syntax element from the input bits in the bitstream, and compare the derived binstring with the available binstrings for the syntax element. If the derived binstring is the same as one of the available binstrings, the value corresponding to that binstring can be derived as the value of the syntax element. If not, the next bit in the bitstream can be further parsed, and the above procedure can be repeated. Through this process, information can be signaled using variable-length bits without using start bits or end bits for specific information (specific syntax elements) in the bitstream. This allows for the allocation of relatively fewer bits to lower values, thereby improving overall coding efficiency.

[0180] The decoding device can perform context-based or bypass-based decoding from the bitstream to each bin in the bin string based on an entropy coding technique such as CABAC or CAVLC. The entropy decoding procedure can be performed by the entropy decoding processing unit 212 within the entropy decoding unit 210. The bitstream may contain various information for image / video decoding as described above. As described above, the bitstream can be transmitted to the decoding device via a (digital) storage medium or a network.

[0181] In this disclosure, a table containing syntax elements (syntax table) can be used to indicate the signaling of information from an encoding device to a decoding device. The order of the syntax elements in the table containing syntax elements used in this disclosure can indicate the parsing order of the syntax elements from a bitstream. An encoding device can configure and encode a syntax table so that the syntax elements can be parsed by a decoding device according to the parsing order, and a decoding device can obtain the values ​​of the syntax elements by parsing and decoding the syntax elements of the syntax table from a bitstream according to the parsing order.

[0182] General image / video coding procedure

[0183] In image / video coding, the pictures that make up an image / video can be encoded / decoded according to a set decoding order. The picture order, which corresponds to the output order of the decoded pictures, can be set to be different from the decoding order. Based on this, both forward and reverse prediction can be performed during interpretation.

[0184] Figure 18 shows an example of a schematic picture decoding procedure to which embodiments of this disclosure can be applied. In Figure 18, S1810 may be performed in the entropy decoding unit 210 of the decoding apparatus described in Figure 3, S1820 may be performed in the prediction unit including the intra-prediction unit 265 and the inter-prediction unit 260, S1830 may be performed in the residual processing unit including the inverse quantization unit 220 and the inverse transform unit 230, S1840 may be performed in the addition unit 235, and S1850 may be performed in the filtering unit 240. S1810 may include the information decoding procedure described in this disclosure, S1820 may include the inter / intra-prediction procedure described in this disclosure, S1830 may include the residual processing procedure described in this disclosure, S1840 may include the block / picture restoration procedure described in this disclosure, and S1850 may include the in-loop filtering procedure described in this disclosure.

[0185] Referring to Figure 18, the picture decoding procedure, as shown in the description of Figure 3, can schematically include a procedure for acquiring image / video information (by decoding) from a bitstream (S1810), a picture restoration procedure (S1820-S1840), and an in-loop filtering procedure (S1850) on the restored picture. The picture restoration procedure can be performed based on predicted samples and residual samples obtained through the inter / intra prediction (S1820) and residual processing (S1830, inverse quantization and inverse transformation of quantized transformation coefficients) processes described herein. A modified restored picture can be generated through an in-loop filtering procedure on the restored picture generated by the picture restoration procedure, and the modified restored picture can be output as a decoded picture, or it can be stored in the decoded picture buffer or memory 250 of the decoding device and used as a reference picture in the inter-prediction procedure when decoding subsequent pictures. In some cases, the in-loop filtering procedure can be omitted. In this case, the restored picture can be output as a decoded picture and can be stored in the decoded picture buffer or memory 250 of the decoding device and used as a reference picture in the inter-prediction procedure when decoding subsequent pictures. The in-loop filtering procedure (S1850) may include, as described above, a deblocking filtering procedure, an SAO (sample adaptive offset) procedure, an ALF (adaptive loop filter) procedure, and / or a bi-lateral filter procedure, and some or all of these may be omitted. In addition, one or some of the deblocking filtering procedure, the SAO (sample adaptive offset) procedure, the ALF (adaptive loop filter) procedure, and the bi-lateral filter procedure may be applied sequentially, or all of them may be applied sequentially.For example, the SAO procedure can be performed after the deblocking filtering procedure has been applied to the restored picture. Alternatively, the ALF procedure can be performed after the deblocking filtering procedure has been applied to the restored picture. This can also be done in the encoding device.

[0186] Figure 19 shows an example of a schematic picture coding procedure to which embodiments of the present disclosure can be applied. In Figure 19, S1910 may be performed in a prediction unit including an intra-prediction unit 185 or an inter-prediction unit 180 of the coding apparatus described in Figure 2, S1920 may be performed in a residual processing unit including a conversion unit 120 and / or a quantization unit 130, and S1930 may be performed in an entropy coding unit 190. S1910 may include the inter / intra-prediction procedure described in the present disclosure, S1920 may include the residual processing procedure described in the present disclosure, and S1930 may include the information coding procedure described in the present disclosure.

[0187] Referring to Figure 19, the picture encoding procedure, as shown in the description of Figure 2, may include not only a procedure for encoding information for picture reconstruction (e.g., prediction information, residual information, partitioning information, etc.) and outputting it in bitstream format, but also a procedure for generating a reconstructed picture for the current picture, and an optional procedure for applying in-loop filtering to the reconstructed picture. The encoding device can derive (modified) residual samples from the quantized conversion coefficients via the inverse quantization unit 140 and the inverse transform unit 150, and can generate a reconstructed picture based on the prediction samples, which are the output of S1910, and the (modified) residual samples. The reconstructed picture thus generated may be identical to the reconstructed picture generated by the decoding device described above. A modified reconstructed picture can be generated via an in-loop filtering procedure on the reconstructed picture, which can be stored in the decoding picture buffer or memory 170, and can be used as a reference picture in the inter-prediction procedure when encoding subsequent pictures, as in the case of the decoding device. As described above, in some cases, some or all of the in-loop filtering procedure may be omitted. When the in-loop filtering procedure is performed, the (in-loop) filtering-related information (parameters) can be encoded by the entropy coding unit 190 and output in bitstream format, and the decoding device can perform the in-loop filtering procedure in the same manner as the coding device based on the filtering-related information.

[0188] Through such in-loop filtering procedures, noise that occurs during image / video coding, such as blocking artifacts and ringing artifacts, can be reduced, thereby improving subjective and objective visual quality. Furthermore, by performing in-loop filtering procedures on both the encoding and decoding devices, the encoding and decoding devices can derive the same prediction results, increasing the reliability of picture coding and reducing the amount of data that needs to be transmitted for picture coding.

[0189] As described above, picture restoration procedures can be performed not only in the decoding device but also in the encoding device. Restored blocks can be generated based on intra-prediction / inter-prediction for each block, and a restored picture containing the restored blocks can be generated. If the current picture / slice / tile group is an I-picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based solely on intra-prediction. On the other hand, if the current picture / slice / tile group is a P or B-picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based on intra-prediction or inter-prediction. In this case, inter-prediction may be applied to some of the blocks in the current picture / slice / tile group, and intra-prediction may be applied to the remaining blocks. The color components of a picture may include luminous and chroma components, and unless expressly limited in this disclosure, the methods and embodiments proposed in this disclosure may be applied to luminous and chroma components.

[0190] Examples of coding hierarchy and structure

[0191] The coded video / images provided in this disclosure can be processed, for example, according to the coding hierarchy and structure described below.

[0192] Figure 20 shows the hierarchical structure for coded images. Coated images can be divided into the VCL (video coding layer), which handles the image decoding process and the image itself; the lower system, which transmits and stores the coded information; and the NAL (network abstraction layer), which exists between the VCL and the lower system and is responsible for network adaptation functions.

[0193] VCL can generate VCL data containing compressed image data (slice data), or generate parameter sets containing information such as Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS), or SEI (Supplemental Enhancement Information) messages additionally required for image decoding.

[0194] In NAL, NAL units can be generated by adding header information (NAL unit header) to the RBSP (Raw Byte Sequence Payload) generated by VCL. In this case, the RBSP refers to slice data, parameter sets, SEI messages, etc., generated by VCL. The NAL unit header may include NAL unit type information identified by the RBSP data contained in the NAL unit.

[0195] As illustrated, NAL units can be divided into VCL NAL units and Non-VCL NAL units by the RBSP generated in VCL. VCL NAL units can represent NAL units that contain information about the image (slice data), while Non-VCL NAL units can represent NAL units that contain information necessary to decode the image (parameter set or SEI message).

[0196] The VCL NAL units and Non-VCL NAL units described above can be transmitted over a network with header information added according to the data specifications of the underlying system. For example, NAL units can be transformed into data formats of predetermined standards such as H.266 / VVC file format, RTP (Real-time Transport Protocol), and TS (Transport Stream) and transmitted over various networks.

[0197] As described above, the NAL unit type can be identified according to the RBSP data structure contained within the NAL unit, and information about such NAL unit types can be stored in the NAL unit header and signaled.

[0198] For example, NAL units can be broadly classified into VCL NAL unit types and Non-VCL NAL unit types depending on whether or not they contain information (slice data) about the image. VCL NAL unit types can be further classified by the nature and type of picture they contain, while Non-VCL NAL unit types can be further classified by the type of parameter set.

[0199] The following is a list of examples of NAL unit types identified by the type of parameter set / information included in the Non-VCL NAL unit type.

[0200] -DCI (Decoding capability information) NAL unit: Type for NAL units that include DCI

[0201] -VPS (Video Parameter Set) NAL unit: Type for NAL units including VPS

[0202] -SPS (Sequence Parameter Set) NAL unit: Type for NAL units that include SPS

[0203] -PPS (Picture Parameter Set) NAL unit: Type for NAL units that include PPS

[0204] -APS (Adaptation Parameter Set) NAL unit: Type for NAL units that include APS

[0205] -PH(Picture header) NAL unit:Type for NAL unit including PH

[0206] The NAL unit types described above have syntax information for the NAL unit type, and this syntax information can be stored in the NAL unit header and signaled. For example, the syntax information may be nal_unit_type, and the NAL unit type can be identified by the value of nal_unit_type.

[0207] On the other hand, as mentioned above, a single picture can contain multiple slices, and a single slice can contain a slice header and slice data. In this case, a single picture header can be added to each of the multiple slices (slice headers and slice data sets) within a single picture. The picture header (picture header syntax) can contain information / parameters that are commonly applicable to the picture.

[0208] The slice header (slice header syntax) may include information / parameters that are commonly applicable to the slice. The APS (APS syntax) or PPS (PPS syntax) may include information / parameters that are commonly applicable to one or more slices or pictures. The SPS (SPS syntax) may include information / parameters that are commonly applicable to one or more sequences. The VPS (VPS syntax) may include information / parameters that are commonly applicable to multiple layers. The DCI (DCI syntax) may include information / parameters that are commonly applicable to video in general. The DCI may include information / parameters related to decoding capability. In this disclosure, High-level syntax (HLS) may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DCI syntax, picture header syntax, and slice header syntax. On the other hand, in this disclosure, low-level syntax (LLS) may include, for example, slice data syntax, CTU syntax, coding unit syntax, and transformation unit syntax.

[0209] In this disclosure, the image / video information encoded from the encoding device to the decoding device and signaled in bitstream format includes not only partitioning-related information within the picture, intra / inter prediction information, residual information, and in-loop filtering information, but also information from the slice header, the picture header, the APS, the PPS, the SPS, the VPS, and / or the DCI. Furthermore, the image / video information may further include general constraint information and / or information from the NAL unit header.

[0210] Picture partitioning using sub-pictures, slices, and tiles

[0211] A single picture can be divided into at least one tile row and at least one tile column. A single tile consists of a sequence of CTUs and can cover a rectangular area of ​​a single picture.

[0212] A slice can consist of an integer number of consecutive complete CTU rows or an integer number of complete tiles within a single picture.

[0213] Two modes can be supported for slicing. One can be called raster-scan slice mode, and the other can be called rectangular slice mode. In raster-scan slice mode, a slice can contain a complete sequence of tiles that exist in tile raster scan order within a single picture. In rectangular slice mode, a slice can contain multiple complete tiles assembled to form a rectangular region of a picture, or multiple consecutive complete rows of a single tile assembled to form a rectangular region of a picture. Tiles within a rectangular slice can be scanned in tile raster scan order within the rectangular region corresponding to the slice. A subpicture can contain at least one slice assembled to cover a rectangular region of a picture.

[0214] To explain the picture division relationships in more detail, refer to Figures 21 to 24. Figures 21 to 24 show examples of pictures divided using tiles, slices, and subpictures. Figure 21 shows an example of a picture divided into 12 tiles and 3 raster scan slices. Figure 22 shows an example of a picture divided into 24 tiles (6 tile columns and 4 tile rows) and 9 square slices. Figure 23 shows an example of a picture divided into 4 tiles (2 tile columns and 2 tile rows) and 4 square slices.

[0215] Figure 24 shows an example of a picture being divided into subpictures. In Figure 24, the picture is divided into 12 left-hand tiles, each covering one slice consisting of a 4x4 CTU, and 6 right-hand tiles, each covering two vertically joined slices consisting of 2x2 CTUs. As a result, one picture is divided into 24 slices and 24 subpictures, each with a different area. In the example in Figure 24, individual slices correspond to individual subpictures.

[0216] Overview of In-Loop Filtering

[0217] An in-loop filtering procedure can be performed on the reconstructed picture generated by the procedure described above. A modified reconstructed picture can be generated through the in-loop filtering procedure, and the decoder can output the modified reconstructed picture as a decoded picture. Furthermore, it can be stored in the decoded picture buffer or memory of the encoding / decoder and used as a reference picture in the inter-prediction procedure when encoding / decoding subsequent pictures. The in-loop filtering procedure may include, as described above, a deblocking filtering procedure, an SAO (sample adaptive offset) procedure, and / or an ALF (adaptive loop filter) procedure. In this case, one or part of the deblocking filtering procedure, SAO (sample adaptive offset) procedure, ALF (adaptive loop filter) procedure, and bi-lateral filter procedure may be applied sequentially, or all of them may be applied sequentially. For example, after the deblocking filtering procedure is applied to the reconstructed picture, the SAO procedure may be performed. Or, for example, after the deblocking filtering procedure is applied to the reconstructed picture, the ALF procedure may be performed. This can also be done in the encoding device.

[0218] Deblocking filtering is a filtering technique that removes distortions that occur at the boundaries between blocks in a restored picture. The deblocking filtering procedure can be performed, for example, by deriving a target boundary from the restored picture, determining a boundary strength (bS) for the target boundary, and performing deblocking filtering on the target boundary based on the bS. The bS can be determined based on the prediction modes of two blocks adjacent to the target boundary, the difference in motion vectors, whether the reference picture is identical, and the presence or absence of a non-zero effectiveness factor.

[0219] SAO is a method for compensating for the offset difference between a restored picture and the original picture on a sample-by-sample basis, and can be applied based on types such as Band Offset and Edge Offset. According to SAO, each SAO type can classify samples into different categories, and an offset value can be added to each sample based on the category. Filtering information for SAO can include information on whether SAO is applied, SAO type information, SAO offset value information, etc. SAO can also be applied to the restored picture after the deblocking filtering has been applied.

[0220] Adaptive Loop Filtering (ALF) is a technique that filters a reconstructed picture on a sample-by-sample basis based on filter coefficients determined by the filter's shape. The encoding device can determine whether to apply ALF, the ALF shape, and / or ALF filtering coefficients by comparing the reconstructed picture with the original picture, and can signal this to the decoding device. In other words, filtering information for ALF can include information on whether to apply ALF, ALF filter shape information, ALF filtering coefficient information, etc. ALF can also be applied to the reconstructed picture after the deblocking filtering has been applied.

[0221] Signaling with or without in-loop filtering

[0222] As previously stated, HLS can be encoded and / or signaled for video and / or image coding. As previously stated, video / image information in this specification can be included in HLS. And image / video coding methods can be performed based on such image / video information.

[0223] In one embodiment, a single picture can be partitioned into sub-pictures, slices, and / or tiles. The application of filtering to the boundaries of these sub-pictures can then be signaled. For example, a single picture can be divided into multiple tiles. In this case, in-loop filtering across the tile boundaries can be performed. Alternatively, a single picture can be divided into multiple slices. In this case, in-loop filtering across the slice boundaries can be performed. In this case, whether or not in-loop filtering across the tile boundaries and / or slice boundaries is performed can be signaled via HLS.

[0224] Figure 25 shows an example of PPS syntax for signaling whether in-loop filtering is performed at tile boundaries and / or slice boundaries. In the syntax of Figure 25, the syntax element no_pic_partition_flag2510 can indicate whether picture partitioning is applied to individual pictures that reference the PPS. For example, a first value of no_pic_partition_flag2510 (e.g., 0) can indicate that individual pictures that reference the PPS can be divided into more than one tile or slice. A second value of no_pic_partition_flag2510 (e.g., 1) can indicate that picture partitioning is not performed on individual pictures that reference the PPS. On the other hand, no_pic_partition_flag2510 can be restricted to having the same value for all PPS present in a single sequence.

[0225] In one embodiment, the syntax element no_pic_partition_flag2510 can be used as a condition to signal information for tile and / or slice partitioning when a picture is partitioned into more than one tile and / or slice. This can be included in the syntax as a conditional statement, such as the drawing reference numeral 2520 in Figure 25. For example, the encoding device can use no_pic_partition_flag2510 to signal to the decoding device whether or not information regarding tile and / or slice partitioning is included in the bitstream. The decoding device can then choose not to parse the tile and / or slice partitioning information from the bitstream if the value of no_pic_partition_flag2510 is 1. If the value of no_pic_partition_flag2510 is 0, the decoding device can parse the tile and / or slice partitioning information from the bitstream according to additional information.

[0226] To achieve the above, if no_pic_partition_flag2510 indicates that the picture can be divided into tiles or slices, the PPS syntax can achieve the following syntax elements being obtained from the bitstream, as shown in the example in Figure 25.

[0227] The syntax element pps_log2_ctu_size_minus5 can indicate the lumacoding tree block size of an individual CTU. More specifically, the encoding device can determine the value of pps_log2_ctu_size_minus5 by subtracting 5 from the lumacoding tree block size of the individual CTU. The decoding device can determine the lumacoding tree block size of an individual CTU by adding 5 to pps_log2_ctu_size_minus5.

[0228] The syntax element num_exp_tile_columns_minus1 can indicate the number of width values ​​in a tile column that is explicitly signaled. For example, a decoder can determine the number of width values ​​in a tile column that is explicitly signaled by adding 1 to num_exp_tile_columns_minus1. The value of num_exp_tile_columns_minus1 can range from 0 to PicWidthInCtbY-1, where PicWidthInCtbY can indicate the width of a picture expressed in units of the luma coding block width. On the other hand, if the value of no_pic_partition_flag is 1, the value of num_exp_tile_columns_minus1 can be induced to be 0.

[0229] The syntax element num_exp_tile_rows_minus1 can indicate the number of tile row height values ​​that are explicitly signaled. For example, a decoder can determine the number of tile row height values ​​that are explicitly signaled by adding 1 to num_exp_tile_rows_minus1. The value of num_exp_tile_rows_minus1 can range from 0 to PicHeightInCtbY-1. PicHeightInCtbY can indicate the height of a picture expressed in Lumacoding block height units. If the value of no_pic_partition_flag is 1, the value of num_exp_tile_rows_minus1 can be induced to be 0.

[0230] The syntax element tile_column_width_minus1[i] can indicate the width of the i-th tile column of a picture referencing a PPS. For example, a decoder might determine the width of the i-th tile column as tile_column_width_minus1[i] plus 1. The syntax element tile_column_width_minus1[i] can be obtained from a bitstream based on the value of num_exp_tile_columns_minus1, as shown in the syntax in Figure 25.

[0231] The syntax element tile_row_height_minus1[i] can indicate the height of the i-th tile row of the picture referencing the PPS. For example, a decoder might determine the height of the i-th tile row as tile_row_height_minus1[i] plus 1. The syntax element tile_row_height_minus1[i] can be obtained from the bitstream based on the value of num_exp_tile_rows_minus1, as shown in the syntax in Figure 25.

[0232] On the other hand, the variable NumTilesInPic can be calculated based on the values ​​of num_exp_tile_columns_minus1 and num_exp_tile_rows_minus1. In one embodiment, the decoding device can determine the value of the variable NumTilesInPic, which indicates the number of tiles in the picture that reference the PPS, as the value of (num_exp_tile_columns_minus1+1)*(num_exp_tile_rows_minus1+1).

[0233] If the value of NumTilesInPic is greater than 1, the syntax element rect_slice_flag can be obtained from the bitstream. For example, if a picture is divided into two or more tiles, the syntax element rect_slice_flag can be obtained.

[0234] The syntax element rect_slice_flag can indicate whether raster-scan slice mode or rectangular scan slice mode is applied to individual pictures referencing the PPS. For example, a first value of rect_slice_flag (e.g., 0) can indicate that raster-scan slice mode is applied to individual pictures referencing the PPS, in which case signaling of the slice layout is omitted. A second value of rect_slice_flag (e.g., 1) can indicate that rectangular slice mode is used for individual pictures referencing the PPS. In this case, the slice layout can be signaled via the PPS as described below. If rect_slice_flag is not signaled, the decoder can induce a value of rect_slice_flag to 1.

[0235] If the value of the syntax element rect_slice_flag is 1, the syntax element single_slice_per_subpic_flag can be signaled. The first value of single_slice_per_subpic_flag (e.g., 0) can indicate that an individual subpicture can consist of more than one rectangular slice. The second value of single_slice_per_subpic_flag (e.g., 1) can indicate that an individual subpicture consists of only one rectangular slice.

[0236] On the other hand, if the value of rect_slice_flag is 1 and the value of single_slice_per_subpic_flag is 0, the syntax element num_slices_in_pic_minus1 can be signaled. For example, if the picture is divided into two or more rectangular slices, the syntax element num_slices_in_pic_minus1, which indicates the number of rectangular slices in the individual picture referencing the PPS, can be signaled to signal the layout of the rectangular slices. For example, a decoder can determine the number of rectangular slices in the picture by adding 1 to num_slices_in_pic_minus1. The value of num_slices_in_pic_minus1 can range from 0 to MaxSlicePerAu-1. The variable MaxSlicePerAu can indicate the maximum number of slices allowed per access unit, for example, the maximum number of slices currently allowed in the picture.

[0237] Based on the value of num_slices_in_pic_minus1, the syntax elements tile_idx_delta_present_flag, slice_width_in_tiles_minus1, slice_height_in_tiles_minus1, num_exp_slices_in_tile, exp_slice_height_in_ctus_minus1, and tile_idx_delta can be obtained as shown in the syntax of Figure 25. Here, the syntax element tile_idx_delta_present_flag can indicate whether the syntax element tile_idx_delta, which is used as an index to identify the rectangular slices in the picture, can be obtained from the bitstream. The value of the syntax element slice_width_in_tiles_minus1[i] plus 1 can represent the width of the i-th rectangular slice in tile columns. The value of the syntax element slice_height_in_tiles_minus1[i] plus 1 can represent the height of the i-th rectangular slice in tile rows.

[0238] The syntax element num_exp_slices_in_tile[i] can represent the number of slice heights explicitly provided for a slice in the tile containing the i-th slice. The value of the syntax element exp_slice_height_in_ctus_minus1[j] plus 1 can represent the height of the j-th rectangular slice in the tile containing the i-th slice, and the unit can be the unit of the CTU row. The syntax element tile_idx_delta[i] can represent the difference between the index of the tile containing the 1st CTU in the i-th rectangular slice and the index of the tile containing the 1st CTU in the i+1-th rectangular slice.

[0239] The syntax element loop_filter_across_tiles_enabled_flag2530 can indicate whether filtering is performed across the boundaries of tiles within a picture that references a PPS. For example, the first value of loop_filter_across_tiles_enabled_flag (e.g., 0) indicates that in-loop filtering is not performed across the boundaries of tiles within a picture that references a PPS containing this syntax element. The second value of loop_filter_across_tiles_enabled_flag (e.g., 1) indicates that in-loop filtering is performed across the boundaries of tiles within a picture that references a PPS containing this syntax element.

[0240] Here, the in-loop filtering operation may include deblocking filtering, SAO (Sample Adaptive Offset) filtering, and / or ALF (Adaptive Loop Filter). If the value of loop_filter_across_tiles_enabled_flag is not obtained from the bitstream (e.g., not provided), the value of the syntax element may be induced to a second value (e.g., 1). On the other hand, in other embodiments, if the value of loop_filter_across_tiles_enabled_flag is not obtained from the bitstream (e.g., not provided), the value of the syntax element may also be induced to a first value (e.g., 0).

[0241] The syntax element loop_filter_across_slices_enabled_flag2540 can indicate whether filtering is performed across the boundaries of slices within a picture that references a PPS. For example, the first value of loop_filter_across_slices_enabled_flag (e.g., 0) indicates that in-loop filtering is not performed across the boundaries of slices within a picture that references a PPS containing this syntax element.

[0242] The second value of loop_filter_across_slices_enabled_flag (for example, 1) can indicate that in-loop filtering can be performed across the boundaries of slices in a picture that references the PPS containing this syntax element.

[0243] Here, the in-loop filtering operation may include deblocking filtering, SAO filtering, and / or ALF, as described above. If the value of loop_filter_across_slices_enabled_flag is not obtained from the bitstream (e.g., not provided), the value of the syntax element may be induced to a first value (e.g., 0).

[0244] In one embodiment, each picture can be partitioned on a tile basis. In such a case, there can be two or more tiles within a single picture, and whether or not an in-loop filter is applied to the boundary portion of each tile area can be determined by loop_filter_across_tiles_enabled_flag2530, which is signaled from the individual PPS.

[0245] In the example in Figure 25, if there are two or more tiles in a single picture (for example, no_pic_partition_flag==0), the value of loop_filter_across_tiles_enabled_flag is always signaled to determine whether or not to apply the in-loop filter in the boundary region. However, considering that this flag is also signaled when there are many tiles in a single picture, it can be improved to signal the flag considering the number of tiles present in order to reduce the amount of bits transmitted.

[0246] For example, in the example in Figure 25, if the value of no_pic_partition_flag is a first value (e.g., 0), this means that the picture is currently divided into more than one part in tile or slice units. In one embodiment, if the value of no_pic_partition_flag is a first value (e.g., 0), the value of loop_filter_across_tiles_enabled_flag can always be signaled to determine whether to apply in-loop filtering in the tile boundary region, in that there can be two or more tiles within a single picture. However, if the value of no_pic_partition_flag is a first value (e.g., 0), it also includes cases where a single picture is not divided into tiles, but only into slices. This means that loop_filter_across_tiles_enabled_flag can also be signaled even when there are no tiles within a single picture, but only many slices. Taking these points into consideration, it is possible to improve the signaling of the flag to take into account the number of tiles present in order to reduce the amount of bits transmitted.

[0247] Figure 26 shows an example of syntax that signals the syntax element loop_filter_across_tiles_enabled_flag considering the number of tiles in order to solve the above problem. As shown in Figure 26, the syntax element loop_filter_across_tiles_enabled_flag2620 can be signaled based on the number of tiles belonging to the picture (e.g., NumTilesInPic). For example, as shown in Figure 26, loop_filter_across_tiles_enabled_flag2620 can be signaled via a bitstream only if the number of tiles belonging to the picture is greater than 1 (2610). This allows the encoder to encode loop_filter_across_tiles_enabled_flag2620 as a bitstream only if the number of tiles belonging to the picture is greater than 1 (2610), and the decoder can obtain loop_filter_across_tiles_enabled_flag2620 from the bitstream.

[0248] On the other hand, in the embodiment shown in Figure 25 above, each picture can be partitioned in slice units. In such a case, there can be two or more slices within a single picture, and whether or not an in-loop filter is applied to the boundary portion of each slice region can be determined by the loop_filter_across_slices_enabled_flag signaled from the individual PPS. More specifically, if there are two or more slices within a single picture (for example, no_pic_partition_flag==0), the value of loop_filter_across_slices_enabled_flag is always signaled to determine whether or not an in-loop filter is applied to the boundary region. However, considering that this flag is also signaled when there are many slices within a single picture, it is possible to improve the system so that the flag is signaled considering the number of slices present in order to reduce the amount of bits transmitted.

[0249] In one embodiment, even when the value of no_pic_partition_flag is a first value (e.g., 0), the picture referencing the PPS may be divided only into tiles and not into slices. However, in the example in Figure 25, even when the picture is divided only into tiles and not into slices, the value of loop_filter_across_slices_enabled_flag is always signaled. Considering that loop_filter_across_slices_enabled_flag is signaled even when there are many tiles within a single picture, the embodiment in Figure 25 can be improved to signal the flag considering the number of slices present in order to reduce the amount of bits transmitted.

[0250] Figure 27 shows an example of syntax that signals the syntax element loop_filter_across_slices_enabled_flag considering the number of slices in order to solve the above problem. As shown in Figure 27, the syntax element loop_filter_across_slices_enabled_flag2720 can be signaled based on the number of slices belonging to the picture (e.g., num_slices_in_pic_minus1). For example, as shown in Figure 27, loop_filter_across_slices_enabled_flag2720 can be signaled via the bitstream only if the number of slices belonging to the picture is greater than 1 (2710). As a result, only if the number of slices belonging to the picture is greater than 1 (2710), the encoder can encode loop_filter_across_slices_enabled_flag2720 and generate a bitstream, and the decoder can obtain loop_filter_across_slices_enabled_flag2720 from the bitstream.

[0251] On the other hand, num_slices_in_pic_minus1 is a syntax element that signals the number of rectangular slices in an individual picture that references PPS to signal the layout of rectangular slices when the picture is divided into two or more rectangular slices, and when raster-scan slice mode is applied to an individual picture, the individual picture can still be divided into multiple slices. This allows the syntax to be modified so that loop_filter_across_slices_enabled_flag2720 is signaled even when the value of rect_slice_flag is a first value (e.g., 0) that represents raster-scan slice mode.

[0252] Furthermore, if the value of single_slice_per_subpic_flag is a secondary value (e.g., 1), the picture can be divided into multiple subpictures, meaning that a single picture can consist of multiple slices. This allows the syntax to be modified so that loop_filter_across_slices_enabled_flag2720 is signaled even when the value of single_slice_per_subpic_flag is a secondary value (e.g., 1).

[0253] Thus, if the value of num_slices_in_pic_minus1 is not obtained from the bitstream, it is not signaled whether the picture is currently divided into several slices. Therefore, the syntax can be modified so that loop_filter_across_slices_enabled_flag2720 is signaled even if the value of num_slices_in_pic_minus1 is not obtained from the bitstream. For example, in the example in Figure 27, the conditions for obtaining num_slices_in_pic_minus1 from the bitstream are that the value of rect_slice_flag is 1 and the value of single_slice_per_subpic_flag is 0. In this respect, if the value of rect_slice_flag is 0 or the value of single_slice_per_subpic_flag is 1, the value of loop_filter_across_slices_enabled_flag can be obtained from the bitstream regardless of whether the value of num_slices_in_pic_minus1 is greater than 1. For this type of processing, the PPS syntax can be modified beforehand as shown in Figure 28.

[0254] Figure 28 shows an example of the PPS syntax to which the signaling of loop_filter_across_tiles_enabled_flag and loop_filter_across_slices_enabled_flag, as described with reference to Figures 25 to 27, is applied. In the example in Figure 28, some of the names of the syntax elements described with reference to Figures 25 and 28 are named with pps_ appended. For example, the aforementioned syntax no_pic_partition_flag is named pps_no_pic_partition_flag.

[0255] Referring to Figure 28, if pps_no_pic_pration_flag has a value (e.g., 0) indicating that the picture can be divided into two or more parts of at least one type of tile or slice, then the number of tiles into which the picture is currently divided can be signaled using the syntax elements pps_num_exp_tile_columns_minus1, which indicates how many tile columns the picture referencing the PPS has, and pps_num_exp_tile_rows_minus1, which indicates how many tile rows the picture referencing the PPS has. The number of tiles currently contained in the picture can then be calculated as (pps_num_exp_tile_columns_minus1+1)*(pps_num_exp_tile_rows_minus1+1) and recorded in the variable NumTileInPic.

[0256] As mentioned above, the pps_loop_filter_across_tiles_enabled_flag and pps_rect_slice_flag syntax elements, which indicate whether filtering can be applied across tile boundaries, can only be obtained if the value of NumTileInPic is greater than 1, that is, if there are currently more than one tile in the picture. If the value of pps_rect_slice_flag is 1, then pps_single_slice_per_subpic_flag can be obtained from the bitstream. If the value of pps_rect_slice_flag is 1 and the value of pps_single_slice_per_subpic_flag is 0, then the pps_num_slices_in_pic_minus1 syntax element can be obtained from the bitstream. Then, if the value of pps_rect_slice_flag is 0, or the value of pps_single_slice_per_subpic_flag is 1, or the value of pps_num_slices_in_pic_minus1 is greater than 1, the value of pps_loop_filter_across_slices_enabled_flag can be obtained from the bitstream.

[0257] Encoding and Decoding Methods

[0258] The following describes an image encoding method and an image decoding method performed by an image encoding device and an image decoding device according to one embodiment.

[0259] First, the operation of the decoding device will be explained. An image decoding device according to one embodiment includes a memory and a processor, and the decoding device can perform decoding by the operation of the processor. Figure 29 shows the decoding method of the decoding device according to one embodiment.

[0260] A decoding device according to one embodiment can determine the number of tiles in the current picture (e.g., NumTilesInPic) based on whether the division of the current picture is not restricted (S2910). For example, the decoding device can obtain a partition restriction flag (e.g., no_pic_partition_flag) from the bitstream indicating whether the division of the current picture is restricted, and can determine whether the division of the current picture is not restricted based on the partition restriction flag.

[0261] Next, the decoding device can obtain a first flag (e.g., loop_filter_across_tiles_enabled_flag) from the bitstream indicating whether filtering is available for tile boundaries, based on the fact that there are currently multiple tiles in the picture (S2920). Here, the number of tiles in the picture can be determined based on tile count information indicating the number of tiles that divide the picture. Here, tile count information can be obtained from the bitstream based on the fact that the division of the picture is not restricted. The tile count information may include information indicating the number of tile columns in the picture (e.g., num_exp_tile_columns_minus1) and information indicating the number of tile rows in the picture (e.g., num_exp_tile_rows_minus1).

[0262] Next, the decoding device can decide whether or not to perform filtering on the boundaries of tiles currently belonging to the picture, based on the value of the first flag (S2930). Here, the type of filtering can be any one of the deblocking filter, SAO filter, and ALF filter, as described above. For example, if the first flag indicates that filtering is not available, then none of the filters used to decode the image among the deblocking filter, SAO filter, and ALF filter may be applied to the boundaries of the tile.

[0263] Furthermore, the decoding device can obtain a second flag (e.g., loop_filter_across_slices_enabled_flag) from the bitstream indicating whether filtering is available for the slice boundaries, based on the fact that the splitting of the picture is currently not restricted (S2940).

[0264] For example, the decoding device can obtain information about the slices that make up the picture from the bitstream, based on the fact that the division of the picture is currently not restricted. Then, based on the fact that the information about the slices does not indicate that the picture consists of a single slice, the decoding device can obtain a second flag from the bitstream.

[0265] Alternatively, the decoder can obtain a second flag from the bitstream based on slice information indicating that rectangular slice mode is not applied to the picture (e.g., rect_slice_flag==0). Alternatively, the decoder can obtain a second flag from the bitstream based on slice information indicating that the picture's subpictures consist of only one rectangular slice (e.g., rect_slice_flag==0 or pps_single_slice_per_subpic_flag==1).

[0266] Alternatively, the decoding device can also obtain the second flag from the bitstream based on the information regarding the slice indicating that the number of slices in the current picture is plural (e.g., num_slices_in_pic_minus1>0). For example, based on the fact that the division of the current picture is not restricted, it is determined whether the slices constituting the picture are rectangular slices, and based on the fact that the slices constituting the picture are rectangular slices, it is determined whether the subpicture of the picture is composed of only one rectangular slice, and based on the fact that the subpicture of the picture is composed of a plurality of rectangular slices more than one, information indicating the number of slices in the current picture is obtained from the bitstream, and based on the information indicating the number of slices in the current picture, it can be determined whether the number of slices in the current picture is plural.

[0267] Then, the decoding device can determine whether to perform filtering on the boundaries of the slices belonging to the current picture based on the value of the second flag (S2950). For example, when the second flag indicates that filtering is not available, for the boundary of the slice, it can be that the filter used for decoding the image among the deblocking filter, SAO filter, and ALF filter is not applied.

[0268] Next, the operation of the encoding device will be described. An image encoding device according to an embodiment includes a memory and a processor, and the encoding device can perform encoding in a manner corresponding to the decoding of the decoding device by the operation of the processor. For example, as shown in FIG. 30, the encoding device can determine the number of tiles in the current picture (e.g., NumTilesInPic) based on the fact that the partitioning of the current picture is not restricted (S3010). Next, based on the fact that the number of tiles in the current picture is plural, the encoding device can determine the value of a first flag (e.g., loop_filter_across_tiles_enabled_flag) indicating whether filtering is available for the tile boundaries (S3020). On the other hand, based on the fact that the partitioning of the current picture is not restricted, the encoding device can further determine whether the current picture is composed of one slice (S3030). Then, based on the fact that the current picture is not composed of one slice, the encoding device can determine the value of a second flag (e.g., loop_filter_across_slices_enabled_flag) indicating whether filtering is available for the slice boundaries (S3040).

[0269] Next, the encoding device can generate a bitstream including at least one of the first flag and the second flag, or not including either of them (S3050). For example, based on the fact that the number of tiles in the picture is not plural and the current picture is composed of one slice, the encoding device may not need to determine the values of both the first flag and the second flag, and can generate a bitstream not including the first flag and the second flag. In addition to this, the value of a partitioning restriction flag (e.g., no_pic_partition_flag) can be set according to whether there is a partitioning restriction in the current picture, and the partitioning restriction flag can also be included in the bitstream.

[0270] As mentioned above, the no_pic_partition_flag is used in the encoding and decoding methods to signal to the no_pic_partition_flag whether or not the picture is currently divided into tiles and / or slices. Then, information on the division of tiles and the number of slices are signaled accordingly. In this respect, if the signaling of loop_filter_across_tiles_enabled_flag and loop_filter_across_slices_enabled_flag is determined solely based on the value of no_pic_partition_flag, loop_filter_across_tiles_enabled_flag and loop_filter_across_slices_enabled_flag will be unnecessarily signaled if the picture is currently divided into slices only or tiles only.

[0271] Furthermore, in order to reduce the signaling of loop_filter_across_tiles_enabled_flag and loop_filter_across_slices_enabled_flag, separately signaling a flag indicating whether the current picture is divided into tiles and a flag indicating whether the current picture is divided into slices, along with no_pic_partition_flag, does not help in terms of bit reduction.

[0272] However, as described herein, a configuration that determines how to signal loop_filter_across_tiles_enabled_flag and loop_filter_across_slices_enabled_flag based on the number of tiles and slices that divide the picture, along with no_pic_partition_flag, allows for determining how to signal loop_filter_across_tiles_enabled_flag and loop_filter_across_slices_enabled_flag from tile and slice parsing information without signaling additional flags. Thus, the technical idea described herein can reduce the frequency with which these flags are generated in the bitstream in encoding / decoding environments where pictures can currently be divided into tiles and / or slices, thereby reducing the size of the bitstream.

[0273] Application Examples

[0274] The exemplary methods in this disclosure are presented as a series of actions for clarity of explanation, but this is not intended to restrict the order in which the steps are performed, and each step may be performed simultaneously or in a different order, if necessary. To implement the methods according to this disclosure, the exemplary steps may be further varied, including the remaining steps with some exceptions, or including additional steps with some exceptions.

[0275] In this disclosure, an image encoding device or image decoding device that performs a predetermined operation (step) may perform an operation (step) to confirm the conditions or status of the execution of said operation (step). For example, if it is stated that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or image decoding device may perform an operation to confirm whether or not the predetermined condition is satisfied, and then perform the predetermined operation.

[0276] The various embodiments of this disclosure are not intended to list all possible combinations, but rather to illustrate representative aspects of this disclosure. The matters described in the various embodiments may be applied independently or in combination of two or more.

[0277] Furthermore, various embodiments of this disclosure can be implemented by hardware, firmware, software, or a combination thereof. In the case of hardware implementation, it can be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), general processors, controllers, microcontrollers, microprocessors, etc.

[0278] Furthermore, the image decoding and image encoding devices to which the embodiments of this disclosure are applied can be included in multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video conferencing equipment, real-time communication equipment such as video communications, mobile streaming equipment, storage media, camcorders, video-on-demand (VoD) service providers, over-the-top (OTT) video equipment, internet streaming service providers, 3D video equipment, image-phone video equipment, and medical video equipment, and can be used to process video signals or data signals. For example, over-the-top (OTT) video equipment can include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).

[0279] Figure 31 illustrates a content streaming system to which the embodiments of this disclosure can be applied.

[0280] As shown in Figure 31, a content streaming system to which an embodiment of the present disclosure is applied may broadly include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.

[0281] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and transmitting this bitstream to the streaming server. In other cases, if a multimedia input device such as a smartphone, camera, or video camera directly generates the bitstream, the encoding server can be omitted.

[0282] The bitstream can be generated by an image encoding method and / or image encoding apparatus to which an embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0283] The streaming server transmits multimedia data to the user's device based on the user's request via a web server, and the web server can act as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server can transmit multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server can play a role in controlling the commands and responses between the devices within the content streaming system.

[0284] The streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0285] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices such as smartwatches, smart glasses, HMDs (head-mounted displays), digital TVs, desktop computers, and digital signage.

[0286] Each server within the aforementioned content streaming system can be operated as a distributed server, in which case the data received from each server can be processed in a distributed manner.

[0287] The scope of this disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) that enable the operation of various embodiments to be performed on a device or computer, and non-transitory computer-readable medium on which such software or commands etc. are stored and can be executed on a device or computer. [Industrial applicability]

[0288] The embodiments described herein can be used for encoding / decoding images.

Claims

1. An image decoding method performed by an image decoding device, Based on the fact that there are currently no restrictions on the division of the picture, the step of determining the number of tiles in the current picture, Based on the fact that there are multiple tiles in the current picture, the steps include obtaining a first flag from the bitstream for filtering against tile boundaries, The steps include determining to perform filtering on the boundaries of the tiles belonging to the current picture based on the value of the first flag, Based on the fact that the division of the current picture is not restricted, the steps include obtaining a second flag from the bitstream for filtering against the slice boundaries, The step of determining to filter the boundaries of the slices belonging to the current picture based on the value of the second flag, Based on the fact that the division of the current picture is not restricted, first information relating to the slices constituting the current picture is obtained from the bitstream. Based on the fact that the first information relating to the slice does not indicate that the current picture consists of a single slice, the second flag is obtained from the bitstream. Based on second information regarding the slice indicating whether to apply rectangular slice mode to the current picture, the second flag is obtained from the bitstream. An image decoding method in which the second information is obtained from the bitstream based on the fact that the number of tiles in the current picture is multiple.

2. The division limit flag for the division of the current picture is obtained from the bitstream. The image decoding method according to claim 1, wherein it is determined that the division of the current picture is restricted based on the division restriction flag.

3. The image decoding method according to claim 1, wherein the number of tiles in the current picture is determined based on tile number information for the number of tiles to divide the current picture.

4. The image decoding method according to claim 3, wherein the tile count information includes information about the number of tile columns in the current picture and information about the number of tile rows in the current picture.

5. The image decoding method according to claim 3, wherein the tile count information is obtained from the bitstream based on the fact that the division of the current picture is not restricted.

6. The image decoding method according to claim 1, wherein the second flag is obtained from the bitstream based on third information relating to the slice indicating that the subpicture of the current picture consists of only one rectangular slice.

7. The image decoding method according to claim 1, wherein the second flag is obtained from the bitstream based on the first information relating to the slices indicating that the number of slices in the current picture is multiple.

8. Based on the fact that the division of the current picture is not restricted, it is determined whether the slices constituting the current picture are rectangular slices. Based on whether the slice constituting the current picture is a rectangular slice, it is determined whether the sub-picture of the current picture consists of only one rectangular slice. Based on the fact that the sub-picture of the current picture consists of more than one rectangular slice, information about the number of slices in the current picture is obtained from the bitstream. The image decoding method according to claim 7, wherein, based on information regarding the number of slices in the current picture, it is determined whether the number of slices in the current picture is multiple.

9. An image encoding method performed by an image encoding device, Based on the fact that there are currently no restrictions on the division of the picture, the step of determining the number of tiles in the current picture, Based on the fact that there are multiple tiles in the current picture, the step of determining the value of a first flag for filtering against tile boundaries, The steps include generating a bitstream that includes the first flag, Based on the fact that the division of the current picture is not restricted, the first step is to determine whether the current picture consists of a single slice. The step of determining the value of a second flag for filtering against slice boundaries, based on the fact that the current picture does not consist of a single slice, Based on the fact that the division of the current picture is not restricted, the first information relating to the slices constituting the current picture is included in the bitstream. Based on second information relating to the slice indicating whether to apply rectangular slice mode to the current picture, the second flag is included in the bitstream. An image encoding method in which the second information is included in the bitstream based on the fact that the number of tiles in the current picture is multiple.

10. A method for transmitting a bitstream, Based on the fact that there are currently no restrictions on the division of the picture, the step of determining the number of tiles in the current picture, Based on the fact that there are multiple tiles in the current picture, the step of determining the value of a first flag for filtering against tile boundaries, Based on the fact that the division of the current picture is not restricted, the first step is to determine whether the current picture consists of a single slice. Based on the fact that the current picture does not consist of a single slice, the step of determining the value of a second flag for filtering against the slice boundary, A step of generating the bitstream including the first flag, the second flag, and the first information, The step of transmitting the bitstream includes, Based on the fact that the division of the current picture is not restricted, the first information relating to the slices constituting the current picture is included in the bitstream. Based on second information relating to the slice indicating whether to apply rectangular slice mode to the current picture, the second flag is included in the bitstream. A method wherein the second information is included in the bitstream based on the fact that the number of tiles in the current picture is multiple.