Image encoding / decoding method and apparatus for selectively signaling filtering available information, and method for transmitting a bitstream
By selectively signaling the available information of filtering, determining whether the tiling boundary needs to be filtered, the problem of increasing the amount of information in high-resolution and high-quality image transmission is solved, improving encoding/decoding efficiency and reducing costs.
Patent Information
- Application Number
- CN202180014308.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-14
- Filing Date
- 2021-02-15
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-02-15
AI Technical Summary
The prior art has cost problems caused by the increase in the amount of information in the transmission and storage of high-resolution and high-quality images, and it is necessary to improve image encoding/decoding efficiency.
By selectively signaling the filtering available information, the number of tiles in the current screen is determined, and the boundary filtering flag is obtained from the bitstream, and whether to perform filtering on the tiles boundary is determined.
Improve image encoding/decoding efficiency and reduce transmission and storage costs.
Smart Images

Figure CN115088256B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly, to an image encoding / decoding method and apparatus for selectively signaling filtering availability information, and a method of transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure. Background Art
[0002] Recently, the demand for high-resolution and high-quality images (such as high-definition (HD) images and ultra-high-definition (UHD) images) is increasing in various fields. As the resolution and quality of image data improve, the amount of information or bits to be transmitted increases relatively compared to existing image data. The increase in the amount of information or bits to be transmitted results in an increase in transmission cost and storage cost.
[0003] Therefore, there is a need for efficient image compression techniques to effectively transmit, store, and reproduce information on high-resolution and high-quality images. Summary of the Invention
[0004] Technical Problem
[0005] An object of the present disclosure is to provide an image encoding / decoding method and apparatus having improved encoding / decoding efficiency.
[0006] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus capable of improving encoding / decoding efficiency by selectively signaling filtering availability information.
[0007] Another object of the present disclosure is to provide a method of transmitting a bitstream generated by an image encoding method or apparatus according to the present disclosure.
[0008] Another object of the present disclosure is to provide a recording medium storing a bitstream generated by an image encoding method or apparatus according to the present disclosure.
[0009] Another object of the present disclosure is to provide a recording medium storing a bitstream received, decoded, and used for reconstructing an image by an image decoding apparatus according to the present disclosure.
[0010] The technical problems solved by the present disclosure are not limited to the above technical problems, and other technical problems not described herein will be apparent to those skilled in the art from the following description.
[0011] Technical Solution
[0012] An image decoding method performed by an image decoding device according to an aspect of the present disclosure may include the following steps: determining the number of tiles in the current picture based on the unrestricted segmentation of the current picture; obtaining, from a bitstream, a first flag indicating whether filtering of the boundaries of the tiles is available based on the number of tiles in the current picture being multiple; and determining whether to perform filtering on the boundaries of the tiles belonging to the current picture based on the value of the first flag.
[0013] An image decoding device according to an aspect of the present disclosure may include a memory and at least one processor. The at least one processor may: determine the number of tiles in the current picture based on the unrestricted segmentation of the current picture; obtain, from a bitstream, a first flag indicating whether filtering of the boundaries of the tiles is available based on the number of tiles in the current picture being multiple; and determine whether to perform filtering on the boundaries of the tiles belonging to the current picture based on the value of the first flag.
[0014] An image encoding method performed by an image encoding device according to an aspect of the present disclosure may include the following steps: determining the number of tiles in the current picture based on the unrestricted segmentation of the current picture; determining the value of a first flag indicating whether filtering of the boundaries of the tiles is available based on the number of tiles in the current picture being multiple; and generating a bitstream including the first flag.
[0015] In addition, a transmission method according to an aspect of the present disclosure may transmit a bitstream generated by the image encoding device or method of the present disclosure.
[0016] Furthermore, a computer-readable recording medium according to an aspect of the present disclosure may store a bitstream generated by the image encoding device or image encoding method of the present disclosure.
[0017] The features briefly outlined above regarding the present disclosure are merely exemplary aspects of the following detailed description of the present disclosure and do not limit the scope of the present disclosure.
[0018] Advantageous Effects
[0019] According to the present disclosure, an image encoding / decoding method and device with improved encoding / decoding efficiency can be provided.
[0020] In addition, according to the present disclosure, an image encoding / decoding method and device capable of improving encoding / decoding efficiency by selectively signaling filtering availability information can be provided.
[0021] Furthermore, according to the present disclosure, a method of transmitting a bitstream generated by the image encoding method or device according to the present disclosure can be provided.
[0022] In addition, according to the present disclosure, a recording medium storing a bitstream generated by an image encoding method or apparatus according to the present disclosure can be provided.
[0023] In addition, according to the present disclosure, a recording medium storing a bitstream received, decoded, and used for reconstructing an image by an image decoding apparatus according to the present disclosure can be provided.
[0024] Those skilled in the art will understand that the effects achievable through the present disclosure are not limited to those specifically described above, and other advantages of the present disclosure will be more clearly understood from the detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a view schematically showing a video coding system to which the present disclosure is applicable.
[0026] Figure 2 is a view schematically showing an image encoding apparatus to which an embodiment of the present disclosure is applicable.
[0027] Figure 3 is a view schematically showing an image decoding apparatus to which an embodiment of the present disclosure is applicable.
[0028] Figure 4 is a view showing a segmentation structure of an image according to an embodiment.
[0029] Figure 5 is a view showing an embodiment of a segmentation type of a block according to a multi-type tree structure.
[0030] Figure 6 is a view showing a signaling mechanism of block partitioning information in a quadtree having a nested multi-type tree structure according to the present disclosure.
[0031] Figure 7 is a view showing an embodiment of dividing a CTU into a plurality of CUs.
[0032] Figure 8 is a view showing neighboring reference samples according to an embodiment.
[0033] Figures 9 to 10 is a view illustrating intra prediction according to an embodiment.
[0034] Figure 11 is a view illustrating an encoding method using inter prediction according to an embodiment.
[0035] Figure 12 is a view illustrating a decoding method using inter prediction according to an embodiment.
[0036] Figure 13It is a block diagram of CABAC according to an embodiment for encoding a syntax element.
[0037] Figures 14 to 17 It is a view exemplifying entropy encoding and entropy decoding according to an embodiment.
[0038] Figure 18 and Figure 19 It is a view showing an example of a picture decoding and encoding process according to an embodiment.
[0039] Figure 20 It is a view showing the layer structure of an encoded image.
[0040] Figures 21 to 24 It is a view showing an embodiment of dividing a picture using tiles, slices, and sub-pictures.
[0041] Figures 25 to 28 It is a view showing an individual embodiment of the syntax for a picture parameter set.
[0042] Figure 29 and Figure 30 It is a view showing an embodiment of a decoding method and an encoding method.
[0043] Figure 31 It is a view showing a content stream system to which an embodiment of the present disclosure is applicable. Detailed Embodiments
[0044] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so as to be easily implemented by those skilled in the art. However, the present disclosure can be implemented in various different forms and is not limited to the embodiments described herein.
[0045] When describing the present disclosure, if it is determined that a detailed description of a related known function or configuration makes the scope of the present disclosure unnecessarily ambiguous, its detailed description will be omitted. In the drawings, parts irrelevant to the description of the present disclosure are omitted, and similar reference numerals are given to similar parts.
[0046] In the present disclosure, when a component "connects", "couples", or "links" to another component, it may include not only a direct connection relationship but also an indirect connection relationship with an intermediate component. In addition, when a component "includes" or "has" other components, unless otherwise specified, it means that other components may also be included, rather than excluding other components.
[0047] In the present disclosure, terms such as first and second may be used only for the purpose of differentiating one component from other components, and do not limit the order or importance of the components, unless otherwise specified. Accordingly, within the scope of the present disclosure, the first component in one embodiment may be referred to as the second component in another embodiment, and similarly, the second component in one embodiment may be referred to as the first component in another embodiment.
[0048] In the present disclosure, components that are mutually distinguishable are intended to clearly describe each feature, and do not mean that the components must be separated. That is, multiple components may be integrated and implemented in one hardware or software unit, or one component may be distributed and implemented in multiple hardware or software units. Therefore, even if not otherwise specified, embodiments in which components are integrated or components are distributed are also included within the scope of the present disclosure.
[0049] In the present disclosure, the components described in each embodiment are not necessarily essential components, and some components may be optional components. Therefore, embodiments consisting of a subset of the components described in the embodiments are also included within the scope of the present disclosure. In addition, embodiments that include other components in addition to the components described in the various embodiments are included within the scope of the present disclosure.
[0050] The present disclosure relates to the encoding and decoding of images. Unless redefined in the present disclosure, the terms used in the present disclosure may have the general meanings commonly used in the technical field to which the present disclosure pertains.
[0051] In the present disclosure, "video" may refer to a collection of a series of images over time. A "picture" generally refers to a unit representing an image within a specific time period, and a slice / tile is an encoding unit that forms part of a picture during encoding. A picture can consist of one or more slices / tiles. Additionally, a slice / tile can include one or more coding tree units (CTUs). A CTU can be divided into one or more coding units (CUs). A picture can consist of one or more slices / tiles. A tile is a rectangular area within a specific tile row and specific tile column in a picture and can be composed of multiple CTUs. A tile column can be defined as a rectangular area of CTUs, can have the same height as the picture, and can have a width specified by syntax elements signaled from a bitstream part such as a picture parameter set. A tile row can be defined as a rectangular area of CTUs, can have the same width as the picture, and can have a height specified by syntax elements signaled from a bitstream part such as a picture parameter set. Tile scanning is a certain consecutive sorting method for dividing the CTUs of a picture. Here, the CTUs can be sorted in sequence according to the CTU raster scan within a tile, and the tiles in a picture can be sorted in sequence according to the raster scan order of the tiles of the picture. A slice can contain an integer number of complete tiles or can contain an integer number of consecutive complete CTU rows within a tile of a picture. A slice can be included only in a single NAL unit.
[0052] A picture can consist of one or more tile groups. A tile group can include one or more tiles. A brick can indicate a rectangular area of CTU rows within a tile in a picture. A tile can include one or more bricks. A brick can represent a rectangular area of CTU rows in a tile. A tile can be divided into multiple bricks, and each brick can include one or more CTU rows belonging to the tile. A tile that is not divided into multiple bricks can also be regarded as a brick.
[0053] Furthermore, a picture can be divided into two or more sub-pictures. A sub-picture can be a rectangular area of one or more slices in a picture.
[0054] "Pixel" or "pel" can refer to the smallest unit that constitutes a picture (or image). Additionally, "sample" can be used as a term corresponding to a pixel. A sample generally can represent a pixel or the value of a pixel, or can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0055] In the present disclosure, a "unit" may represent a basic unit for image processing. The unit may include at least one of a specific area of a picture and information related to the area. A unit may include one luminance block and two chrominance (e.g., Cb and Cr) blocks. In some cases, the unit may be used interchangeably with terms such as "sample array", "block", or "area". In general, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients having M columns and N rows.
[0056] In the present disclosure, a "current block" may mean one of "current coding block", "current coding unit", "coding target block", "decoding target block", or "processing target block". When performing prediction, the "current block" may mean "current prediction block" or "prediction target block". When performing transform (inverse transform) / quantization (dequantization), the "current block" may mean "current transform block" or "transform target block". When performing filtering, the "current block" may mean "filtering target block".
[0057] Additionally, in the present disclosure, unless explicitly stated as a chrominance block, a "current block" may mean the luminance block of the "current block". The chrominance block of the "current block" may be expressed by including an explicit description of a chrominance block such as "chrominance block" or "current chrominance block".
[0058] In the present disclosure, the slashes " / " or "," should be interpreted as indicating "and / or". For example, the expressions "A / B" and "A, B" may mean "A and / or B". Additionally, "A / B / C" and "A / B / C" may mean at least one of "A, B, and / or C".
[0059] In the present disclosure, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" may include 1) only "A", 2) only "B", and / or 3) both "A and B". In other words, in the present disclosure, the term "or" should be interpreted as indicating "additionally or alternatively".
[0060] Overview of the video coding system
[0061] Figure 1 is a view showing a video coding system according to the present disclosure.
[0062] A video coding system according to an embodiment may include a source device 10 and a receiving device 20. The source device 10 may deliver encoded video and / or image information or data in the form of a file or a stream to the receiving device 20 via a digital storage medium or a network.
[0063] The source device 10 according to an embodiment may include a video source generator 11, an encoding device 12, and a transmitter 13. The receiving device 20 according to an embodiment may include a receiver 21, a decoding device 22, and a renderer 23. The encoding device 12 may be referred to as a video / image encoding device, and the decoding device 22 may be referred to as a video / image decoding device. The transmitter 13 may be included in the encoding device 12. The receiver 21 may be included in the decoding device 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.
[0064] The video source generator 11 may obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source generator 11 may include a video / image capturing device and / or a video / image generating device. The video / image capturing device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generating device may include, for example, a computer, a tablet computer, and a smart phone, and may generate (electronically) video / images. For example, virtual video / images may be generated by a computer or the like. In this case, the video / image capturing process may be replaced by a process of generating relevant data.
[0065] The encoding device 12 may encode the input video / images. For compression and encoding efficiency, the encoding device 12 may perform a series of processes such as prediction, transformation, and quantization. The encoding device 12 may output the encoded data (encoded video / image information) in the form of a bitstream.
[0066] The transmitter 13 may send the encoded video / image information or data output in the form of a bitstream to the receiver 21 of the receiving device 20 in the form of a file or a stream via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter 13 may include an element for generating a media file in a predetermined file format and may include an element for sending via a broadcast / communication network. The receiver 21 may extract / receive the bitstream from the storage medium or the network and send the bitstream to the decoding device 22.
[0067] The decoding device 22 may decode the video / images by performing a series of processes (such as dequantization, inverse transformation, and prediction) corresponding to the operations of the encoding device 12.
[0068] The renderer 23 may render the decoded video / images. The rendered video / images may be displayed via a display.
[0069] Overview of the image coding device
[0070] Figure 2is a view schematically showing an image encoding apparatus to which embodiments of the present disclosure are applicable.
[0071] As Figure 2 shown, the image encoding apparatus 100 may include an image splitter 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoder 190. The inter prediction unit 180 and the intra prediction unit 185 may be collectively referred to as a “prediction unit”. The transformer 120, the quantizer 130, the dequantizer 140, and the inverse transformer 150 may be included in a residual processor. The residual processor may further include the subtractor 115.
[0072] In some embodiments, all or at least some of the plurality of components configuring the image encoding apparatus 100 may be configured by one hardware component (e.g., an encoder or a processor). Additionally, the memory 170 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium.
[0073] The image splitter 110 may split an input image (or picture or frame) input to the image encoding apparatus 100 into one or more processing units. For example, the processing unit may be referred to as a coding unit (CU). The coding unit may be obtained by recursively splitting a coding tree unit (CTU) or a largest coding unit (LCU) according to a quadtree binary tree ternary tree (QT / BT / TT) structure. For example, one coding unit may be split into a plurality of coding units with a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the splitting of the coding unit, the quadtree structure may be applied first, and later the binary tree structure and / or the ternary tree structure may be applied. The encoding process according to the present disclosure may be performed based on the final coding unit that is no longer split. The largest coding unit may be used as the final coding unit, or the coding unit with a deeper depth obtained by splitting the largest coding unit may be used as the final coding unit. Here, the encoding process may include processes of prediction, transformation, and reconstruction to be described later. As another example, the processing unit of the encoding process may be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit may be divided or split from the final coding unit. The prediction unit may be a sample prediction unit, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0074] A prediction unit (inter-frame prediction unit 180 or intra-frame prediction unit 185) may perform prediction on a block to be processed (current block) and generate a prediction block including prediction samples of the current block. The prediction unit may determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. The prediction unit may generate various information related to the prediction of the current block and send the generated information to the entropy encoder 190. The information about the prediction may be encoded in the entropy encoder 190 and output in the form of a bitstream.
[0075] The intra-frame prediction unit 185 may predict the current block by referring to samples in the current picture. Depending on the intra-frame prediction mode and / or intra-frame prediction technique, the reference samples may be located in the neighbors of the current block or may be placed separately. The intra-frame prediction mode may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, the DC mode and the planar mode. Depending on the level of detail of the prediction direction, the directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes may be used according to the settings. The intra-frame prediction unit 185 may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0076] The inter-frame prediction unit 180 may derive a prediction block of a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to as a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block may be referred to as a collocated picture (colPic). For example, the inter-frame prediction unit 180 may configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or the reference picture index of the current block. The inter-frame prediction may be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter-frame prediction unit 180 may use the motion information of neighboring blocks as the motion information of the current block. In the case of the skip mode, different from the merge mode, the residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of a neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be signaled by encoding the motion vector difference and an indicator of the motion vector predictor. The motion vector difference may mean the difference between the motion vector of the current block and the motion vector predictor.
[0077] The prediction unit may generate a prediction signal based on various prediction methods and prediction techniques described below. For example, the prediction unit may not only apply intra-frame prediction or inter-frame prediction, but also apply both intra-frame prediction and inter-frame prediction simultaneously to predict the current block. The prediction method of applying both intra-frame prediction and inter-frame prediction simultaneously to predict the current block may be referred to as combined inter-frame and intra-frame prediction (CIIP). In addition, the prediction unit may perform intra-block copy (IBC) to predict the current block. Intra-block copy may be used for content image / video coding such as games, for example, screen content coding (SCC). IBC is a method of predicting the current picture using a previously reconstructed reference block in the current picture at a position separated from the current block by a predetermined distance. When IBC is applied, the position of the reference block in the current picture may be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction in the current picture, but may be performed similar to inter-frame prediction because the reference block is derived within the current picture. That is, IBC may use at least one of the inter-frame prediction techniques described in the present disclosure.
[0078] The prediction signal generated by the prediction unit can be used to generate a reconstructed signal or a residual signal. The subtractor 115 can generate a residual signal (residual block or residual sample array) by subtracting the prediction signal (prediction block or prediction sample array) output from the prediction unit from the input image signal (original block or original sample array). The generated residual signal can be sent to the transformer 120.
[0079] The transformer 120 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a karhunen-loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, the GBT refers to a transform obtained from a graph when the relationship information between pixels is represented by the graph. The CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. Additionally, the transform process can be applied to a square pixel block of the same size or can be applied to a block having a variable size rather than a square.
[0080] The quantizer 130 can quantize the transform coefficients and send them to the entropy encoder 190. The entropy encoder 190 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 130 can rearrange the block-type quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0081] The entropy encoder 190 can perform various coding methods, such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 190 can encode information required for video / image reconstruction other than the quantized transform coefficients (e.g., values of syntax elements, etc.) together or separately. The encoded information (e.g., encoded video / image information) can be sent or stored in the form of a bitstream in units of a network abstraction layer (NAL). The video / image information can also include information about various parameter sets, such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information can also include general constraint information. The information signaled, sent, and / or syntax elements described in the present disclosure can be encoded through the above coding process and included in the bitstream.
[0082] The bitstream can be sent over a network or stored in a digital storage medium. The network can include a broadcast network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that sends the signal output from the entropy encoder 190 and / or a storage unit (not shown) that stores the signal can be included as internal / external elements of the image encoding device 100. Alternatively, a transmitter can be provided as a component of the entropy encoder 190.
[0083] The quantized transform coefficients output from the quantizer 130 can be used to generate a residual signal. For example, the quantized transform coefficients can be dequantized and inverse-transformed by the dequantizer 140 and the inverse-transformer 150 to reconstruct the residual signal (residual block or residual samples).
[0084] The adder 155 adds the reconstructed residual signal to the prediction signal output from the inter-frame prediction unit 180 or the intra-frame prediction unit 185 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). If the block to be processed has no residual, such as in the case of applying the skip mode, the predicted block can be used as the reconstructed block. The adder 155 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, and can be used for inter-frame prediction of the next picture through filtering as described below.
[0085] The filter 160 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 160 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. The filter 160 can generate various information related to the filtering and send the generated information to the entropy encoder 190, as described later in the description of each filtering method. The information related to the filtering can be encoded by the entropy encoder 190 and output in the form of a bitstream.
[0086] The modified reconstructed picture sent to the memory 170 can be used as a reference picture in the inter-frame prediction unit 180. When inter-frame prediction is applied by the image encoding device 100, prediction mismatch between the image encoding device 100 and the image decoding device can be avoided and the coding efficiency can be improved.
[0087] The DPB of memory 170 may store the modified reconstructed picture to be used as a reference picture in the inter prediction unit 180. Memory 170 may store the motion information of the blocks for deriving (or encoding) the motion information in the current picture and / or the motion information of the blocks that have been reconstructed in the picture. The stored motion information may be sent to the inter prediction unit 180 and used as the motion information of spatially neighboring blocks or temporally neighboring blocks. Memory 170 may store the reconstructed samples of the reconstructed blocks in the current picture and may transfer the reconstructed samples to the intra prediction unit 185.
[0088] Overview of the image decoding device
[0089] Figure 3 is a view schematically showing an image decoding device to which an embodiment of the present disclosure is applicable.
[0090] As Figure 3 shown, the image decoding device 200 may include an entropy decoder 210, a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 may be collectively referred to as a "prediction unit". The dequantizer 220 and the inverse transformer 230 may be included in a residual processor.
[0091] According to an embodiment, all or at least some of the multiple components configuring the image decoding device 200 may be configured by hardware components (e.g., a decoder or a processor). Additionally, the memory 250 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium.
[0092] The image decoding device 200 that has received a bitstream including video / image information may reconstruct the image by performing processing corresponding to the processing performed by the Figure 2 image encoding device 100. For example, the image decoding device 200 may perform decoding using the processing units applied in the image encoding device. Thus, the decoding processing unit may be, for example, an encoding unit. The encoding unit may be obtained by splitting a coding tree unit or a maximum coding unit. The reconstructed image signal decoded and output by the image decoding device 200 may be reproduced by a reproduction device (not shown).
[0093] The image decoding device 200 may receive, in the form of a bitstream, from Figure 2The signal output by the image encoding device. The received signal can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can also include information about various parameter sets, such as adaptive parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). Additionally, the video / image information can also include general constraint information. The image decoding device can also decode the picture based on the information about the parameter set and / or the general constraint information. The information and / or syntax elements signaled / received described in this disclosure can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 210 decodes the information in the bitstream based on an encoding method such as exponential Golomb coding, CAVLC, or CABAC, and outputs the values of the syntax elements required for image reconstruction and the quantization values of the transformed coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive the bins corresponding to each syntax element in the bitstream, use the decoding target syntax element information, neighboring blocks, and the decoding information of the decoding target block or the information of the symbols / bins decoded in the previous stage to determine the context model, perform arithmetic decoding on the bins by predicting the occurrence probability of the bins according to the determined context model, and generate symbols corresponding to the values of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin. The information related to prediction among the information decoded by the entropy decoder 210 can be provided to the prediction units (inter-frame prediction unit 260 and intra-frame prediction unit 265), and the residual values (i.e., the quantized transform coefficients and related parameter information) for which entropy decoding is performed in the entropy decoder 210 can be input to the dequantizer 220. Additionally, the information about filtering among the information decoded by the entropy decoder 210 can be provided to the filter 240. Furthermore, the receiver (not shown) for receiving the signal output by the image encoding device can be further configured as an internal / external component of the image decoding device 200, or the receiver can be a component of the entropy decoder 210.
[0094] Furthermore, the image decoding device according to the present disclosure can be referred to as a video / image / picture decoding device. The image decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 210. The sample decoder can include at least one of the dequantizer 220, inverse transformer 230, adder 235, filter 240, memory 250, inter-frame prediction unit 260, or intra-frame prediction unit 265.
[0095] The dequantizer 220 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 220 may rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement may be performed based on the coefficient scan order executed in the image coding device. The dequantizer 220 may perform dequantization on the quantized transform coefficients by using a quantization parameter (e.g., quantization step information) and obtain the transform coefficients.
[0096] The inverse transformer 230 may perform an inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0097] The prediction unit may perform prediction on a current block and generate a prediction block including prediction samples of the current block. The prediction unit may determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 210, and may determine a specific intra / inter prediction mode (prediction technique).
[0098] Similar to that described for the prediction unit in the image coding apparatus 100, the prediction unit may generate a prediction signal based on various prediction methods (techniques) to be described later.
[0099] The intra prediction unit 265 may predict the current block by referring to samples in the current picture. The description of the intra prediction unit 185 is equally applicable to the intra prediction unit 265.
[0100] The inter prediction unit 260 may derive a prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information about an inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatially neighboring blocks present in the current picture and temporally neighboring blocks present in the reference picture. For example, the inter prediction unit 260 may configure a motion information candidate list based on neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. The inter prediction may be performed based on various prediction modes, and the information about prediction may include information indicating the inter prediction mode of the current block.
[0101] The adder 235 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter-frame prediction unit 260 and / or the intra-frame prediction unit 265). If there is no residual for the block to be processed, such as when the skip mode is applied, the predicted block can be used as the reconstructed block. The description of the adder 155 equally applies to the adder 235. The adder 235 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture and can be used for inter-frame prediction of the next picture through filtering as described below.
[0102] The filter 240 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 240 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 250, specifically, in the DPB of the memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc.
[0103] The (modified) reconstructed picture stored in the DPB of the memory 250 can be used as a reference picture in the inter-frame prediction unit 260. The memory 250 can store the motion information of the blocks for deriving (or decoding) the motion information in the current picture and / or the motion information of the blocks that have been reconstructed in the picture. The stored motion information can be sent to the inter-frame prediction unit 260 to be used as the motion information of spatially neighboring blocks or temporally neighboring blocks. The memory 250 can store the reconstructed samples of the reconstructed blocks in the current picture and send the reconstructed samples to the intra-frame prediction unit 265.
[0104] In the present disclosure, the embodiments described in the filter 160, the inter-frame prediction unit 180, and the intra-frame prediction unit 185 of the image encoding device 100 can be equally or correspondingly applied to the filter 240, the inter-frame prediction unit 260, and the intra-frame prediction unit 265 of the image decoding device 200.
[0105] Overview of image segmentation
[0106] The video / image encoding method according to the present disclosure can be performed based on the following image segmentation structure. Specifically, processes such as prediction, residual processing ((inverse) transformation, (de) quantization, etc.), syntax element encoding, and filtering, which will be described later, can be performed based on CTUs, CUs (and / or TUs, PUs) derived according to the image segmentation structure. The image can be segmented in block units, and the block segmentation process can be performed in the image segmenter 110 of the encoding device. The segmentation-related information can be encoded by the entropy encoder 190 and sent to the decoding device in the form of a bitstream. The entropy decoder 210 of the decoding device can derive the block segmentation structure of the current picture based on the segmentation-related information obtained from the bitstream, and based on this, a series of processes (such as prediction, residual processing, block / picture reconstruction, in-loop filtering, etc.) can be performed for image decoding.
[0107] The picture can be segmented into a sequence of coding tree units (CTUs). Figure 4 An example of a picture segmented into CTUs is shown. A CTU can correspond to a coding tree block (CTB). Alternatively, a CTU can include a coding tree block of luminance samples and two coding tree blocks of corresponding chrominance samples. For example, for a picture containing three sample arrays, a CTU can include an N×N block of luminance samples and two corresponding blocks of chrominance samples. The maximum allowable size of the CTU for encoding and prediction can be different from the maximum allowable size of the CTU for transformation. For example, even if the maximum size of the luminance transform block is 64×64, the maximum allowable size of the luminance block in the CTU can be 128×128.
[0108] Overview of CTU segmentation
[0109] As described above, coding units can be obtained by recursively segmenting coding tree units (CTUs) or largest coding units (LCUs) according to a quadtree / binary tree / trinary tree (QT / BT / TT) structure. For example, a CTU can be first segmented into a quadtree structure. Thereafter, the leaf nodes of the quadtree structure can be further segmented by a multi-type tree structure.
[0110] Segmentation according to the quadtree means that the current CU (or CTU) is equally segmented into four. By segmentation according to the quadtree, the current CU can be segmented into four CUs with the same width and the same height. When the current CU is no longer segmented into a quadtree structure, the current CU corresponds to the leaf node of the quadtree structure. The CU corresponding to the leaf node of the quadtree structure can no longer be segmented and can be used as the final coding unit described above. Alternatively, the CU corresponding to the leaf node of the quadtree structure can be further segmented by a multi-type tree structure.
[0111] Figure 5It is a view showing an embodiment of the split type of blocks according to a multi-type tree structure. The split according to the multi-type tree structure may include two types of divisions according to a binary tree structure and two types of divisions according to a ternary tree structure.
[0112] The two types of divisions according to the binary tree structure may include a vertical binary split (SPLIT_BT_VER) and a horizontal binary split (SPLIT_BT_HOR). The vertical binary split (SPLIT_BT_VER) means that the current CU is equally divided into two in the vertical direction. As Figure 4 shown, through the vertical binary split, two CUs with the same height as the current CU and a width that is half of the width of the current CU can be generated. The horizontal binary split (SPLIT_BT_HOR) means that the current CU is equally divided into two in the horizontal direction. As Figure 5 shown, through the horizontal binary split, two CUs with a height that is half of the height of the current CU and the same width as the current CU can be generated.
[0113] The two types of divisions according to the ternary tree structure may include a vertical ternary split (SPLIT_TT_VER) and a horizontal ternary split (SPLIT_TT_HOR). In the vertical ternary split (SPLIT_TT_VER), the current CU is divided in the vertical direction at a ratio of 1:2:1. As Figure 5 shown, through the vertical ternary split, two CUs with the same height as the current CU and a width that is 1 / 4 of the width of the current CU and one CU with the same height as the current CU and a width that is half of the width of the current CU can be generated. In the horizontal ternary split (SPLIT_TT_HOR), the current CU is divided in the horizontal direction at a ratio of 1:2:1. As Figure 5 shown, through the horizontal ternary split, two CUs with a height that is 1 / 4 of the height of the current CU and the same width as the current CU and one CU with a height that is half of the height of the current CU and the same width as the current CU can be generated.
[0114] Figure 6 It is a view showing a signaling mechanism of block partition information in a quadtree having a nested multi-type tree structure according to the present disclosure.
[0115] Here, the CTU is regarded as the root node of the quadtree and is first divided into a quadtree structure. Information indicating whether to perform quadtree partitioning on the current CU (the CTU of the quadtree or the node (QT_node) of the quadtree), such as qt_split_flag, can be signaled. For example, when the qt_split_flag has a first value (e.g., "1"), the current CU can be quadtree partitioned. Additionally, when the qt_split_flag has a second value (e.g., "0"), the current CU is not quadtree partitioned but becomes a leaf node (QT_leaf_node) of the quadtree. Then, each quadtree leaf node can be further divided into a multi-type tree structure. That is, the leaf node of the quadtree can become a node (MTT_node) of the multi-type tree. In the multi-type tree structure, a first flag (e.g., Mtt_split_cu_flag) can be signaled to indicate whether the current node is additionally partitioned. If the corresponding node is additionally partitioned (e.g., if the first flag is 1), a second flag (e.g., Mtt_split_cu_vertical_flag) can be signaled to indicate the partitioning direction. For example, the partitioning direction can be the vertical direction when the second flag is 1 and the horizontal direction when the second flag is 0. Then, a third flag (e.g., Mtt_split_cu_binary_flag) can be signaled to indicate whether the partitioning type is a binary partitioning type or a ternary partitioning type. For example, the partitioning type can be a binary partitioning type when the third flag is 1 and a ternary partitioning type when the third flag is 0. The nodes of the multi-type tree obtained by binary partitioning or ternary partitioning can be further divided into a multi-type tree structure. However, the nodes of the multi-type tree can not be divided into a quadtree structure. If the first flag is 0, the corresponding node of the multi-type tree is no longer partitioned but becomes a leaf node (MTT_leaf_node) of the multi-type tree. The CU corresponding to the leaf node of the multi-type tree can be used as the above-mentioned final coding unit.
[0116] Based on mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the multi-type tree partitioning mode (MttSplitMode) of the CU can be derived as shown in Table 1 below. In the following description, the multi-type tree partitioning mode can be referred to as the multi-tree partitioning type or the partitioning type.
[0117] [Table 1]
[0118] MttSplitMode mtt_split_cu_vertical_flag mtt_split_cu_binary_flag SPLIT_TT_HOR 0 0 SPLIT_BT_HOR 0 1 SPLIT_TT_VER 1 0 SPLIT_BT_VER 1 1
[0119] Figure 7 is a view showing an example of dividing the CTU into multiple CUs by applying a multi-type tree after applying a quadtree. InFigure 7 Among them, the bold block edges 710 represent quadtree partitioning, while the remaining edges 720 represent multi-type tree partitioning. A CU may correspond to a coding block (CB). In an embodiment, a CU may include a coding block of luminance samples and two coding blocks of chrominance samples corresponding to the luminance samples. The chrominance component (sample) CB or TB size may be derived based on the component ratio according to the color format (chrominance format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.) of the picture / image based on the luminance component (sample) CB or TB size. In the case of the 4:4:4 color format, the chrominance component CB / TB size may be set equal to the luminance component CB / TB size. In the case of the 4:2:2 color format, the width of the chrominance component CB / TB may be set to half of the width of the luminance component CB / TB and the height of the chrominance component CB / TB may be set to the height of the luminance component CB / TB. In the case of the 4:2:0 color format, the width of the chrominance component CB / TB may be set to half of the width of the luminance component CB / TB and the height of the chrominance component CB / TB may be set to half of the height of the luminance component CB / TB.
[0120] In an embodiment, when the size of the CTU is 128 based on luminance sample units, the size of the CU may have sizes ranging from 128×128 to 4×4, which is the same size as the CTU. In one embodiment, in the case of the 4:2:0 color format (or chrominance format), the chrominance CB size may have sizes ranging from 64×64 to 2×2.
[0121] In addition, in an embodiment, the CU size and the TU size may be the same. Alternatively, there may be multiple TUs in the CU region. The TU size generally may represent the luminance component (sample) transform block (TB) size.
[0122] The TU size may be derived based on the maximum allowable TB size maxTbSize which is a predetermined value. For example, when the CU size is greater than maxTbSize, multiple TUs (TBs) with maxTbSize may be derived from the CU, and the transform / inverse transform may be performed in units of TUs (TBs). For example, the maximum allowable luminance TB size may be 64×64 and the maximum allowable chrominance TB size may be 32×32. If the width or height of the CB divided according to the tree structure is greater than the maximum transform width or height, the CB may be automatically (or implicitly) divided until the TB size limits in the horizontal and vertical directions are met.
[0123] In addition, for example, when intra prediction is applied, the intra prediction mode / type may be derived on a CU (or CB) basis, and the neighboring reference sample derivation and prediction sample generation processes may be performed on a TU (or TB) basis. In this case, there may be one or more TUs (or TBs) in one CU (or CB) region, and in this case, multiple TUs or (TBs) may share the same intra prediction mode / type.
[0124] Furthermore, for a quadtree coding tree scheme with a nested multi-type tree, the following parameters may be signaled from an encoding device to a decoding device as SPS syntax elements. For example, the CTU size as a parameter representing the root node size of the quadtree, the MinQTSize as a parameter representing the minimum allowable quadtree leaf node size, the MaxBtSize as a parameter representing the maximum allowable binary tree root node size, the MaxTtSize as a parameter representing the maximum allowable ternary tree root node size, the MaxMttDepth as a parameter representing the maximum allowable hierarchical depth of multi-type tree partitioning starting from a quadtree leaf node, the MinBtSize as a parameter representing the minimum allowable binary tree leaf node size, or the MinTtSize as a parameter representing the minimum allowable ternary tree leaf node size, or at least one of them.
[0125] As an implementation using the 4:2:0 chroma format, the CTU size can be set to 128×128 luma blocks and two 64×64 chroma blocks corresponding to these luma blocks. In this case, MinOTSize can be set to 16×16, MaxBtSize can be set to 128×128, MaxTtSzie can be set to 64×64, MinBtSize and MinTtSize can be set to 4×4, and MaxMttDepth can be set to 4. Quadtree partitioning can be applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes can be referred to as leaf QT nodes. The size of the quadtree leaf nodes can range from a 16×16 size (e.g., MinOTSize) to a 128×128 size (e.g., CTU size). If the leaf QT node is 128×128, it may not be additionally partitioned into a binary tree / trinary tree. This is because, in this case, even if partitioned, it exceeds MaxBtsize and MaxTtszie (e.g., 64×64). In other cases, the leaf QT node can be further partitioned into a multi-type tree. Thus, the leaf QT node is the root node of the multi-type tree, and the leaf QT node can have a multi-type tree depth (mttDepth) of 0 value. If the multi-type tree depth reaches MaxMttdepth (e.g., 4), further partitioning may not be further considered. If the width of the multi-type tree node is equal to MinBtSize and less than or equal to 2×MinTtSize, further horizontal partitioning may not be considered. If the height of the multi-type tree node is equal to MinBtSize and less than or equal to 2×MinTtSize, further vertical partitioning may not be considered. When partitioning is not considered, the encoding device can skip signaling of the partitioning information. In this case, the decoding device can derive the partitioning information with a predetermined value.
[0126] In addition, a CTU may include an encoded block of luma samples (hereinafter referred to as a "luma block") and two encoded blocks of chroma samples corresponding thereto (hereinafter referred to as "chroma blocks"). The above-described coding tree scheme may be equally or separately applied to the luma block and the chroma block of the current CU. Specifically, the luma block and the chroma block in a CTU may be divided into the same block tree structure, and in this case, the tree structure may be represented as SINGLE_TREE. Alternatively, the luma block and the chroma block in a CTU may be divided into separate block tree structures, and in this case, the tree structure may be represented as DUAL_TREE. That is, when a CTU is divided into a dual tree, the block tree structure for the luma block and the block tree structure for the chroma block may exist separately. In this case, the block tree structure for the luma block may be referred to as DUAL_TREE_LUMA, and the block tree structure for the chroma component may be referred to as DUAL_TREE_CHROMA. For P and B slices / tile groups, the luma block and the chroma block in a CTU may be restricted to have the same coding tree structure. However, for I slices / tile groups, the luma block and the chroma block may have separate block tree structures from each other. If a separate block tree structure is applied, the luma CTB may be divided into CUs based on a specific coding tree structure, and the chroma CTB may be divided into chroma CUs based on another coding tree structure. That is, this means that a CU in an I slice / tile group to which a separate block tree structure is applied may include an encoded block of a luma component or two encoded blocks of chroma components, and a CU of a P or B slice / tile group may include blocks of three color components (one luma component and two chroma components).
[0127] Although a quadtree coding tree structure having a nested multi-type tree has been described, the structure for dividing a CU is not limited thereto. For example, the BT structure and the TT structure may be interpreted as concepts included in a multi-partition tree (MPT) structure, and a CU may be interpreted as being divided by a QT structure and an MPT structure. In an example of dividing a CU by a QT structure and an MPT structure, a syntax element (e.g., MPT_split_type) including information on how many blocks a leaf node of the QT structure is divided into and a syntax element (e.g., MPT_split_mode) including information on which of the vertical direction and the horizontal direction a leaf node of the QT structure is divided into may be signaled to determine the division structure.
[0128] In another example, the CU may be split in a manner different from the QT structure, the BT structure, or the TT structure. That is, different from splitting a lower-depth CU into 1 / 4 of a higher-depth CU according to the QT structure, splitting a lower-depth CU into 1 / 2 of a higher-depth CU according to the BT structure, or splitting a lower-depth CU into 1 / 4 or 1 / 2 of a higher-depth CU according to the TT structure, in some cases, a lower-depth CU may be split into 1 / 5, 1 / 3, 3 / 8, 3 / 5, 2 / 3, or 5 / 8 of a higher-depth CU, and the method of splitting the CU is not limited thereto.
[0129] The quadtree coding block structure with multi-type trees can provide a very flexible block splitting structure. Due to the splitting types supported in the multi-type trees, in some cases, different splitting patterns may potentially result in the same coding block structure. In the encoding device and the decoding device, by restricting the occurrence of such redundant splitting patterns, the amount of data of the splitting information can be reduced.
[0130] In addition, in the video / image encoding and decoding according to the present disclosure, the image processing unit may have a hierarchical structure. A picture may be divided into one or more tiles, patches, slices, and / or tile groups. A slice may include one or more patches. A patch may include one or more CTU rows in a tile. A slice may include an integer number of patches in a picture. A tile group may include one or more tiles. A tile may include one or more CTUs. A CTU may be divided into one or more CUs. A tile may be a rectangular area including a specific tile row and a specific tile column composed of multiple CTUs in a picture. According to the raster scan of the tiles in the picture, a tile group may include an integer number of tiles. The slice header may carry information / parameters applicable to the slice (the patches in the slice). When the encoding device or the decoding device has a multi-core processor, the encoding / decoding processes of the tiles, slices, patches, and / or tile groups may be executed in parallel.
[0131] In the present disclosure, the names or concepts of slices or tile groups may be mixed. That is, the tile group header may be referred to as the slice header. Here, a slice may have one of the slice types including an intra (I) slice, a predictive (P) slice, and a bi-predictive (B) slice. For the blocks in the I slice, inter prediction is not used for prediction, and only intra prediction may be used. Of course, even in this case, the original sample values may be encoded and signaled without prediction. For the blocks in the P slice, intra prediction or inter prediction may be used, and when inter prediction is used, only single prediction may be used. In addition, for the blocks in the B slice, intra prediction or inter prediction may be used, and when inter prediction is used, up to maximum bi-prediction may be used.
[0132] Based on the characteristics of the video image (e.g., resolution), or considering encoding efficiency and parallel processing, the encoding device can determine the tile / tile group, tile, slice, and maximum and minimum coding unit sizes. Additionally, information about this or information used to derive it can be included in the bitstream.
[0133] The decoding device can obtain information indicating the tile / tile group, tile, and slice of the current picture and whether the CTU in the tile is divided into multiple coding units. The encoding device and the decoding device can improve encoding efficiency by signaling such information only under specific conditions.
[0134] The slice header (slice header syntax) can include information / parameters commonly applied to the slice. The APS (APS syntax) or PPS (PPS syntax) can include information / parameters commonly applied to one or more pictures. The SPS (SPS syntax) can include information / parameters commonly applied to one or more sequences. The VPS (VPS syntax) can include information / parameters that can be commonly applied to multiple layers. The DPS (DPS syntax) can include information / parameters commonly applied to the entire video. The DPS can include information / parameters related to the combination of the coded video sequence (CVS).
[0135] Additionally, for example, information about the division and configuration of the tile / tile group / tile / slice can be configured at the encoding end through a higher-level syntax and sent to the decoding device in the form of a bitstream.
[0136] Overview of intra prediction
[0137] Hereinafter, the intra prediction performed by the above encoding device and decoding device will be described in more detail. Intra prediction can represent a prediction of generating a prediction sample of the current block based on reference samples in the picture (hereinafter referred to as the current picture) to which the current block belongs.
[0138] Reference will be Figure 8 made to give a description. When intra prediction is applied to the current block 801, the neighboring reference samples to be used for the intra prediction of the current block 801 can be derived. The neighboring reference samples of the current block can include: a total of 2×nh samples including the sample 811 with a size of nW×nH adjacent to the left boundary of the current block and the sample 812 adjacent to the lower left, a total of 2×nW samples including the sample 821 adjacent to the upper boundary of the current block and the sample 822 adjacent to the upper right, and one sample 831 adjacent to the upper left of the current block. Alternatively, the neighboring reference samples of the current block can include multiple columns of upper neighboring samples and multiple rows of left neighboring samples.
[0139] In addition, the neighboring reference samples of the current block may include: a total of nH samples 841 of size nW×nH adjacent to the right boundary of the current block, a total of nW samples 851 adjacent to the lower boundary of the current block, and one sample 842 adjacent to the lower right of the current block.
[0140] However, some of the neighboring reference samples of the current block have not been decoded or may be unavailable. In this case, the decoding device may construct the neighboring reference samples to be used for prediction by replacing the unavailable samples with the available samples. Alternatively, the neighboring reference samples to be used for prediction may be constructed by interpolation of the available samples.
[0141] When deriving neighboring reference samples, (i) a predicted sample may be derived based on an average or interpolation of neighboring reference samples of a current block, and (ii) a predicted sample may be derived based on reference samples among the neighboring reference samples of the current block that exist in a specific (prediction) direction for the predicted sample. The case of (i) may be referred to as a non-directional mode or non-corner mode, and the case of (ii) may be referred to as a directional mode or corner mode. Additionally, a predicted sample may be generated by interpolation of a first neighboring sample and a second neighboring sample among the neighboring reference samples based on which the predicted sample of the current block is located in a direction opposite to the prediction direction of the intra prediction mode of the current block. The above case may be referred to as linear interpolation intra prediction (LIP). Additionally, a chrominance prediction sample may be generated based on luminance samples using a linear model. This case may be referred to as the LM mode. Additionally, a temporary predicted sample of a current block may be derived based on filtered neighboring reference samples, and a predicted sample of the current block may be derived by weighted summation of at least one reference sample derived according to the intra prediction mode among existing neighboring reference samples (i.e., unfiltered neighboring reference samples) and the temporary predicted sample. The above case may be referred to as position-dependent intra prediction (PDPC). Additionally, a reference sample row with the highest prediction accuracy is selected from among multiple neighboring reference sample rows of the current block to derive a predicted sample using the reference samples located in the prediction direction on the corresponding row, and at this time, intra prediction coding may be performed in such a way that the used reference sample row is indicated (signaled) to a decoding device. The above case may be referred to as multi-reference row (MRL) intra prediction or MRL-based intra prediction. Additionally, a current block is divided into vertical or horizontal sub-partitions to perform intra prediction based on the same intra prediction mode, but neighboring reference samples may be derived and used on a sub-partition basis. That is, in this case, the intra prediction mode of the current block is also applied to the sub-partitions, but in some cases, the intra prediction performance may be improved by deriving and using neighboring reference samples on a sub-partition basis. This prediction method may be referred to as intra sub-partition (ISP) or ISP-based intra prediction. These intra prediction methods may be referred to as intra prediction types in order to be distinguished from intra prediction modes (e.g., DC mode, planar mode, or directional mode). Intra prediction types may be referred to by various terms such as intra prediction techniques or additional intra prediction modes. For example, an intra prediction type (or additional intra prediction mode, etc.) may include at least one of the above LIP, PDPC, MRL, and ISP. A general intra prediction method other than specific intra prediction types such as LIP, PDPC, MRL, and ISP may be referred to as a normal intra prediction type. The normal intra prediction type may refer to a case where the above specific intra prediction types are not applied, and prediction may be performed based on the above intra prediction modes. Additionally, when necessary, post-filtering may be performed on the derived predicted samples.
[0142] Specifically, the intra prediction process may include an intra prediction mode / type determination step, a neighboring reference sample derivation step, and a predicted sample derivation step based on the intra prediction mode / type. Additionally, when necessary, a post-filtering step may be performed on the derived predicted samples.
[0143] In addition to the above-mentioned intra prediction types, affine linear weighted intra prediction (ALWIP) may also be used. ALWIP may be referred to as linear weighted intra prediction (LWIP) or matrix weighted intra prediction or matrix-based intra prediction (MIP). When applying MIP to a current block, i) neighboring reference samples that have undergone an averaging process may be used, ii) a matrix-vector multiplication process may be performed, and iii) when necessary, a horizontal / vertical interpolation process may be further performed to derive the predicted samples of the current block. The intra prediction mode for MIP may be configured differently from the intra prediction modes used in the above-mentioned LIP, PDPC, MRL, and ISP intra predictions or normal intra prediction. The intra prediction mode for MIP may be referred to as the MIP intra prediction mode, the MIP prediction mode, or the MIP mode. For example, the matrix and offset used in the matrix-vector multiplication may be set differently according to the intra prediction mode for MIP. Here, the matrix may be referred to as the (MIP) weighted matrix, and the offset may be referred to as the (MIP) offset vector or the (MIP) bias vector. A specific MIP method will be described later.
[0144] The block reconstruction process based on intra prediction and intra prediction units in an encoding device may schematically include, for example, the following. S910 may be executed by the intra prediction unit 185 of the encoding device, and S920 may be executed by a residual processor including at least one of the subtractor 115, the transformer 120, the quantizer 130, the dequantizer 140, and the inverse transformer 150 of the encoding device. Specifically, S920 may be executed by the subtractor 115 of the encoding device. In S930, prediction information may be derived by the intra prediction unit 185 and encoded by the entropy encoder 190. In S930, residual information may be derived by the residual processor and encoded by the entropy encoder 190. Residual information is information about residual samples. Residual information may include information about the quantized transform coefficients of the residual samples. As described above, the residual samples may be derived as transform coefficients by the transformer 120 of the encoding device, and the transform coefficients may be derived as quantized transform coefficients by the quantizer 130. Information about the quantized transform coefficients may be encoded by the entropy encoder 190 through a residual encoding process.
[0145] The encoding device may perform intra prediction on the current block (S910). The encoding device may derive the intra prediction mode / type of the current block, derive the neighboring reference samples of the current block, and generate prediction samples in the current block based on the intra prediction mode / type and the neighboring reference samples. Here, the process for determining the intra prediction mode / type, the process for deriving the neighboring reference samples, and the process for generating the prediction samples may be executed simultaneously, or any one of the processes may be executed before another process. For example, although not shown, the intra prediction unit 185 of the encoding device may include an intra prediction mode / type determination unit, a reference sample derivation unit, and a prediction sample derivation unit. The intra prediction mode / type determination unit may determine the intra prediction mode / type of the current block, the reference sample derivation unit may derive the neighboring reference samples of the current block, and the prediction sample derivation unit may derive the prediction samples of the current block. In addition, when performing the following prediction sample filtering process, the intra prediction unit 185 may further include a prediction sample filter. The encoding device may determine the mode / type to be applied to the current block from among multiple intra prediction modes / types. The encoding device may compare the RD costs of the intra prediction modes / types and determine the best intra prediction mode / type for the current block.
[0146] In addition, the encoding device may perform a prediction sample filtering process. Prediction sample filtering may be referred to as post-filtering. Some or all of the prediction samples may be filtered through the prediction sample filtering process. In some cases, the prediction sample filtering process may be omitted.
[0147] The encoding device may generate residual samples for the current block based on the (filtered) prediction samples (S920). The encoding device may compare the prediction samples with the original samples of the current block based on the phase and derive the residual samples.
[0148] The encoding device may encode the image information including information about intra prediction (prediction information) and residual information of the residual samples (S930). The prediction information may include intra prediction mode information and intra prediction type information. The encoding device may output the encoded image information in the form of a bitstream. The output bitstream may be sent to the decoding device via a storage medium or a network.
[0149] The residual information may include the following residual coding syntax. The encoding device may perform transformation / quantization on the residual samples to derive quantized transform coefficients. The residual information may include information about the quantized transform coefficients.
[0150] In addition, as described above, the encoding device may generate a reconstructed picture (including reconstructed samples and reconstructed blocks). To this end, the encoding device may perform dequantization / inverse transformation again on the quantized transform coefficients to derive (modified) residual samples. The residual samples are transformed / quantized and then dequantized / inverse transformed so as to derive the same residual samples as the residual samples derived in the decoding device as described above. The encoding device may generate a reconstructed block including the reconstructed samples of the current block based on the prediction samples and the (modified) residual samples. The reconstructed picture of the current picture may be generated based on the reconstructed block. As described above, the in-loop filtering process is further applicable to the reconstructed picture.
[0151] The video / image decoding process based on intra prediction and the intra prediction unit in the decoding device may schematically include, for example, the following. The decoding device may perform operations corresponding to the operations performed by the encoding device.
[0152] S1010 to S1030 may be performed by the intra prediction unit 265 of the decoding device, and the prediction information of S1010 and the residual information of S1040 may be obtained by the entropy decoder 210 of the decoding device from the bitstream. The residual processor including the dequantizer 220 or the inverse transformer 230 of the decoding device may derive the residual samples of the current block based on the residual information. Specifically, the dequantizer 220 of the residual processor may perform dequantization based on the quantized transform coefficients derived according to the residual information to derive the transform coefficients, and the dequantizer 220 of the residual processor may perform inverse transformation on the transform coefficients to derive the residual samples of the current block. S1050 may be performed by the adder 235 or the reconstructor of the decoding device.
[0153] Specifically, the decoding device may derive the intra prediction mode / type of the current block based on the received prediction information (intra prediction mode / type information) (S1010). The decoding device may derive the neighboring reference samples of the current block (S1020). The decoding device may generate prediction samples in the current block based on the intra prediction mode / type and the neighboring reference samples (S1030). In this case, the decoding device may perform a prediction sample filtering process. The prediction sample filtering may be referred to as post-filtering. Some or all of the prediction samples in the prediction samples may be filtered through the prediction sample filtering process. In some cases, the prediction sample filtering process may be omitted.
[0154] The decoding device generates the residual samples of the current block based on the received residual information. The decoding device may generate the reconstructed samples of the current block based on the prediction samples and the residual samples, and derive a reconstructed block including the reconstructed samples (S1040). The reconstructed picture of the current picture may be generated based on the reconstructed block. As described above, the in-loop filtering process is further applicable to the reconstructed picture.
[0155] Here, the intra prediction unit 265 of the decoding device may include an intra prediction mode / type determination unit, a reference sample derivation unit, and a predicted sample derivation unit. The intra prediction mode / type determination unit may determine the intra prediction mode / type of the current block based on the intra prediction mode / type information obtained by the entropy decoder 210. The reference sample derivation unit may derive the neighboring reference samples of the current block. The predicted sample derivation unit may derive the predicted samples of the current block. In addition, although not shown, when the above-described predicted sample filtering process is performed, the intra prediction unit 265 may further include a predicted sample filter.
[0156] The intra prediction mode information may include flag information (e.g., intra_luma_mpm_flag) specifying whether the most probable mode (MPM) or the residual mode is applied to the current block. When the MPM is applied to the current block, the prediction mode information may further include index information (e.g., intra_luma_mpm_idx) specifying one of the intra prediction mode candidates (MPM candidates). The intra prediction mode candidates (MPM candidates) may be configured as an MPM candidate list or an MPM list. Additionally, when the MPM is not applied to the current block, the intra prediction mode information may further include residual mode information (e.g., intra_luma_mpm_remainder) specifying one of the remaining intra prediction modes other than the intra prediction mode candidates (MPM candidates). The decoding device may determine the intra prediction mode of the current block based on the intra prediction mode information. For the above MIP, a separate MPL list may be constructed.
[0157] In addition, the intra prediction type information may be implemented in various forms. For example, the intra prediction type information may include intra prediction type index information specifying one of the intra prediction types. As another example, the intra prediction type information may include reference sample row information (e.g., intra_luma_ref_idx) specifying whether the MRL is applied to the current block and which reference sample row to use if applied, ISP flag information (e.g., intra_subpartitions_mode_flag) specifying whether the ISP is applied to the current block, ISP type information (e.g., intra_subpartitions_split_flag) specifying the partitioning type of the subpartitions when the ISP is applied, flag information specifying whether the PDCP is applied, or flag information specifying whether the LIP is applied, or at least one of them. Additionally, the intra prediction type information may include an MIP flag specifying whether the MIP is applied to the current block.
[0158] Intra prediction mode information and / or intra prediction type information can be encoded / decoded by the encoding method described in the present disclosure. For example, the intra prediction mode information and / or intra prediction type information can be encoded / decoded by entropy coding (e.g., CABAC or CAVLC) based on truncated (Rice) binary code.
[0159] Overview of inter prediction
[0160] In the following, the detailed techniques of the inter prediction method in the description of encoding and decoding with reference to Figure 2 and Figure 3 will be described. In the case of a decoding device, the inter prediction-based video / image decoding method and the inter prediction unit in the decoding device can operate according to the following description. In the case of an encoding device, the inter prediction-based video / image encoding method and the inter prediction unit in the encoding device can operate according to the following description. Additionally, the data encoded according to the following description can be stored in the form of a bitstream.
[0161] The prediction units of the encoding device and the decoding device may perform inter-frame prediction in units of blocks to derive prediction samples. Inter-frame prediction may mean prediction derived in a manner that depends on data elements (e.g., sample values, motion information, etc.) of a picture other than the current picture. When applying inter-frame prediction to a current block, a prediction block (prediction sample array) of the current block may be derived based on a reference block (reference sample array) specified by a motion vector on a reference picture indicated by a reference picture index. In this case, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information of the current block may be predicted in units of blocks, sub-blocks, or samples based on the motion information correlation between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. When applying inter-frame prediction, neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to as a collocated reference block, a collocated CU (ColCU), and the reference picture including the temporal neighboring block may be referred to as a collocated picture (colPic). For example, a motion information candidate list may be constructed based on neighboring blocks of the current block, and in order to derive the motion vector and / or the reference picture index of the current block, a flag or index information indicating which candidate is selected (used) may be signaled. Inter-frame prediction may be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the motion information of the current block may be equal to the motion information of the selected neighboring block. In the case of the skip mode, a residual signal may not be transmitted as in the merge mode. In the case of the motion vector prediction (MVP) mode, the motion vector of the selected neighboring block may be used as a motion vector predictor, and a motion vector difference may be signaled. In this case, the sum of the motion vector predictor and the motion vector difference may be used to derive the motion vector of the current block.
[0162] Depending on the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.), the motion information may include L0 motion information and / or L1 motion information. The motion vector in the L0 direction may be referred to as the L0 motion vector or MVL0, and the motion vector in the L1 direction may be referred to as the L1 motion vector or MVL1. The prediction based on the L0 motion vector may be referred to as L0 prediction, the prediction based on the L1 motion vector may be referred to as L1 prediction, and the prediction based on both the L0 motion vector and the L1 motion vector may be referred to as bi-prediction. Here, the L0 motion vector may indicate the motion vector associated with the reference picture list L0 (L0), and the L1 motion vector may indicate the motion vector associated with the reference picture list L1 (L1). The reference picture list L0 may include pictures before the current picture in the output order as reference pictures, and the reference picture list L1 may include pictures after the current picture in the output order. The previous picture may be referred to as the forward (reference) picture, and the subsequent picture may be referred to as the backward (reference) picture. The reference picture list L0 may also include pictures after the current picture in the output order as reference pictures. In this case, within the reference picture list L0, the previous picture may be indexed first, and then the subsequent picture may be indexed. The reference picture list L1 may also include pictures before the current picture in the output order as reference pictures. In this case, within the reference picture list L1, the subsequent picture may be indexed first, and then the previous picture may be indexed. Here, the output order may correspond to the picture order count (POC) order.
[0163] The video / image encoding process based on inter-frame prediction and the inter-frame prediction unit in the encoding device may schematically include, for example, the following. The reference Figure 11A description thereof will be given. The encoding device performs inter prediction on the current block (S1100). The image encoding device may derive an inter prediction mode and motion information of the current block and generate a prediction sample of the current block. Here, the inter prediction mode determination, motion information derivation, and prediction sample generation processes may be performed simultaneously, or any one of them may be performed before the other processes. For example, the inter prediction unit of the encoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit, and the prediction mode determination unit may determine the prediction mode of the current block, the motion information derivation unit may derive the motion information of the current block, and the prediction sample derivation unit may derive the prediction sample of the current block. For example, the inter predictor of the encoding device may search for a block similar to the current block within a predetermined region (search region) of the reference picture by motion estimation and derive a reference block whose difference from the current block is equal to or less than a predetermined criterion or minimum value. Based on this, a reference picture index indicating the reference picture in which the reference block is located may be derived, and a motion vector may be derived based on the positional difference between the reference block and the current block. The encoding device may determine the mode applied to the current block among various prediction modes. The encoding device may compare the RD costs of various prediction modes and determine the best prediction mode of the current block.
[0164] For example, when the skip mode or the merge mode is applied to the current block, the encoding device may construct a merge candidate list and derive a reference block whose difference from the current block is equal to or less than a predetermined criterion or minimum value among the reference blocks indicated by the merge candidates included in the merge candidate list. In this case, a merge candidate associated with the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and signaled to the decoding device. The motion information of the current block may be derived using the motion information of the selected merge candidate.
[0165] As another example, when the (A)MVP mode is applied to the current block, the encoding device may construct an (A)mvp candidate list and derive a motion vector of an mvp candidate selected from among the mvp candidates included in the (A)MVP candidate list. In this case, for example, the motion vector indicating the reference block derived by the above motion estimation may be used as the motion vector of the current block, and the mvp candidate having the motion vector with the smallest difference from the motion vector of the current block among the mvp candidates may be the selected mvp candidate. A motion vector difference (MVD) that is the difference obtained by subtracting the mvp from the motion vector of the current block may be derived. In this case, information about the MVD may be signaled to the decoding device. Additionally, when the (A)MVP mode is applied, the value of the reference picture index may be constructed as reference picture index information and signaled separately to the decoding device.
[0166] The encoding device may derive a residual sample based on a prediction sample (S1120). The encoding device may derive the residual sample by comparing the original sample of the current block with the prediction sample.
[0167] The encoding device may encode the image information including the prediction information and the residual information (S1130). The encoding device may output the encoded image information in the form of a bitstream. The prediction information may include prediction mode information (e.g., skip flag, merge flag, or mode index, etc.) and information about motion information as information related to the prediction process. The information about motion information may include candidate selection information (e.g., merge index, mvp flag, or mvp index) as information for deriving a motion vector. Additionally, the information about motion information may include information about the above MVD and / or reference picture index information. Additionally, the information about motion information may include information indicating whether to apply L0 prediction, L1 prediction, or bi-prediction. The residual information is information about the residual sample. The residual information may include information about the quantization transform coefficients for the residual sample.
[0168] The output bitstream may be stored in a (digital) storage medium and sent to the decoding device, or may be sent to the decoding device via a network.
[0169] As described above, the encoding device may generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference sample and the residual sample. This enables the encoding device to derive the same prediction result as the prediction result executed by the decoding device, thereby improving the encoding efficiency. Therefore, the encoding device may store the reconstructed picture (or reconstructed samples and reconstructed blocks) in the memory and use it as a reference picture for inter-frame prediction. As described above, the in-loop filtering process is further applicable to the reconstructed picture.
[0170] The video / image decoding process in the decoding device based on inter-frame prediction and the inter-frame prediction unit may schematically include, for example, the following.
[0171] The decoding device may perform operations corresponding to those performed by the encoding device. The decoding device may perform prediction for the current block based on the received prediction information and derive a prediction sample.
[0172] Specifically, the decoding device may determine the prediction mode of the current block based on the received prediction information (S1210). The image decoding device may determine which inter-frame prediction mode to apply to the current block based on the prediction mode information in the prediction information.
[0173] For example, it is possible to determine whether to apply a merge mode or an (A)MVP mode to the current block based on a merge flag. Alternatively, one of various inter prediction mode candidates can be selected based on a mode index. The inter prediction mode candidates can include a skip mode, a merge mode, and / or an (A)MVP mode, or can include various inter prediction modes to be described below.
[0174] The decoding device can derive motion information of the current block based on the determined inter prediction mode (S1220). For example, when the skip mode or the merge mode is applied to the current block, the decoding device can construct a merge candidate list to be described below and select one of the merge candidates included in the merge candidate list. The selection can be performed based on the above candidate selection information (merge index). The motion information of the selected merge candidate can be used to derive the motion information of the current block. The motion information of the selected merge candidate can be used as the motion information of the current block.
[0175] As another example, when the (A)MVP mode is applied to the current block, the decoding device can construct an (A)MVP candidate list and use the motion vector of the mvp candidate selected from among the mvp candidates included in the (A)MVP candidate list as the mvp of the current block. The selection can be performed based on the above candidate selection information (mvp flag or mvp index). In this case, the MVD of the current block can be derived based on the information about the MVD, and the motion vector of the current block can be derived based on the mvp and the MVD of the current block. In addition, the reference picture index of the current block can be derived based on the reference picture index information. The picture indicated by the reference picture index in the reference picture list of the current block can be derived as the reference picture referred to for the inter prediction of the current block.
[0176] In addition, as described below, the motion information of the current block can be derived without constructing a candidate list, and in this case, the motion information of the current block can be derived according to the process disclosed under the following prediction mode. In this case, the above candidate list construction can be omitted.
[0177] The image decoding device can generate a prediction sample of the current block based on the motion information of the current block (S1230). In this case, the reference picture can be derived based on the reference picture index of the current block, and the prediction sample of the current block can be derived using the samples of the reference block indicated by the motion vector of the current block on the reference picture. In this case, as described below, in some cases, a prediction sample filtering process can be further performed for all or some of the prediction samples of the current block.
[0178] For example, the inter-frame prediction unit of the decoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit may determine the prediction mode of the current block based on the received prediction mode information. The motion information derivation unit may derive the motion information (such as motion vector and / or reference picture index, etc.) of the current block based on the received motion information. And the prediction sample derivation unit may derive the prediction samples of the current block.
[0179] The decoding device may generate residual samples of the current block based on the received residual information (S1240). The decoding device may generate reconstructed samples of the current block based on the prediction samples and the residual samples, and generate a reconstructed picture based on this (S1250). Thereafter, the in-loop filtering process is applied to the reconstructed picture as described above.
[0180] As described above, the inter-frame prediction process may include steps of determining an inter-frame prediction mode, deriving motion information according to the determined prediction mode, and performing prediction (generating prediction samples) based on the derived motion information. The inter-frame prediction process may be performed by the encoding device and the decoding device as described above.
[0181] Quantization / dequantization
[0182] As described above, the quantizer of the encoding device may derive quantized transform coefficients by applying quantization to the transform coefficients, and the dequantizer of the encoding device or the dequantizer of the decoding device may derive the transform coefficients by applying dequantization to the quantized transform coefficients.
[0183] In the encoding and decoding of moving images / still images, the quantization rate may be changed and the compression rate may be adjusted using the changed quantization rate. From an implementation perspective, considering complexity, instead of directly using the quantization rate, a quantization parameter (QP) may be used. For example, a quantization parameter with an integer value from 0 to 63 may be used, and each quantization parameter value may correspond to an actual quantization rate. Additionally, the quantization parameter QP of the luminance component (luminance samples) and the quantization parameter QP of the chrominance component (chrominance samples) may be set differently. Y and the quantization parameter QP of the chrominance component (chrominance samples) C .
[0184] In the quantization process, the transform coefficient C may be received and divided by the quantization rate Qstep to obtain a quantized transform. In this case, considering the computational complexity, the quantization rate may be multiplied by a ratio to form an integer, and a shift operation may be performed according to a value corresponding to the ratio value. The quantization ratio may be derived based on the product of the quantization rate and the ratio value. That is, the quantization ratio may be derived according to QP. By applying the quantization ratio to the transform coefficient C, the quantized transform coefficient C' may be derived.
[0185] Inverse quantization is the inverse process of quantization. By multiplying the quantized transform coefficient C’ by the quantization rate Qstep, the reconstructed transform coefficient C” is obtained. Additionally, a level ratio can be derived based on quantization parameters, and the level ratio can be applied to the quantized transform coefficient C’ to derive the reconstructed transform coefficient C”. Due to losses in the transform and / or quantization process, the reconstructed transform coefficient C” may slightly differ from the initial transform coefficient C. Therefore, inverse quantization can be performed in the encoding device in the same manner as in the decoding device.
[0186] In addition, an adaptive frequency-weighted quantization technique that adjusts the quantization strength according to frequency can be applied. The adaptive frequency-weighted quantization technique refers to a method of applying the quantization strength differently according to frequency. In adaptive frequency-weighted quantization, a predefined quantization scaling matrix can be used to apply the quantization strength differently according to frequency. That is, the above quantization / inverse quantization process can be performed based on the quantization scaling matrix. For example, different quantization scaling matrices can be used according to the size of the current block and / or whether the prediction mode applied to the current block to generate the residual signal of the current block is inter prediction or intra prediction. The quantization scaling matrix can be referred to as a quantization matrix or a scaling matrix. The quantization scaling matrix can be predefined. Additionally, for frequency-adaptive scaling, the frequency quantization ratio information of the quantization scaling matrix can be constructed / encoded in the encoding device and signaled to the decoding device. The frequency quantization ratio information can be referred to as quantization scaling information. The frequency quantization ratio information can include scaling list data scaling_list_data. The (modified) quantization scaling matrix can be derived based on the scaling list data. Additionally, the frequency quantization ratio information can include a presence flag indicating whether there is scaling list data. Alternatively, when signaling the scaling list data at a higher level (e.g., SPS), information indicating whether the scaling list data is modified at a lower level (e.g., PPS or tile group header, etc.) can be further included.
[0187] Transformation / inverse transformation
[0188] As described above, the encoding device can derive a residual block (residual samples) based on a block (predicted block) predicted by intra / inter / IBC prediction, and derive quantized transform coefficients by applying transform and quantization to the derived residual samples. Information about the quantized transform coefficients (residual information) can be included and encoded in the residual encoding syntax and output in the form of a bitstream. The decoding device can obtain and decode the information about the quantized transform coefficients (residual information) from the bitstream to derive the quantized transform coefficients. The decoding device can derive the residual samples based on the quantized transform coefficients by dequantization / inverse transform. As described above, at least one of quantization / dequantization and / or transform / inverse transform can be skipped. When the transform / inverse transform is skipped, the transform coefficients can be referred to as coefficients or residual coefficients, or for the sake of unified expression, can still be referred to as transform coefficients. Whether the transform / inverse transform is skipped can be signaled based on a transform skip flag (e.g., transform_skip_flag).
[0189] The transform / inverse transform can be performed based on a transform kernel. For example, a multi-transform selection (MTS) scheme for performing the transform / inverse transform is applicable. In this case, some of a plurality of transform kernel sets can be selected and applied to the current block. The transform kernel can be referred to by various terms such as a transform matrix or a transform type. For example, a transform kernel set can indicate a combination of a vertical direction transform kernel (vertical transform kernel) and a horizontal direction transform kernel (horizontal transform kernel).
[0190] The transform / inverse transform can be performed in units of a CU or a TU. That is, the transform / inverse transform is applicable to the residual samples in the CU or the residual samples in the TU. The CU size can be equal to the TU size, or there can be a plurality of TUs in the CU region. In addition, the CU size can generally indicate the luminance component (samples) CB size. The TU size can generally indicate the luminance component (samples) TB size. The chrominance component (samples) CB or TB size can be derived based on the component ratio according to the color format (chrominance format) (e.g., 4:4:4, 4:2:2, 4:2:0, etc.) based on the luminance component (samples) CB or TB size. The TU size can be derived based on maxTbSize. For example, when the CU size is larger than maxTbSize, a plurality of TUs (TBs) of maxTbSize can be derived from the CU, and the transform / inverse transform can be performed in units of the TUs (TBs). Various intra prediction types such as ISP can be determined considering maxTbSize. Information about maxTbSize can be predetermined or can be generated and encoded in the encoding device and signaled to the encoding device.
[0191] Entropy coding
[0192] All or some of the video / image information can be as described above with reference toFigure 2 The described entropy encoder 190 performs entropy encoding, with reference to Figure 3 All or some of the described video / image information can be entropy decoded by the entropy decoder 310. In this case, the video / image information can be encoded / decoded in units of syntax elements. In the present disclosure, the encoded / decoded information can include encoding / decoding performed by the methods described in this paragraph.
[0193] Figure 13 is a block diagram of CABAC for encoding a syntax element. In the encoding process of CABAC, first, when the input signal is a syntax element other than a binary value, the input signal can be transformed into a binary value through binarization. When the input signal already has a binary value, binarization can be bypassed. Here, the binary number 0 or 1 that configures the binary value can be referred to as a bin. For example, when the binarized binary string (bin string) is 110, each of 1, 1, and 0 can be referred to as a bin. The bin of a syntax element can represent the value of the corresponding syntax element.
[0194] The binarized bin can be input to a normal encoding engine or a bypass encoding engine. The normal encoding engine can assign a context model reflecting a probability value to the corresponding bin and encode the corresponding bit based on the assigned context model. The normal encoding engine can encode each bin and then update the probability model of the corresponding bin. The bin encoded in this way can be referred to as a context-encoded bin. The bypass encoding engine can bypass the process of estimating the probability for the input bin and the process of updating the probability pattern applied to the corresponding bin after encoding. Instead of assigning a context, the bypass encoding engine can encode the input bin by applying a uniform probability distribution (e.g., 50:50), thereby improving the encoding speed. The bin encoded in this way can be referred to as a bypass bin. A context model can be assigned and updated for each context-encoded (normal-encoded) bin, and the context model can be indicated based on ctxidx or ctxInc. ctxidx can be derived based on ctxInc. Specifically, for example, the context index ctxidx indicating the context model for each normally encoded bin can be derived as the sum of the context index increment ctxInc and the context index offset ctxIdxOffset. Here, ctxInc can be derived differently for each bin. ctxIdxOffset can be represented by the lowest value of ctxIdx. The lowest value of ctxIdx can be referred to as the initial value initValue of ctxIdx. ctxIdxOffset is generally a value used to distinguish the context model from other syntax elements, and the context model of a syntax element can be distinguished / derived based on ctxinc.
[0195] During the entropy encoding process, it is possible to determine whether to perform encoding through a regular encoding engine or a bypass encoding engine, and it is possible to switch the encoding path. Entropy decoding can be performed in the reverse order of the same processing as entropy encoding.
[0196] For example, as Figure 14 and Figure 15 shown, the above-mentioned entropy encoding can be performed. Referring to Figure 14 and Figure 15 , an encoding device (entropy encoder) can perform an entropy encoding process on image / video information. The image / video information can include segmentation-related information, prediction-related information (e.g., inter-frame / intra-frame prediction discrimination information, intra-frame prediction mode information, inter-frame prediction mode information, etc.), residual information, in-loop filtering-related information, etc., or can include various syntax elements related thereto. Entropy encoding can be performed on a syntax element-by-syntax element basis. Figure 14 The steps S1410 to S1420 of Figure 2 can be performed by the entropy encoder 190 of the encoding device of
[0197] The encoding device can perform binarization (S1410) on a target syntax element. Here, the binarization can be based on various binarization methods such as truncated Rice binarization processing, fixed-length binarization processing, etc., and the binarization method for the target syntax element can be predefined. The binarization process can be performed by the binarization unit 191 in the entropy encoder 190.
[0198] The encoding device can perform entropy encoding (S1420) on a target syntax element. The encoding device can perform context-based (context-based) or bypass-encoding-based encoding on the bin string of the target syntax element based on entropy encoding techniques such as CABAC (Context Adaptive Binary Arithmetic Coding) or CAVLC (Context Adaptive Variable Length Coding), and its output can be included in the bitstream. The entropy encoding process can be performed by the entropy encoding processor 192 in the entropy encoder 190. As described above, the bitstream can be sent to the decoding device via a (digital) storage medium or a network.
[0199] Referring to Figure 16 and Figure 17 , a decoding device (entropy decoder) can decode the encoded image / video information. The image / video information can include segmentation-related information, prediction-related information (e.g., inter-frame / intra-frame prediction discrimination information, intra-frame prediction mode information, inter-frame prediction mode information, etc.), residual information, in-loop filtering-related information, etc., or can include various syntax elements related thereto. Entropy encoding can be performed on a syntax element-by-syntax element basis. Figure 16 The steps S1610 to S1620 of Figure 3 can be performed by the entropy decoder 210 of the decoding device of
[0200] The decoding device may perform binarization on a target syntax element (S1610). Here, the binarization may be based on various binarization methods such as truncated Rice binarization processing, fixed-length binarization processing, etc., and the binarization method for the target syntax element may be predefined. The decoding device may derive an available bin string (bin string candidate) of available values of the target syntax element through the binarization process. The binarization process may be performed by the binarization unit 211 in the entropy decoder 210.
[0201] The decoding device may perform entropy decoding on the target syntax element (S1620). The decoding device may compare the derived bin string with the available bin strings of the corresponding syntax element, and at the same time decode and parse the bin of the target syntax element sequentially from the input bits in the bitstream. If the derived bin string is equal to one of the available bin strings, the value corresponding to the corresponding bin string may be derived as the value of the corresponding syntax element. If not, the above process may be performed again after further parsing the next bit in the bitstream. Through this processing, variable-length bits may be used to signal the corresponding information without using the start bit or end bit of a specific piece of information (specific syntax element) in the bitstream. Thus, relatively fewer bits may be allocated to low values, and the overall coding efficiency may be improved.
[0202] The decoding device may perform context-based or bypass-based decoding on the bins in the bin string from the bitstream based on entropy coding techniques such as CABAC or CAVLC. The entropy decoding process may be performed by the entropy decoding processor 212 in the entropy decoder 210. The bitstream may include various information for image / video decoding as described above. As described above, the bitstream may be sent to the decoding device through a (digital) storage medium or a network.
[0203] In the present disclosure, a table (syntax table) including syntax elements may be used to indicate information signaling from the encoding device to the decoding device. The order of the syntax elements in the table including the syntax elements used in the present disclosure may indicate the parsing order of the syntax elements from the bitstream. The encoding device may construct and encode the syntax elements such that the decoding device parses the syntax elements in the parsing order, and the decoding device may parse and decode the syntax elements of the corresponding syntax table from the bitstream according to the parsing order and obtain the values of the syntax elements.
[0204] General image / video coding process
[0205] In image / video coding, pictures configuring an image / video may be encoded / decoded according to the decoding order. The picture order corresponding to the output order of the decoded pictures may be set differently from the decoding order, and based on this, not only forward prediction but also backward prediction may be performed during inter-frame prediction.
[0206] Figure 18 An example of a schematic picture decoding process to which embodiments of the present disclosure are applicable is shown. In Figure 18 , S1810 may be performed in the entropy decoder 210 of the decoding device, S1820 may be performed in a prediction unit including an intra prediction unit 265 and an inter prediction unit 260, S1830 may be performed in a residual processor including a dequantizer 220 and an inverse transformer 230, S1840 may be performed in an adder 235, and S1850 may be performed in a filter 240. S1810 may include the information decoding process described in the present disclosure, S1820 may include the inter / intra prediction process described in the present disclosure, S1830 may include the residual processing process described in the present disclosure, S1840 may include the block / picture reconstruction process described in the present disclosure, and S1850 may include the in-loop filtering process described in the present disclosure.
[0207] Referring to Figure 18 , the picture decoding process may schematically include a process (S1810) for obtaining image / video information (by decoding) from a bitstream, a picture reconstruction process (S1820 to S1840), and an in-loop filtering process (S1850) for the reconstructed picture. The picture reconstruction process may be performed based on prediction samples and residual samples obtained through inter / intra prediction (S1820) and residual processing (S1830) (dequantization and inverse transformation of quantization transform coefficients) described in the present disclosure. The modified reconstructed picture may be generated by an in-loop filtering process for the reconstructed picture generated by the picture reconstruction process, and the modified reconstructed picture may be output as a decoded picture, stored in the decoded picture buffer or memory 250 of the decoding device, and used as a reference picture in the inter prediction process when decoding a picture later. In some cases, the in-loop filtering process may be omitted. In this case, the reconstructed picture may be output as a decoded picture, stored in the decoded picture buffer or memory 250 of the decoding device, and used as a reference picture in the inter prediction process when decoding a picture later. The in-loop filtering process (S1850) may include a deblocking filtering process, a sample adaptive offset (SAO) process, an adaptive loop filtering (ALF) process, and / or a bilateral filtering process, as described above, some or all of which may be omitted. Additionally, one or some of the deblocking filtering process, the sample adaptive offset (SAO) process, the adaptive loop filtering (ALF) process, and / or the bilateral filtering process may be applied sequentially, or all of them may be applied sequentially. For example, after applying the deblocking filtering process to the reconstructed picture, the SAO process may be performed. Alternatively, for example, after applying the deblocking filtering process to the reconstructed picture, the ALF process may be performed. This may even be similarly performed in an encoding device.
[0208] Figure 19 An example of a schematic picture encoding process to which embodiments of the present disclosure are applicable is shown. In Figure 19 , S1910 may be performed in a prediction unit including the intra prediction unit 185 or the inter prediction unit 180 of the encoding device described above with reference to Figure 2 . S1920 may be performed in a residual processor including the transformer 120 and / or the quantizer 130. S1930 may be performed in the entropy encoder 190. S1910 may include the inter / intra prediction process described in the present disclosure. S1920 may include the residual process described in the present disclosure. S1930 may include the information encoding process described in the present disclosure.
[0209] Referring to Figure 19 , the picture encoding process may schematically include not only a process for encoding and outputting information for picture reconstruction (e.g., prediction information, residual information, segmentation information, etc.) in the form of a bitstream, but also a process for generating a reconstructed picture of the current picture and a process for applying in-loop filtering to the reconstructed picture (optional), as described with respect to Figure 2 . The encoding device may derive (modified) residual samples from the quantized transform coefficients through the dequantizer 140 and the inverse transformer 150, and generate a reconstructed picture based on the prediction samples and the (modified) residual samples that are the output of S1910. The reconstructed picture generated in this way may be the same as the reconstructed picture generated in the decoding device. Similar to the decoding device, the modified reconstructed picture may be generated through an in-loop filtering process for the reconstructed picture, may be stored in the decoded picture buffer or the memory 170, and may be used as a reference picture in the inter prediction process when encoding a picture later. As described above, in some cases, some or all of the in-loop filtering process may be omitted. When the in-loop filtering process is performed, the (in-loop) filtering-related information (parameters) may be encoded in the entropy encoder 190 and output in the form of a bitstream, and the decoding device may perform the in-loop filtering process using the same method as the encoding device based on the filtering-related information.
[0210] Through this in-loop filtering process, noise (e.g., block artifacts and ringing artifacts) that appears during image / video encoding can be reduced, and the subjective / objective visual quality can be improved. In addition, by performing the in-loop filtering process in both the encoding device and the decoding device, the encoding device and the decoding device can derive the same prediction result, the picture encoding reliability can be increased, and the amount of data to be transmitted for picture encoding can be reduced.
[0211] As described above, the picture reconstruction process can be performed not only in a decoding device but also in an encoding device. Reconstruction blocks can be generated in units of blocks based on intra prediction / inter - prediction, and a reconstructed picture including the reconstruction blocks can be generated. When the current picture / slice / tile group is an I picture / slice / tile group, the blocks included in the current picture / slice / tile group can be reconstructed based only on intra prediction. In addition, when the current picture / slice / tile group is a P or B picture / slice / tile group, the blocks included in the current picture / slice / tile group can be reconstructed based on intra prediction or inter - prediction. In this case, inter - prediction can be applied to some of the blocks in the current picture / slice / tile group, and intra prediction can be applied to the remaining blocks. The color components of a picture can include a luminance component and a chrominance component, and unless explicitly restricted in the present disclosure, the methods and embodiments of the present disclosure are applicable to both the luminance component and the chrominance component.
[0212] Example of coding layer and structure
[0213] For example, the encoded video / image according to the present disclosure can be processed according to the encoding layer and structure to be described below.
[0214] Figure 20 FIG. is a diagram showing the layer structure of an encoded image. The encoded image can be classified into a video coding layer (VCL) for image decoding processing and its own processing, a lower - layer system for transmitting and storing the encoded information, and a network abstraction layer (NAL) that exists between the VCL and the lower - layer system and is responsible for network adaptation functions.
[0215] In the VCL, VCL data including compressed image data (slice data) can be generated, or a supplementary enhancement information (SEI) message that is additionally required for the decoding processing of the image or a parameter set including information such as a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS) can be generated.
[0216] In the NAL, header information (NAL unit header) can be added to the raw byte sequence payload (RBSP) generated in the VCL to generate a NAL unit. In this case, the RBSP refers to slice data, parameter sets, and SEI messages generated in the VCL. The NAL unit header can include NAL unit type information specified according to the RBSP data included in the corresponding NAL unit.
[0217] As shown in the figure, the NAL unit can be classified into a VCL NAL unit and a non - VCL NAL unit according to the RBSP generated in the VCL. The VCL NAL unit can mean a NAL unit including information about the image (slice data), and the non - VCL NAL unit can mean a NAL unit including information required for decoding the image (parameter set or SEI message).
[0218] VCL NAL units and non-VCL NAL units can attach header information and be sent over a network according to the data standards of the underlying system. For example, the NAL unit can be modified to a data format of a predetermined standard such as the H.266 / VVC file format, RTP (Real-Time Transport Protocol), or TS (Transport Stream), and sent over various networks.
[0219] As described above, in the NAL unit, the NAL unit type can be specified according to the RBSP data structure included in the corresponding NAL unit, and information about the NAL unit type can be stored in the NAL unit header and signaled.
[0220] For example, this can be roughly classified into VCL NAL unit types and non-VCL NAL unit types according to whether the NAL unit includes information about an image (slice data). The VCL NAL unit type can be classified according to the nature and type of the picture included in the VCL NAL unit, and the non-VCL NAL unit type can be classified according to the type of parameter set.
[0221] Examples of NAL unit types specified according to the type of parameter set / information included in the non-VCL NAL unit type will be listed below.
[0222] - DCI (Decoding Capability Information) NAL unit: The NAL unit type including DCI
[0223] - VPS (Video Parameter Set) NAL unit: The NAL unit type including VPS
[0224] - SPS (Sequence Parameter Set) NAL unit: The NAL unit type including SPS
[0225] - PPS (Picture Parameter Set) NAL unit: The NAL unit type including PPS
[0226] - APS (Adaptive Parameter Set) NAL unit: The NAL unit type including APS
[0227] - PH (Picture Header) NAL unit: The NAL unit type including PH
[0228] The above NAL unit types can have syntax information of the NAL unit type, and this syntax information can be stored in the NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified as the nal_unit_type value.
[0229] In addition, as described above, a picture may include a plurality of slices, and a slice may include a slice header and slice data. In this case, a picture header may be further added to the plurality of slices (slice header and slice data set) in a picture. The picture header (picture header syntax) may include information / parameters that are commonly applicable to the picture.
[0230] The slice header (slice header syntax) may include information / parameters that are commonly applicable to the slice. APS (APS syntax) or PPS (PPS syntax) may include information / parameters that are commonly applicable to one or more slices or pictures. SPS (SPS syntax) may include information / parameters that are commonly applicable to one or more sequences. VPS (VPS syntax) may include information / parameters that are commonly applicable to multiple layers. DCI (DCI syntax) may include information / parameters that are commonly applicable to the entire video. DCI may include information / parameters related to decoding capabilities. In the present disclosure, the high-level syntax (HLS) may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DCI syntax, picture header syntax, or slice header syntax. In addition, in the present disclosure, the low-level syntax (LLS) may include, for example, slice data syntax, CTU syntax, coding unit syntax, transform unit syntax, and the like.
[0231] In the present disclosure, the image / video information encoded in an encoding device and signaled to a decoding device in the form of a bitstream may include not only information related to intra-picture segmentation, intra / inter-frame prediction information, residual information, in-loop filtering information, but also information about the slice header, information about the picture header, information about APS, information about PPS, information about SPS, information about VPS, and / or information about DCI. Additionally, the image / video information may further include general constraint information and / or information about the NAL unit header.
[0232] Using sub - pictures, slices and tiles to segment the picture
[0233] A picture may be divided into at least one tile row and at least one tile column. A tile may be composed of a sequence of CTUs and may cover a rectangular area of a picture.
[0234] A slice may be composed of an integer number of complete tiles or an integer number of consecutive complete CTU rows in a picture.
[0235] For slicing, two modes can be supported: one mode can be called the raster scan slicing mode, and the other mode can be called the rectangular slicing mode. In the raster scan slicing mode, a slice can include a complete sequence of tiles present in a picture in tile raster scan order. In the rectangular slicing mode, a slice can include multiple complete tiles assembled to form a rectangular area of the picture or multiple consecutive complete CTU rows of a single tile assembled to form a rectangular area of the picture. The tiles in a rectangular slice can be scanned in tile raster scan order in the rectangular area corresponding to the slice. A sub-picture can include at least one slice assembled to cover a rectangular area of the picture.
[0236] To describe the segmentation relationship of the picture in more detail, reference will be made to Figures 21 to 24 give a description. Figures 21 to 24 An embodiment of segmenting a picture using tiles, slices, and sub-pictures is shown. Figure 21 An example of a picture segmented into 12 tiles and three raster scan slices is shown. Figure 22 An example of a picture segmented into 24 tiles (six tile columns and four tile rows) and nine rectangular slices is shown. Figure 23 An example of a picture segmented into four tiles (two tile columns and two tile rows) and four rectangular slices is shown.
[0237] Figure 24 An example of segmenting a picture into sub-pictures is shown. In Figure 24 , the picture can be segmented into 12 left tiles covering a slice composed of 4×4 CTUs and six right tiles covering two vertically assembled slices composed of 2×2 CTUs, such that a picture is segmented into 24 slices and 24 sub-pictures with different regions. In the example of Figure 24 , an individual slice corresponds to an individual sub-picture.
[0238] Overview of in - loop filtering
[0239] An in-loop filtering process can be performed on the reconstructed picture generated through the above process. A modified reconstructed picture can be generated through the in-loop filtering process, and the modified reconstructed picture can be output as a decoded picture from the decoding device, can be stored in the memory of the encoding device / decoding device or the decoded picture buffer, and can be used as a reference picture in the inter prediction process when encoding / decoding a picture. As described above, the in-loop filtering process can include a deblocking filtering process, a sample adaptive offset (SAO) process, and / or an adaptive loop filtering (ALF) process. In this case, one or some of the deblocking filtering process, the sample adaptive offset (SAO) process, the adaptive loop filtering (ALF) process, and the bilateral filtering process can be applied in sequence, or all of them can be applied in sequence. For example, after applying the deblocking filtering process to the reconstructed picture, the SAO process can be performed. Alternatively, for example, after applying the deblocking filtering process to the reconstructed picture, the ALF process can be performed. This can also be performed in the encoding device.
[0240] Deblocking filtering is a filtering technique for removing distortions at the boundaries between blocks in the reconstructed picture. The deblocking filtering process can, for example, derive a target boundary from the reconstructed picture, determine the boundary strength bS of the target boundary, and perform deblocking filtering on the target boundary based on bS. The bS can be determined based on the prediction modes of two blocks adjacent to the target boundary, the motion vector difference, whether the reference pictures are the same, whether there are non-zero valid coefficients, etc.
[0241] SAO is a method for compensating the offset difference between the reconstructed picture and the original picture on a sample-by-sample basis, and can be applied based on types such as band offset and edge offset. According to SAO, samples can be classified into different categories according to each SAO type, and an offset value can be added to each sample based on the category. The filtering information of SAO can include information on whether SAO is applied, SAO type information, SAO offset value information, etc. After applying deblocking filtering, SAO can be applied to the reconstructed picture.
[0242] Adaptive loop filtering (ALF) is a technique for filtering the reconstructed picture on a sample-by-sample basis based on filtering coefficients according to the filtering shape. The encoding device can determine whether to apply ALF, the ALF shape, and / or the ALF filtering coefficients, etc. by comparing the reconstructed picture with the original picture, and can signal them to the decoding device. That is, the filtering information of ALF can include information on whether ALF is applied, ALF filtering shape information, ALF filtering coefficient information, etc. After applying deblocking filtering, ALF can be applied to the reconstructed picture.
[0243] Signaling of whether to apply in - loop filtering
[0244] As described above, HLS can be encoded and / or signaled for video and / or image encoding. As described above, the video / image information in this specification can be included in HLS. Additionally, an image / video encoding method can be performed based on such image / video information.
[0245] In an embodiment, a picture can be divided into segmentation units such as sub-pictures, slices, and / or tiles. Additionally, it can be signaled whether to apply filtering to the boundaries of such segmentation units. For example, a picture can be divided into multiple tiles. In this case, in-loop filtering across the boundaries of the tiles can be performed. Alternatively, a picture can be divided into multiple slices. In this case, in-loop filtering across the boundaries of the slices can be performed. In this case, whether to perform in-loop filtering across the boundaries of the tiles and / or slices can be signaled via HLS.
[0246] Figure 25 is a view illustrating an embodiment of the syntax of PPS for signaling whether in-loop filtering is performed at the boundaries of tiles and / or slices. In Figure 25 In the syntax, the syntax element no_pic_partition_flag2510 can indicate whether picture segmentation is applied to an individual picture of the reference PPS. For example, a first value (e.g., 0) of no_pic_partition_flag2510 can indicate that an individual picture of the reference PPS can be divided into more than one tile or slice. A second value (e.g., 1) of no_pic_partition_flag 2510 can indicate that an individual picture of the reference PPS is not divided. Additionally, no_pic_partition_flag 2510 can be restricted to have the same value for all PPSs present in a sequence.
[0247] In an embodiment, the syntax element no_pic_partition_flag 2510 can be used as a condition for signaling information for dividing tiles and / or slices when a picture is divided into more than one tile and / or slice, and this can be included in the syntax as a conditional statement, as in Figure 252520 of reference numeral 2520. For example, the encoding device may use no_pic_partition_flag 2510 to signal the decoding device whether the information about the partition of the patch and / or slice is included in the bitstream. In addition, when the value of no_pic_partition_flag 2510 is 1, the decoding device may not parse the information about the partition of the patch and / or slice from the bitstream. When the value of no_pic_partition_flag 2510 is 0, the decoding device may parse the information about the partition of the patch and / or slice from the bitstream according to the additional information.
[0248] For the above implementation, when no_pic_partition_flag 2510 indicates that the picture can be partitioned into tiles or slices, the PPS syntax indicating that the following syntax elements can be obtained from the bitstream can be implemented as Figure 25 .
[0249] The syntax element pps_log2_ctu_size_minus5 may indicate the luma coding tree block size of the individual CTU. Specifically, the encoding device may determine the value of pps_log2_ctu_size_minus5 as a value obtained by subtracting 5 from the luma coding tree block size of the individual CTU. The decoding device may determine the value obtained by adding 5 to pps_log2_ctu_size_minus5 as the luma coding tree block size of the individual CTU.
[0250] The syntax element num_exp_tile_columns_minus1 may indicate the numerical value of the width value of the tile column explicitly signaled. For example, the decoding device may determine the numerical value of the width value of the tile column explicitly signaled as a value obtained by adding 1 to num_exp_tile_columns_minus1. The value of num_exp_tile_columns_minus1 may have a value from 0 to PicWidthInCtbY-1. Here, PicWidthInCtbY may indicate the width of the picture expressed in units of the width of the luma coding block. In addition, when the value of no_pic_partition_flag is 1, the value of num_exp_tile_columns_minus1 may be derived as 0.
[0251] The syntax element num_exp_tile_rows_minus1 can indicate the numerical value of the height value of the explicitly signaled tile rows. For example, the decoding device can determine the numerical value of the height value of the explicitly signaled tile rows as the value obtained by adding 1 to num_exp_tile_rows_minus1. The value of num_exp_tile_rows_minus1 can have a value from 0 to PicHeightInCtbY - 1. PicHeightInCtbY can indicate the height of the picture represented in units of the height of the luminance coded blocks. When the value of no_pic_partition_flag is 1, the value of num_exp_tile_rows_minus1 can be derived as 0.
[0252] The syntax element tile_column_width_minus1[i] can indicate the width of the i-th tile column of the picture with reference to the PPS. For example, the decoding device can determine the value obtained by adding 1 to tile_column_width_minus1[i] as the width of the i-th tile column. The syntax element tile_column_width_minus1[i] can be obtained from the bitstream based on the value of num_exp_tile_columns_minus1, as in the syntax of Figure 25 as follows.
[0253] The syntax element tile_row_height_minus1[i] can indicate the height of the i-th tile row of the picture with reference to the PPS. For example, the decoding device can determine the value obtained by adding 1 to tile_row_height_minus1[i] as the height of the i-th tile row. The syntax element tile_row_height_minus1[i] can be obtained from the bitstream based on the value of num_exp_tile_rows_minus1, as in the syntax of Figure 25 as follows.
[0254] In addition, the variable NumTilesInPic can be calculated based on the values of num_exp_tile_columns_minus1 and num_exp_tile_rows_minus1. In an implementation, the decoding device can determine the value of the variable NumTilesInPic indicating the number of tiles in the picture with reference to the PPS as the value of (num_exp_tile_columns_minus1 + 1) × (num_exp_tile_rows_minus1 + 1).
[0255] When the value of NumTilesInPic is greater than 1, the syntax element rect_slice_flag can be obtained from the bitstream. For example, when the picture is divided into two or more tiles, the syntax element rect_slice_flag can be obtained.
[0256] The syntax element rect_slice_flag can indicate whether to apply the raster scan slice mode or the rectangular slice mode to an individual picture of the reference PPS. For example, the first value of rect_slice_flag (e.g., 0) can indicate that the raster scan slice mode is applied to an individual picture of the reference PPS. In this case, the signaling of the slice layout can be omitted. The second value of rect_slice_flag (e.g., 1) can indicate that the rectangular slice mode is used for an individual picture of the reference PPS. In this case, as described below, the slice layout can be signaled by the PPS. When rect_slice_flag is not signaled, the decoding device can deduce the value of rect_slice_flag as 1.
[0257] When the value of the syntax element rect_slice_flag is 1, the syntax element single_slice_per_subpic_flag can be signaled. The first value of single_slice_per_subpic_flag (e.g., 0) can indicate that an individual subpicture can consist of more than one rectangular slice. The second value of single_slice_per_subpic_flag (e.g., 1) can indicate that an individual subpicture consists of only one rectangular slice.
[0258] In addition, when the value of rect_slice_flag is 1 and the value of single_slice_per_subpic_flag is 0, the syntax element num_slices_in_pic_minus1 can be signaled. For example, when the picture is divided into two or more rectangular slices, the syntax element num_slices_in_pic_minus1, which indicates the number of rectangular slices in an individual picture of the reference PPS, can be signaled to signal the rectangular slice layout. For example, the decoding device can determine the number of rectangular slices in the picture as the value obtained by adding 1 to num_slices_in_pic_minus1. The value of num_slices_in_pic_minus1 can have a value from 0 to MaxSlicePerAu - 1. The variable MaxSlicePerAu can indicate the maximum number of slices allowed per access unit, for example, the maximum number of slices allowed in the current picture.
[0259] Based on the value of num_slices_in_pic_minus1, it is also possible to obtain the syntax elements tile_idx_delta_present_flag, slice_width_in_tiles_minus1, slice_height_in_tiles_minus1, num_exp_slices_in_tile, exp_slice_height_in_ctus_minus1, and tile_idx_delta in the same syntax as Figure 25 . Here, the syntax element tile_idx_delta_present_flag can indicate whether the syntax element tile_idx_delta, which is used as an index for identifying rectangular slices in the picture, is obtained from the bitstream. The value obtained by adding 1 to the syntax element slice_width_in_tiles_minus1[i] can represent the width of the i-th rectangular slice in terms of tile columns. The value obtained by adding 1 to the syntax element slice_height_in_tiles_minus1[i] can indicate the height of the i-th rectangular slice in terms of tile rows.
[0260] The syntax element num_exp_slices_in_tile[i] can indicate the numerical value of the slice height explicitly provided for the slices in the tile including the i-th slice. The value obtained by adding 1 to the syntax element exp_slice_height_in_ctus_minus1[j] can indicate the height of the j-th rectangular slice in the tile including the i-th slice, and the unit can be in terms of CTU rows. The syntax element tile_idx_delta[i] can indicate the difference between the index of the tile including the first CTU in the i-th rectangular slice and the index of the tile including the first CTU in the (i + 1)-th rectangular slice.
[0261] The syntax element loop_filter_across_tiles_enabled_flag 2530 can indicate whether a filtering operation is performed across the boundaries of tiles in the picture of the reference PPS. For example, the first value of loop_filter_across_tiles_enabled_flag (e.g., 0) indicates that within-loop filtering operations are not performed across the boundaries of tiles in the picture of the PPS that includes this syntax element. The second value of loop_filter_across_tiles_enabled_flag (e.g., 1) indicates that within-loop filtering operations can be performed across the boundaries of tiles in the picture of the PPS that includes this syntax element.
[0262] Here, the in-loop filtering operation may include deblocking filtering, sample adaptive offset (SAO) filtering, and / or adaptive loop filtering (ALF). When the value of loop_filter_across_tiles_enabled_flag is not obtained from the bitstream (e.g., not provided), the value of the syntax element may be derived as a second value (e.g., 1). Additionally, in another embodiment, when the value of loop_filter_across_tiles_enabled_flag is not obtained from the bitstream (e.g., not provided), the value of the syntax element may be derived as a first value (e.g., 0).
[0263] The syntax element loop_filter_across_slices_enabled_flag 2540 may indicate whether to perform a filtering operation across the boundaries of slices in a picture of a reference PPS. For example, a first value of loop_filter_across_slices_enabled_flag (e.g., 0) may indicate that the in-loop filtering operation is not performed across the boundaries of slices in a picture of the reference PPS that includes this syntax element.
[0264] A second value of loop_filter_across_slices_enabled_flag (e.g., 1) may indicate that the in-loop filtering operation may be performed across the boundaries of slices in a picture of the reference PPS that includes this syntax element.
[0265] Here, as described above, the in-loop filtering operation may include deblocking filtering, SAO filtering, and / or ALF. When the value of loop_filter_across_slices_enabled_flag is not obtained from the bitstream (e.g., not provided), the value of the syntax element may be derived as a first value (e.g., 0).
[0266] In an embodiment, each picture may be divided in units of tiles. In this case, there may be two or more tiles in a picture, and it may be determined whether to apply in-loop filtering to the boundary portions of each tile region by the loop_filter_across_tiles_enabled_flag 2530 signaled in an individual PPS.
[0267] In Figure 25In the example, when there are two or more tiles in a picture (e.g., no_pic_partition_flag == 0), the value of loop_filter_across_tiles_enabled_flag can always be signaled to determine whether to apply in-loop filtering in the boundary region. However, considering the fact that the flag is signaled even when there are only multiple tiles in a picture, an improvement can be made such that the flag can be signaled considering the number of tiles present, in order to reduce the amount of bits sent.
[0268] For example, in Figure 25 the example, when the value of no_pic_partition_flag is the first value (e.g., 0), this means that the current picture is partitioned into more than one tile or slice. In an embodiment, when the value of no_pic_partition_flag is the first value (e.g., 0), since there can be two or more tiles in a picture, the value of loop_filter_across_tiles_enabled_flag can always be signaled to determine whether to apply in-loop filtering in the tile boundary region. However, the case where the value of no_pic_partition_flag is the first value (e.g., 0) includes the following situation: a picture is not partitioned into tiles but only into slices. Therefore, even when there are no tiles and only multiple slices in a picture, loop_filter_across_tiles_enabled_flag can be signaled. Considering this, an improvement can be made to signal the flag considering the number of tiles present, in order to reduce the amount of bits sent.
[0269] Figure 26 is a view illustrating an embodiment of the syntax for signaling the syntax element loop_filter_across_tiles_enabled_flag considering the number of tiles to solve the above problems. As Figure 26 shown, the syntax element loop_filter_across_tiles_enabled_flag 2620 can be signaled based on the number of tiles belonging to the picture (e.g., NumTilesInPic). For example, as Figure 26As shown, only when the number of tiles belonging to a picture is greater than 1 (2610), can loop_filter_across_tiles_enabled_flag 2620 be signaled via a bitstream. Therefore, only when the number of tiles belonging to a picture is greater than 1 (2610), can an encoding device encode loop_filter_across_tiles_enabled_flag 2620 into a bitstream, and a decoding device can obtain loop_filter_across_tiles_enabled_flag 2620 from the bitstream.
[0270] In addition, in Figure 25 the embodiment, each picture can be divided in units of slices. In this case, two or more slices can exist in a picture, and it can be determined whether to apply in-loop filtering to the boundary parts of each slice region by loop_filter_across_slices_enabled_flag signaled in an individual PPS. More specifically, when two or more slices exist in a picture (e.g., no_pic_partition_flag == 0), the value of loop_filter_across_slices_enabled_flag can always be signaled to determine whether to apply in-loop filtering in the boundary region. However, considering the fact that the flag is signaled even when only multiple slices exist in a picture, an improvement can be made such that the flag can be signaled considering the number of existing slices, in order to reduce the amount of bits transmitted.
[0271] In the embodiment, even when the value of no_pic_partition_flag is the first value (e.g., 0), the picture of the reference PPS can be divided only into tiles and not into slices. However, in Figure 25 the example, even when the picture is divided into tiles without being divided into slices as described above, the value of loop_filter_across_slices_enabled_flag is always signaled. In this way, considering the fact that loop_filter_across_slices_enabled_flag is signaled even when only multiple tiles exist in a picture, the Figure 25 embodiment can be improved such that the flag can be signaled considering the number of existing slices, in order to reduce the amount of bits transmitted.
[0272] Figure 27This is a view of an implementation of the syntax that signals the syntax element loop_filter_across_slices_enabled_flag while considering the number of slices in order to solve the above problems. As Figure 27 shown, the syntax element loop_filter_across_slices_enabled_flag 2720 can be signaled based on the number of slices belonging to a picture (e.g., num_slices_in_pic_minus1). For example, as Figure 27 shown, the loop_filter_across_slices_enabled_flag 2720 can be signaled through the bitstream only when the number of slices belonging to a picture is greater than 1 (2710). Therefore, only when the number of slices belonging to a picture is greater than 1 (2710), the encoding device can encode the loop_filter_across_slices_enabled_flag 2720 to generate a bitstream, and the decoding device can obtain the loop_filter_across_slices_enabled_flag 2720 from the bitstream.
[0273] In addition, num_slices_in_pic_minus1 is a syntax element that is signaled when a picture is divided into two or more rectangular slices to indicate the number of rectangular slices in an individual picture that refers to the PPS, thereby signaling the layout of the rectangular slices. When the raster scan slice mode is applied to an individual picture, the individual picture can still be divided into multiple slices. Therefore, the syntax can be modified so that the loop_filter_across_slices_enabled_flag 2720 can be signaled even when the value of rect_slice_flag indicates the first value (e.g., 0) that indicates the raster scan slice mode.
[0274] Furthermore, when the value of single_slice_per_subpic_flag is the second value (e.g., 1), since a picture can be divided into multiple sub-pictures, a picture can be composed of multiple slices. Therefore, the syntax can be modified so that the loop_filter_across_slices_enabled_flag 2720 can be signaled even when the value of single_slice_per_subpic_flag indicates the second value (e.g., 1).
[0275] In this way, when the value of num_slices_in_pic_minus1 is not obtained from the bitstream, since it is not signaled how many slices the current picture is partitioned into, the syntax can be modified so that the loop_filter_across_slices_enabled_flag 2720 can be signaled even when the value of num_slices_in_pic_minus1 is not obtained from the bitstream. For example, in Figure 27 's example, as a condition for obtaining num_slices_in_pic_minus1 from the bitstream, it is required that the value of rect_slice_flag is 1 and the value of single_slice_per_subpic_flag is 0. In this regard, when the value of rect_slice_flag is 0 or the value of single_slice_per_subpic_flag is 1, the value of loop_filter_across_slices_enabled_flag can be obtained from the bitstream regardless of whether the value of num_slices_in_pic_minus1 is greater than 1. For this process, the PPS syntax can be modified as shown in Figure 28 .
[0276] Figure 28 is an example application reference Figures 25 to 27 describes a view of an implementation of the syntax of the PPS that signals the loop_filter_across_tiles_enabled_flag and the loop_filter_across_slices_enabled_flag. In Figure 28 's implementation, in some of the names of the syntax elements described in reference Figure 25 and Figure 28 , pps_ is added. For example, the above syntax no_pic_partition_flag is named pps_no_pic_partition_flag.
[0277] Reference Figure 28, if the pps_no_pic_partition_flag has a value (e.g., 0) indicating that the picture can be partitioned into at least one of patches or slices, the syntax elements pps_num_exp_tile_columns_minus1 indicating how many patch columns the picture of the reference PPS has and pps_num_exp_tile_rows_minus1 indicating how many patch rows the picture of the reference PPS has can be used to signal how many patches the current picture is partitioned into. Then, the number of patches included in the current picture can be calculated as (pps_num_exp_tile_columns_minus1 + 1) × (pps_num_exp_tile_rows_minus1 + 1), and is recorded in the variable NumTileInPic.
[0278] As described above, only when the value of NumTileInPic is greater than 1, that is, only when there is more than one patch in the current picture, can the pps_loop_filter_across_tiles_enabled_flag syntax element and the pps_rect_slice_flag syntax element indicating whether filtering can be applied across the boundaries of patches be obtained. Additionally, when the value of pps_rect_slice_flag is 1, the pps_single_slice_per_subpic_flag can be obtained from the bitstream. When the value of pps_rect_slice_flag is 1 and the value of pps_single_slice_per_subpic_flag is 0, the pps_num_slice_in_pic_minus1 syntax element can be obtained from the bitstream. Additionally, when the value of pps_rect_slice_flag is 0, the value of pps_single_slice_per_subpic_flag is 1 or the value of pps_num_slices_in_pic_minus1 is greater than 1, the value of pps_loop_filter_across_slices_enabled_flag can be obtained from the bitstream.
[0279] Coding and decoding methods
[0280] Hereinafter, an image encoding and decoding method performed by an image encoding and decoding device according to an embodiment will be described.
[0281] First, the operation of the decoding device will be described. An image decoding device according to an embodiment includes a memory and a processor, and the decoding device can perform decoding according to the operation of the processor. Figure 29Illustrates a decoding method of a decoding device according to an embodiment.
[0282] The decoding device according to the embodiment may determine the number of tiles (e.g., NumTilesInPic) in the current picture based on the unrestricted segmentation of the current picture (S2910). For example, the decoding device may obtain a segmentation restriction flag (e.g., no_pic_partition_flag) indicating whether the segmentation of the current picture is restricted from the bitstream, and may determine whether the segmentation of the current picture is restricted based on the segmentation restriction flag.
[0283] Next, the decoding device may obtain a first flag (e.g., loop_filter_across_tiles_enabled_flag) indicating whether filtering of the boundaries of the tiles is available from the bitstream based on the number of tiles in the current picture being plural (S2920). Here, the number of tiles in the current picture may be determined based on the tile number information indicating the number of tiles partitioning the current picture. Here, the tile number information may be obtained from the bitstream based on the unrestricted segmentation of the current picture. In addition, the tile number information may include information indicating the number of tile columns in the picture (e.g., num_exp_tile_columns_minus1) and information indicating the number of tile rows in the picture (e.g., num_exp_tile_rows_minus1).
[0284] Next, the decoding device may determine whether to perform filtering on the boundaries of the tiles belonging to the current picture based on the value of the first flag (S2930). Here, the type of filtering may be any one of the deblocking filtering, SAO filtering, and ALF filtering described above. For example, when the first flag indicates that filtering is not available, the filtering used for decoding the corresponding image among the deblocking filtering, SAO filtering, and ALF filtering may not be applied to the boundaries of the tiles.
[0285] In addition, the decoding device may obtain a second flag (e.g., loop_filter_across_slices_enabled_flag) indicating whether filtering of the boundaries of the slices is available from the bitstream based on the unrestricted segmentation of the current picture (S2940).
[0286] For example, the decoding device may obtain information about the slices constituting the picture from the bitstream based on the unrestricted segmentation of the current picture. In addition, the decoding device may obtain the second flag from the bitstream based on the information about the slices not indicating that the picture consists of one slice.
[0287] Alternatively, the decoding device may obtain a second flag from the bitstream based on information about a slice indicating that the rectangular slice mode is not applied to the picture (e.g., rect_slice_flag == 0). Alternatively, the decoding device may obtain a second flag from the bitstream based on information about a slice indicating that a sub-picture of the picture consists of only one rectangular slice (e.g., rect_slice_flag == 0 or pps_single_slice_per_subpic_flag == 1).
[0288] Alternatively, the decoding device may obtain a second flag from the bitstream based on information about a slice indicating that the number of slices in the current picture is more than one (e.g., num_slices_in_pic_minus1 > 0). For example, based on the unrestricted segmentation of the current picture, it can be determined whether the slices constituting the picture are rectangular slices. Based on the slices constituting the picture being rectangular slices, it can be determined whether a sub-picture of the picture consists of only one rectangular slice. Information indicating the number of slices in the current picture can be obtained from the bitstream based on a sub-picture of the picture consisting of more than one rectangular slice, and based on the information indicating the number of slices in the current picture, it can be determined whether the number of slices in the current picture is more than one.
[0289] Then, the decoding device may determine whether to perform filtering on the boundaries of the slices belonging to the current picture based on the value of the second flag (S2950). For example, when the second flag indicates that filtering is not available, the filtering used for decoding the corresponding image among deblocking filtering, SAO filtering, and ALF filtering may not be applied to the boundaries of the slices.
[0290] Next, the operation of the encoding device will be described. The image encoding device according to an embodiment includes a memory and a processor, and the encoding device may perform encoding in a manner corresponding to the decoding of the decoding device through the operation of the processor. For example, as Figure 30 shown, the encoding device may determine the number of tiles in the current picture (e.g., NumTilesInPic) based on the unrestricted segmentation of the current picture (S3010). Next, the encoding device may determine the value of a first flag (e.g., loop_filter_across_tiles_enabled_flag) indicating whether filtering on the boundaries of the tiles is available based on the number of tiles in the current picture being more than one (S3020). In addition, the encoding device may also determine whether the current picture consists of one slice based on the unrestricted segmentation of the current picture (S3030). Additionally, the encoding device may determine the value of a second flag (e.g., loop_filter_across_slices_enabled_flag) indicating whether filtering on the boundaries of the slices is available based on the current picture not consisting of one slice (S3040).
[0291] Next, the encoding device may generate a bitstream that includes at least one of the first flag or the second flag or does not include any of them (S3050). For example, the encoding device may not determine the values of the first flag and the second flag based on the number of picture partitions in the picture not being multiple and the current picture consisting of one slice, and may generate a bitstream that does not include the first flag and the second flag. Additionally, the value of the partitioning restriction flag (e.g., no_pic_partition_flag) may be set according to whether the current picture partitioning is restricted, and the partitioning restriction flag may also be included in the bitstream.
[0292] As described above, since no_pic_partition_flag is utilized in the encoding and decoding methods, it signals whether the current picture is partitioned into partitions and / or slices. Additionally, accordingly, information about the partitioning of partitions and information about the number of slices are signaled. In this regard, if it is only determined whether to signal loop_filter_across_tiles_enabled_flag and loop_filter_across_slices_enabled_flag based only on the value of no_pic_partition_flag, then when the current picture is partitioned into only slices or only partitions, it is not necessary to signal loop_filter_across_tiles_enabled_flag or loop_filter_across_slices.
[0293] Additionally, in order to reduce the signaling of loop_filter_across_tiles_enabled_flag and loop_filter_across_slices_enabled_flag, signaling of the flag indicating whether the current picture is partitioned into partitions and the flag indicating whether the current picture is partitioned into slices together with no_pic_partition_flag does not help in terms of bit reduction.
[0294] However, as described in this specification, together with the no_pic_partition_flag, the configuration for determining whether to signal the loop_filter_across_tiles_enabled_flag and the loop_filter_across_slices_enabled_flag based on the number of tiles of the partitioned picture and the number of slices of the partitioned picture enables the determination of whether the loop_filter_across_tiles_enabled_flag and the loop_filter_across_slices_enabled_flag are signaled from the parsing information of the tiles and slices without an additional flag. Therefore, the technical idea described in this disclosure can reduce the frequency of generating corresponding flags in the bitstream in an encoding / decoding environment where the current picture can be partitioned into tiles and / or slices, thereby reducing the size of the bitstream.
[0295] Application implementation
[0296] Although, for clarity of description, the above-described exemplary methods of the present disclosure are represented as a series of operations, they are not intended to limit the order of execution of the steps, and these steps can be performed simultaneously or in a different order when necessary. To implement the methods according to the present disclosure, the described steps may further include other steps, may include the remaining steps except some steps, or may include other additional steps except some steps.
[0297] In the present disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) may perform an operation (step) of confirming the execution conditions or circumstances of the corresponding operation (step). For example, if it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or the image decoding device may perform the predetermined operation after determining whether the predetermined condition is satisfied.
[0298] The various embodiments of the present disclosure are not a list of all possible combinations and are intended to describe representative aspects of the present disclosure, and the matters described in the various embodiments can be applied independently or in combinations of two or more.
[0299] The various embodiments of the present disclosure can be implemented in hardware, firmware, software, or a combination thereof. In the case where the present disclosure is implemented by hardware, the present disclosure can be implemented by an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a general purpose processor, a controller, a microcontroller, a microprocessor, or the like.
[0300] In addition, the image decoding device and the image encoding device according to the embodiments of the present disclosure may be included in a multimedia broadcast transmission and reception device, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camera, a video-on-demand (VoD) service providing device, an over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone video device, a medical video device, etc., and may be used to process video signals or data signals. For example, the OTT video device may include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.
[0301] Figure 31 is a view showing a content streaming system to which embodiments of the present disclosure can be applied.
[0302] As Figure 31 shown, the content streaming system to which embodiments of the present disclosure are applied may mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.
[0303] The encoding server compresses the content input from a multimedia input device such as a smart phone, a camera, a video camera, etc. into digital data to generate a bitstream and sends the bitstream to the streaming server. As another example, when a multimedia input device such as a smart phone, a camera, a video camera, etc. directly generates a bitstream, the encoding server may be omitted.
[0304] The bitstream may be generated by an image encoding method or an image encoding device according to the embodiments of the present disclosure, and the streaming server may temporarily store the bitstream during the process of sending or receiving the bitstream.
[0305] The streaming server sends multimedia data to the user device based on a request from the user through the network server, and the network server serves as a medium for informing the user of the service. When the user requests a required service from the network server, the network server may deliver it to the streaming server, and the streaming server may send multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server is used to control commands / responses between devices in the content streaming system.
[0306] The streaming server may receive content from the media storage device and / or the encoding server. For example, when receiving content from the encoding server, the content may be received in real time. In this case, in order to provide a smooth streaming service, the streaming server may store the bitstream for a predetermined time.
[0307] Examples of user devices may include mobile phones, smartphones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet computers, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays), digital TVs, desktop computers, digital signage, etc.
[0308] Each server in the content streaming system may operate as a distributed server, in which case the data received from each server may be distributed.
[0309] The scope of the present disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) for enabling the operation of methods according to various embodiments to be performed on a device or computer, and non-transitory computer-readable media having such software or commands stored thereon and executable on the device or computer.
[0310] Industrial Applicability
[0311] Embodiments of the present disclosure may be used to encode or decode images.
Claims
1. An image decoding method performed by an image decoding device, the image decoding method comprising the following steps: Determine the number of tiles in the current picture based on that the segmentation of the current picture is unrestricted; Based on the number of tiles in the current picture being multiple, obtain a first flag for filtering the boundaries of the tiles from the bitstream; Based on the value of the first flag, determine to perform filtering on the boundaries of the tiles belonging to the current picture; Based on the segmentation of the current picture being unrestricted, obtain a second flag for filtering the boundaries of the slices from the bitstream; And Based on the value of the second flag, determine to perform filtering on the boundaries of the slices belonging to the current picture, wherein, based on the segmentation of the current picture being unrestricted, obtain first information about the slices constituting the current picture from the bitstream, wherein, the second flag is obtained from the bitstream based on at least one of that the number of slices in the current picture is multiple indicated by the first information about the slices, that the rectangular slice mode is not applied to the current picture indicated by the second information about the slices, or that the sub - picture of the current picture consists of only one rectangular slice indicated by the third information about the slices.
2. The image decoding method according to claim 1, Among them, Obtain a segmentation restriction flag for the segmentation of the current picture from the bitstream, and wherein, determine that the segmentation of the current picture is restricted based on the segmentation restriction flag.
3. The image decoding method according to claim 1, wherein, Determine the number of tiles in the current picture based on the tile number information of the number of tiles for segmenting the current picture.
4. The image decoding method according to claim 3, wherein, The tile number information includes information about the number of tile columns in the current picture and information about the number of tile rows in the current picture.
5. The image decoding method according to claim 3, wherein, Based on the segmentation of the current picture being unrestricted, obtain the tile number information from the bitstream.
6. The image decoding method according to claim 1, Among them, Based on the segmentation of the current picture being unrestricted, determine whether the slices constituting the current picture are rectangular slices, wherein, based on whether the slices constituting the current picture are rectangular slices, determine whether the sub - picture of the current picture consists of only one rectangular slice, wherein, based on the sub - picture of the current picture consisting of more than one rectangular slice, obtain information about the number of slices in the current picture from the bitstream, and wherein, based on the information about the number of slices in the current picture, determine whether the number of slices in the current picture is multiple.
7. An image encoding method performed by an image encoding device, the image encoding method comprising the following steps: Determine the number of tiles in the current picture based on that the segmentation of the current picture is unrestricted; Based on the number of tiles in the current picture being multiple, determine the value of a first flag for filtering the boundaries of the tiles; Generate a bitstream including the first flag; Based on the segmentation of the current picture being unrestricted, determine whether the current picture consists of one slice as the first information; And Based on the current picture not consisting of one slice, determine the value of a second flag for filtering the boundaries of the slices, Wherein, the second flag is included in the bitstream based on at least one of: the first information about the slice indicating that the number of slices in the current picture is plural, the second information about the slice indicating that the rectangular slice mode is not applied to the current picture, or the third information about the slice indicating that the sub-picture of the current picture consists of only one rectangular slice.
8. A method for transmitting a bitstream generated by an image coding method, the image coding method comprising the steps of: Determining the number of tiles in the current picture based on the unrestricted segmentation of the current picture; Determining the value of a first flag for filtering the boundaries of the tiles based on the number of tiles in the current picture being plural; and Generating a bitstream including the first flag; Determining, as first information, whether the current picture consists of one slice based on the unrestricted segmentation of the current picture; Determining the value of a second flag for filtering the boundaries of the slices based on the current picture not consisting of one slice, Wherein, the second flag is included in the bitstream based on at least one of: the first information about the slice indicating that the number of slices in the current picture is plural, the second information about the slice indicating that the rectangular slice mode is not applied to the current picture, or the third information about the slice indicating that the sub-picture of the current picture consists of only one rectangular slice.