Image decoding and encoding method, method for transmitting bitstream, and medium

By introducing sub-pictures and picture header information logos into the bitstream, the image encoding and decoding process is optimized, and the problem of low encoding/decoding efficiency in high-resolution, high-quality image transmission and storage is solved, reducing transmission and storage costs.

CN120529094APending Publication Date: 2025-08-22LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510642149.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-01-14
Filing Date
2021-01-14
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The prior art has low encoding/decoding efficiency when transmitting and storing high-resolution, high-quality images, resulting in increased transmission and storage costs.

Method used

By introducing information marks about sub-pictures and logos of the picture header into the bitstream, the image encoding and decoding process includes encoding or decoding the sub-picture identifier in the slice header, and notifying the picture header information through a network abstract layer unit with the NAL unit type equal to PH_NUT.

Benefits of technology

Improve the efficiency of image encoding/decoding, reduce transmission and storage costs, and realize efficient information notification and storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120529094A_ABST
    Figure CN120529094A_ABST
Patent Text Reader

Abstract

The invention provides an image decoding and encoding method, a method for transmitting a bitstream, and a medium. An image decoding method according to the present disclosure may comprise the steps of: acquiring a first flag indicating whether information related to a sub-picture exists in a bitstream; acquiring a second mark indicating whether picture header information exists in the slice header or not; and decoding the bitstream based on the first flag and the second flag. If the first flag indicates that information related to a sub-picture is present in the bitstream, the second flag may have a value indicating that picture header information is not present in the slice header.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the original invention patent application with application number 202180019325.2 (International application number: PCT / KR2021 / 000515, application date: January 14, 2021, invention name: Image encoding / decoding method and device for signaling information related to sub-pictures and picture headers and method for sending bit streams). Technical Field

[0002] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly, to an image encoding and decoding method and apparatus for signaling information about a sub-picture and a picture header, and a method for transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure. Background Art

[0003] Recently, demand for high-resolution and high-quality images, such as high-definition (HD) and ultra-high-definition (UHD), is increasing across various fields. As the resolution and quality of image data improve, the amount of information or bits transmitted increases relative to existing image data. This increase in the amount of transmitted information or bits leads to increased transmission and storage costs.

[0004] Therefore, efficient image compression techniques are needed to effectively transmit, store, and reproduce information about high-resolution and high-quality images. Summary of the Invention

[0005] Technical issues

[0006] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0007] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus that improve encoding / decoding efficiency by efficiently signaling information about a sub-picture and a picture header.

[0008] Another object of the present disclosure is to provide a method for transmitting a bit stream generated by the image encoding method or apparatus according to the present disclosure.

[0009] Another object of the present disclosure is to provide a recording medium storing a bit stream generated by the image encoding method or apparatus according to the present disclosure.

[0010] Another object of the present disclosure is to provide a recording medium storing a bit stream received and decoded by the image decoding apparatus according to the present disclosure and used to reconstruct an image.

[0011] The technical problems solved by the present disclosure are not limited to the above-mentioned technical problems, and those skilled in the art will understand other technical problems not described here through the following description.

[0012] Technical Solution

[0013] An image decoding method performed by an image decoding apparatus according to one aspect of the present disclosure may include: acquiring a first flag that specifies whether information about a sub-picture exists in a bitstream; acquiring a second flag that specifies whether picture header information exists in a slice header; and decoding the bitstream based on the first flag and the second flag. When the first flag specifies that information about a sub-picture exists in the bitstream, the second flag may have a value that specifies that picture header information does not exist in the slice header.

[0014] In the image decoding method according to the present disclosure, when the first flag specifies that information about a sub-picture exists in the bitstream, the slice header may include an identifier of the sub-picture including the slice related to the slice header.

[0015] The image decoding method according to the present disclosure may further include: when the second flag specifies that the picture header information exists in the slice header, acquiring the picture header information from the slice header.

[0016] In the image decoding method according to the present disclosure, the second flag may have the same value with respect to all slices in a coding layer video sequence (CLVS).

[0017] In the image decoding method according to the present disclosure, when the second flag specifies that picture header information exists in the slice header, a network abstraction layer (NAL) unit for transmitting the picture header information may not exist in the coding layer video sequence (CLVS).

[0018] In the image decoding method according to the present disclosure, when the second flag specifies that picture header information does not exist in the slice header, picture header information can be acquired from a Network Abstraction Layer (NAL) unit whose NAL unit type is equal to PH_NUT.

[0019] In the image decoding method according to the present disclosure, the first flag may be signaled at a higher level of a slice, and the second flag may be included in a slice header and signaled.

[0020] According to another aspect of the present disclosure, an image decoding device may include a memory and at least one processor. The at least one processor may be configured to: obtain a first flag indicating whether information about a sub-picture exists in a bitstream; obtain a second flag indicating whether picture header information exists in a slice header; and decode the bitstream based on the first flag and the second flag. When the first flag indicates that information about a sub-picture exists in the bitstream, the second flag may have a value indicating that picture header information does not exist in the slice header.

[0021] According to another aspect of the present disclosure, an image encoding method may include encoding a first flag that specifies whether information about a sub-picture exists in a bitstream; encoding a second flag that specifies whether picture header information exists in a slice header; and encoding the bitstream based on the first flag and the second flag. When the first flag specifies that information about the sub-picture exists in the bitstream, the second flag may have a value that specifies that picture header information does not exist in the slice header.

[0022] In the image encoding method according to the present disclosure, when the first flag specifies that information about a sub-picture exists in the bitstream, the slice header may include an identifier of the sub-picture including the slice related to the slice header.

[0023] The image encoding method according to the present disclosure may further include: when the second flag specifies that the picture header information exists in the slice header, encoding the picture header information in the slice header.

[0024] In the image encoding method according to the present disclosure, the second flag may have the same value with respect to all slices in a coding layer video sequence (CLVS).

[0025] In the image encoding method according to the present disclosure, when the second flag specifies that picture header information does not exist in the slice header, the picture header information may be signaled through a Network Abstraction Layer (NAL) unit having a NAL unit type equal to PH_NUT.

[0026] In the image encoding method according to the present disclosure, the first flag may be signaled at a higher level of a slice, and the second flag may be included in a slice header and signaled.

[0027] In addition, a transmission method according to another aspect of the present disclosure may transmit a bit stream generated by the image encoding apparatus or the image encoding method of the present disclosure.

[0028] In addition, a computer-readable recording medium according to another aspect of the present disclosure may store a bit stream generated by the image encoding apparatus or the image encoding method of the present disclosure.

[0029] The features of the above brief summary of the present disclosure are merely exemplary aspects of the following detailed description of the present disclosure and do not limit the scope of the present disclosure.

[0030] Beneficial effects

[0031] According to the present disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.

[0032] In addition, according to the present disclosure, an image encoding / decoding method and apparatus that improve encoding / decoding efficiency by efficiently signaling information about a sub-picture and a picture header may be provided.

[0033] In addition, according to the present disclosure, a method of transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure may be provided.

[0034] In addition, according to the present disclosure, a recording medium storing a bit stream generated by the image encoding method or apparatus according to the present disclosure may be provided.

[0035] In addition, according to the present disclosure, there may be provided a recording medium storing a bit stream received and decoded by the image decoding apparatus according to the present disclosure and used to reconstruct an image.

[0036] Those skilled in the art will understand that the effects that can be achieved through the present disclosure are not limited to the contents that have been specifically described above, and other advantages of the present disclosure will be more clearly understood from the detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 FIG. 1 is a diagram schematically illustrating a video encoding system to which an embodiment of the present disclosure is applicable.

[0038] Figure 2 is a diagram schematically showing an image encoding device to which an embodiment of the present disclosure is applicable.

[0039] Figure 3 FIG. 1 is a diagram schematically showing an image decoding device to which an embodiment of the present disclosure is applicable.

[0040] Figure 4 is a diagram illustrating a segmentation structure of an image according to an embodiment.

[0041] Figure 5 is a view showing an embodiment of a partition type of a block according to a multi-type tree structure.

[0042] Figure 6 is a diagram illustrating a signaling mechanism of block partitioning information in a quadtree having a nested multi-type tree structure according to the present disclosure.

[0043] Figure 7 is a diagram showing an example of partitioning a CTU into a plurality of CUs.

[0044] Figure 8 3 is a flowchart illustrating a method for encoding an image using slices / tiles by an image encoding apparatus according to an embodiment of the present disclosure.

[0045] Figure 9 3 is a flowchart illustrating a method for decoding an image using slices / tiles by an image decoding apparatus according to an embodiment of the present disclosure.

[0046] Figure 10 is a diagram of an example of the present disclosure showing signaling and syntax elements in a picture header.

[0047] Figure 11 is a diagram illustrating a syntax structure of a slice header according to an embodiment of the present disclosure.

[0048] Figure 12 This example illustrates parsing and decoding Figure 11 Flowchart of the slicing header method.

[0049] Figure 13 is an example of Figure 11 Flowchart of a method for encoding a slice header.

[0050] Figure 14 is a diagram illustrating a syntax structure of a slice header according to another embodiment of the present disclosure.

[0051] Figure 15 This example illustrates parsing and decoding Figure 14 Flowchart of the slicing header method.

[0052] Figure 16 is an example of Figure 14 Flowchart of a method for encoding a slice header.

[0053] Figure 17 is a diagram illustrating a syntax structure of a slice header according to another embodiment of the present disclosure.

[0054] Figure 18 This example illustrates parsing and decoding Figure 17 Flowchart of the slicing header method.

[0055] Figure 19 is an example of Figure 17 Flowchart of a method for encoding a slice header.

[0056] Figure 20 is a diagram illustrating a content streaming system to which an embodiment of the present disclosure is applicable. DETAILED DESCRIPTION

[0057] Hereinafter, the embodiments of the present disclosure will be described in detail with reference to the accompanying drawings to facilitate implementation by those skilled in the art. However, the present disclosure can be implemented in various forms and is not limited to the embodiments described herein.

[0058] When describing the present disclosure, if it is determined that the detailed description of related known functions or configurations makes the scope of the present disclosure unnecessarily ambiguous, its detailed description will be omitted. In the drawings, parts not related to the description of the present disclosure are omitted, and like reference numerals are given to like parts.

[0059] In this disclosure, when a component is “connected,” “coupled,” or “linked” to another component, it may include not only a direct connection relationship but also an indirect connection relationship with intermediate components. In addition, when a component “includes” or “has” other components, unless otherwise specified, it means that other components may also be included, rather than excluding other components.

[0060] In the present disclosure, the terms first, second, etc. are used only to distinguish one component from other components and do not limit the order or importance of the components unless otherwise specified. Accordingly, within the scope of the present disclosure, the first component in one embodiment may be referred to as the second component in another embodiment, and similarly, the second component in one embodiment may be referred to as the first component in another embodiment.

[0061] In this disclosure, components that are distinguished from each other are intended to clearly describe each feature and do not necessarily mean that the components must be separated. That is, multiple components can be integrated and implemented in a single hardware or software unit, or a component can be distributed and implemented in multiple hardware or software units. Therefore, even if not specifically stated, implementations in which these components are integrated or distributed are also included in the scope of this disclosure.

[0062] In the present disclosure, the components described in the various embodiments are not necessarily essential components, and some components may be optional components. Therefore, embodiments consisting of a subset of the components described in the embodiments are also included in the scope of the present disclosure. In addition, embodiments that include other components in addition to the components described in the various embodiments are included in the scope of the present disclosure.

[0063] The present disclosure relates to encoding and decoding of images. Unless otherwise defined in the present disclosure, terms used in the present disclosure may have general meanings commonly used in the technical field to which the present disclosure belongs.

[0064] In this disclosure, a "picture" generally refers to a unit representing an image within a specific time period, while a slice / tile is a coding unit that constitutes a portion of a picture. A picture can be composed of one or more slices / tiles. In addition, a slice / tile can include one or more coding tree units (CTUs).

[0065] In the present disclosure, "pixel" or "picture element (pel)" may refer to the smallest unit constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, or may represent only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component.

[0066] In this disclosure, a "unit" may refer to a basic unit of image processing. The unit may include at least one of a specific area of ​​a picture and information related to the area. In some cases, the unit may be used interchangeably with terms such as "sample array," "block," or "area." In general, an M×N block may include M columns and N rows of samples (or sample arrays) or a set (or array) of transform coefficients.

[0067] In the present disclosure, the term "current block" may refer to one of the following: "current coding block," "current coding unit," "encoding target block," "decoding target block," or "processing target block." When prediction is performed, the term "current block" may refer to either the "current prediction block" or the "prediction target block." When transform (inverse transform) / quantization (dequantization) is performed, the term "current block" may refer to either the "current transform block" or the "transform target block." When filtering is performed, the term "current block" may refer to the "filtering target block."

[0068] In addition, in the present disclosure, unless explicitly stated as a chroma block, "current block" may mean "luminance block of the current block." "Chroma block of the current block" may be expressed by including an explicit description of the chroma block such as "chroma block" or "current chroma block."

[0069] In the present disclosure, the slash " / " or "," may be interpreted as indicating "and / or". For example, "A / B" and "A, B" may mean "A and / or B". In addition, "A / B / C" and "A, B, C" may mean "at least one of A, B, and / or C".

[0070] In the present disclosure, the term "or" should be interpreted to mean "and / or". For example, the expression "A or B" may include 1) only "A", 2) only "B", or 3) both "A and B". In other words, in the present disclosure, "or" should be interpreted to mean "additionally or alternatively".

[0071] Video Coding System Overview

[0072] Figure 1 is a diagram schematically illustrating a video encoding system according to the present disclosure.

[0073] The video encoding system according to the embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may deliver encoded video and / or image information or data to the decoding device 20 via a digital storage medium or a network in the form of a file or stream.

[0074] The encoding device 10 according to the embodiment may include a video source generator 11, an encoding unit 12, and a transmitter 13. The decoding device 20 according to the embodiment may include a receiver 21, a decoding unit 22, and a renderer 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmitter 13 may be included in the encoding unit 12. The receiver 21 may be included in the decoding unit 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.

[0075] The video source generator 11 can obtain video / images by capturing, synthesizing, or generating video / images. The video source generator 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured videos / images, etc. The video / image generation device may include, for example, a computer, a tablet computer, and a smartphone, and may generate videos / images (electronically). For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capture process may be replaced by a process for generating relevant data.

[0076] The encoding unit 12 may encode the input video / image. For compression and coding efficiency, the encoding unit 12 may perform a series of processes such as prediction, transformation, and quantization. The encoding unit 12 may output the encoded data (encoded video / image information) in the form of a bitstream.

[0077] Transmitter 13 can transmit the encoded video / image information or data output in the form of a bitstream to receiver 21 of decoding device 20 via a digital storage medium or network in the form of a file or stream. Digital storage media can include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. Transmitter 13 can include components for generating media files in a predetermined file format and can also include components for transmission via a broadcast / communication network. Receiver 21 can extract / receive the bitstream from the storage medium or network and transmit the bitstream to decoding unit 22.

[0078] The decoding unit 22 may decode a video / image by performing a series of processes corresponding to the operations of the encoding unit 12 , such as dequantization, inverse transformation, and prediction.

[0079] The renderer 23 may render the decoded video / image. The rendered video / image may be displayed on a display.

[0080] Overview of Image Coding Device

[0081] Figure 2FIG. 1 is a diagram schematically showing an image encoding device to which an embodiment of the present disclosure is applicable.

[0082] like Figure 2 As shown, the image encoding device 100 may include an image splitter 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-frame predictor 180, an intra-frame predictor 185, and an entropy encoder 190. The inter-frame predictor 180 and the intra-frame predictor 185 may be collectively referred to as a "predictor." The transformer 120, the quantizer 130, the dequantizer 140, and the inverse transformer 150 may be included in a residual processor. The residual processor may further include a subtractor 115.

[0083] In some embodiments, all or at least some of the components configuring the image encoding apparatus 100 may be configured by one hardware component (eg, an encoder or a processor). In addition, the memory 170 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium.

[0084] The image splitter 110 may split the input image (or picture or frame) input to the image encoding device 100 into one or more processing units. For example, a processing unit may be referred to as a coding unit (CU). A coding unit may be obtained by recursively splitting a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree, binary tree, or ternary tree (QT / BT / TT) structure. For example, a coding unit may be split into multiple coding units of a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the splitting of the coding unit, a quadtree structure may be applied first, and then a binary tree structure and / or a ternary tree structure may be applied. The encoding process according to the present disclosure may be performed based on the final coding unit that is no longer split. The maximum coding unit may be used as the final coding unit, or a coding unit of a deeper depth obtained by splitting the maximum coding unit may be used as the final coding unit. Here, the encoding process may include the prediction, transformation, and reconstruction processes described later. As another example, the processing unit of the encoding process may be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit may be divided or partitioned from the final coding unit. The prediction unit may be a sample prediction unit, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from the transform coefficient.

[0085] The predictor (inter-frame predictor 180 or intra-frame predictor 185) can perform prediction on the block to be processed (current block) and generate a prediction block including prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. The predictor can generate various information related to the prediction of the current block and transmit the generated information to the entropy encoder 190. The information about the prediction can be encoded in the entropy encoder 190 and output in the form of a bitstream.

[0086] The intra-frame predictor 185 can predict the current block by referring to samples in the current picture. Depending on the intra-frame prediction mode and / or intra-frame prediction technology, the reference samples can be located in the neighborhood of the current block or can be placed separately. The intra-frame prediction mode may include multiple non-directional modes and multiple directional modes. The non-directional mode may include, for example, a DC mode and a planar mode. Depending on the level of detail of the prediction direction, the directional mode may include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-frame predictor 185 may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.

[0087] The inter-frame predictor 180 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. Motion information can include a motion vector and a reference picture index. Motion information can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc. A reference picture including temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 180 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate to use to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter-frame predictor 180 can use the motion information of the neighboring block as the motion information of the current block. In the case of skip mode, unlike merge mode, the residual signal may not be transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of the neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be signaled by encoding a motion vector difference and an indicator of the motion vector predictor. The motion vector difference may mean the difference between the motion vector of the current block and the motion vector predictor.

[0088] The predictor can generate a prediction signal based on various prediction methods and prediction techniques described below. For example, the predictor can apply not only intra prediction or inter prediction, but also both intra and inter prediction simultaneously to predict the current block. A prediction method that simultaneously applies both intra and inter prediction to predict the current block is referred to as combined inter and intra prediction (CIIP). Furthermore, the predictor can perform intra block copying (IBC) to predict the current block. Intra block copying can be used for content image / video encoding, such as gaming, such as screen content coding (SCC). IBC is a method that predicts the current picture using a previously reconstructed reference block in the current picture at a predetermined distance from the current block. When IBC is applied, the position of the reference block in the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC essentially performs prediction within the current picture, but can be performed similarly to inter prediction because the reference block is derived within the current picture. That is, IBC can use at least one of the inter prediction techniques described in this disclosure. IBC essentially performs prediction within the current picture, but can be performed similarly to inter prediction because the reference block is derived within the current picture. That is, IBC may use at least one of the inter-frame prediction techniques described in this disclosure.

[0089] The prediction signal generated by the predictor can be used to generate a reconstructed signal or a residual signal. The subtractor 115 can generate a residual signal (residual block or residual sample array) by subtracting the prediction signal (prediction block or prediction sample array) output from the predictor from the input image signal (original block or original sample array). The generated residual signal can be transmitted to the transformer 120.

[0090] The transformer 120 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process may be applied to square pixel blocks of the same size or to blocks of variable size other than square.

[0091] The quantizer 130 may quantize the transform coefficients and transmit them to the entropy encoder 190. The entropy encoder 190 may encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 130 may rearrange the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scanning order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.

[0092] The entropy encoder 190 can perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 190 can encode information required for video / image reconstruction (e.g., values ​​of syntax elements, etc.) in addition to quantized transform coefficients, together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of a network abstraction layer (NAL). The video / image information may also include information about various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include general constraint information. The signaled information, transmitted information, and / or syntax elements described in the present disclosure may be encoded through the above-mentioned encoding process and included in the bitstream.

[0093] The bitstream may be transmitted over a network or stored in a digital storage medium. The network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting a signal output from the entropy encoder 190 and / or a storage unit (not shown) for storing the signal may be included as an internal / external element of the image encoding device 100. Alternatively, the transmitter may be provided as a component of the entropy encoder 190.

[0094] The quantized transform coefficients output from the quantizer 130 may be used to generate a residual signal. For example, the residual signal (residual block or residual sample) may be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients through the dequantizer 140 and the inverse transformer 150.

[0095] The adder 155 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 180 or the intra-frame predictor 185 to generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). If the block to be processed has no residual, such as when skip mode is applied, the prediction block can be used as the reconstructed block. The adder 155 can be called a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, and can be used for inter-frame prediction of the next picture through filtering as described below.

[0096] The filter 160 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 160 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. The filter 160 can generate various information related to filtering and transmit the generated information to the entropy encoder 190, as described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoder 190 and output in the form of a bitstream.

[0097] The modified reconstructed picture transferred to the memory 170 may be used as a reference picture in the inter predictor 180. When inter prediction is applied by the image encoding device 100, prediction mismatch between the image encoding device 100 and the image decoding device may be avoided and encoding efficiency may be improved.

[0098] The DPB of the memory 170 may store the modified reconstructed picture for use as a reference picture in the inter-frame predictor 180. The memory 170 may store motion information of a block from which motion information in the current picture was derived (or encoded) and / or motion information of a reconstructed block in the picture. The stored motion information may be transmitted to the inter-frame predictor 180 and used as motion information of a spatially neighboring block or motion information of a temporally neighboring block. The memory 170 may store reconstructed samples of a reconstructed block in the current picture and may transmit the reconstructed samples to the intra-frame predictor 185.

[0099] Overview of Image Decoding Device

[0100] Figure 3 FIG. 1 is a diagram schematically showing an image decoding device to which an embodiment of the present disclosure is applicable.

[0101] like Figure 3As shown, the image decoding apparatus 200 may include an entropy decoder 210, a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame predictor 260, and an intra-frame predictor 265. The inter-frame predictor 260 and the intra-frame predictor 265 may be collectively referred to as a "predictor." The dequantizer 220 and the inverse transformer 230 may be included in a residual processor.

[0102] According to an embodiment, all or at least some of the components configuring the image decoding apparatus 200 may be configured by hardware components (eg, a decoder or a processor). In addition, the memory 250 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium.

[0103] The image decoding apparatus 200 having received a bit stream including video / image information may decode the image by performing the same operation as that performed by Figure 2 The image may be reconstructed by processing corresponding to the processing performed by the image encoding device 100. For example, the image decoding device 200 may perform decoding using the processing unit applied in the image encoding device. Therefore, the processing unit for decoding may be, for example, a coding unit. The coding unit may be obtained by dividing the coding tree unit or the maximum coding unit. The reconstructed image signal decoded and output by the image decoding device 200 may be reproduced by a reproduction device (not shown).

[0104] The image decoding apparatus 200 may receive the image in the form of a bit stream from Figure 2The received signal can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can parse the bitstream to derive information required for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may also include information about various parameter sets, such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include general constraint information. The image decoding device may also decode the picture based on the parameter set information and / or the general constraint information. The signaled / received information and / or syntax elements described in the present disclosure can be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 210 decodes the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and outputs the values ​​of the syntax elements required for image reconstruction and the quantized values ​​of the transform coefficients of the residual. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, use the decoding target syntax element information, the decoding information of the neighboring blocks and the decoding target block, or the information of the symbol / bin decoded in the previous stage to determine the context model, and perform arithmetic decoding on the bin by predicting the probability of occurrence of the bin according to the determined context model to generate a symbol corresponding to the value of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin. The information related to the prediction in the information decoded by the entropy decoder 210 can be provided to the predictor (inter-frame predictor 260 and intra-frame predictor 265), and the residual value on which entropy decoding is performed in the entropy decoder 210, that is, the quantized transform coefficient and related parameter information can be input to the dequantizer 220. In addition, information about filtering in the information decoded by the entropy decoder 210 can be provided to the filter 240. Meanwhile, a receiver (not shown) for receiving a signal output from the image encoding device may be further configured as an internal / external element of the image decoding device 200 , or the receiver may be a component of the entropy decoder 210 .

[0105] Meanwhile, the image decoding apparatus according to the present disclosure may be referred to as a video / image / picture decoding apparatus. The image decoding apparatus may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 210. The sample decoder may include a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame predictor 260, or at least one of an intra-frame predictor 265.

[0106] The dequantizer 220 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 220 may rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the image encoding device. The dequantizer 220 may dequantize the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.

[0107] The inverse transformer 230 may inversely transform the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0108] The predictor may perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the prediction information output from the entropy decoder 210, and may determine a specific intra / inter prediction mode (prediction technique).

[0109] As described in the predictor of the image encoding device 100 , the predictor can generate a prediction signal based on various prediction methods (techniques) described later.

[0110] The intra predictor 265 may predict the current block by referring to samples in the current picture. The description of the intra predictor 185 is also applicable to the intra predictor 265.

[0111] The inter-frame predictor 260 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. Motion information can include a motion vector and a reference picture index. Motion information can also include information on the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter-frame predictor 260 can configure a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and information about the prediction can include information indicating the inter-frame prediction mode of the current block.

[0112] The adder 235 can generate a reconstructed block by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter-frame predictor 260 and / or the intra-frame predictor 265). If the block to be processed has no residual, such as when skip mode is applied, the prediction block can be used as the reconstructed block. The description of the adder 155 also applies to the adder 235. The adder 235 can be called a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, and can be used for inter-frame prediction of the next picture through filtering as described below.

[0113] The filter 240 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 240 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 250, specifically, in the DPB of the memory 250. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc.

[0114] The (modified) reconstructed picture stored in the DPB of the memory 250 can be used as a reference picture in the inter-frame predictor 260. The memory 250 can store the motion information of the block from which the motion information in the current picture was derived (or decoded) and / or the motion information of the reconstructed block in the picture. The stored motion information can be transmitted to the inter-frame predictor 260 to be used as the motion information of the spatially adjacent block or the motion information of the temporally adjacent block. The memory 250 can store the reconstructed samples of the reconstructed block in the current picture and transmit the reconstructed samples to the intra-frame predictor 265.

[0115] In the present disclosure, the embodiments described in the filter 160 , the inter-frame predictor 180 , and the intra-frame predictor 185 of the image encoding device 100 may be equally or correspondingly applied to the filter 240 , the inter-frame predictor 260 , and the intra-frame predictor 265 of the image decoding device 200 .

[0116] Overview of Image Segmentation

[0117] The video / image encoding method according to the present disclosure can be performed based on the image segmentation structure as follows. Specifically, the prediction, residual processing ((inverse) transform, (de)quantization, etc.), syntax element encoding and filtering processes described later can be performed based on the CTU, CU (and / or TU, PU) derived from the image segmentation structure. The image can be segmented in units of blocks and the block segmentation process can be performed in the image segmentor 110 of the encoding device. The segmentation related information can be encoded by the entropy encoder 190 and sent to the decoding device in the form of a bit stream. The entropy decoder 210 of the decoding device can derive the block segmentation structure of the current picture based on the segmentation related information obtained from the bit stream, and based on this, a series of processes (e.g., prediction, residual processing, block / picture reconstruction, in-loop filtering, etc.) can be performed to perform image decoding. The CU size and the TU size can be the same, or there can be multiple TUs in the CU area. At the same time, the CU size can generally represent the luminance component (sample) CB size. The TU size can generally represent the luminance component (sample) TB size. The chroma component (sample) CB or TB size can be derived based on the luminance component (sample) CB or TB size according to the component ratio according to the chroma format (color format, such as 4:4:4, 4:2:2, 4:2:0, etc.) of the picture / image. The TU size can be derived based on the maxTbSize that specifies the maximum available TB size. For example, when the CU size is larger than the maxTbSize, multiple TUs (TBs) of the maxTbSize can be derived from the CU, and transformation / inverse transformation can be performed in units of TUs (TBs). In addition, for example, when intra prediction is applied, the intra prediction mode / type can be derived in units of CUs (or CBs), and the neighboring reference sample derivation and prediction sample generation process can be performed in units of TUs (or TBs). In this case, one or more TUs (or TBs) can exist in one CU (or CB) area, and in this case, multiple TUs (or TBs) can share the same intra prediction mode / type.

[0118] A picture may be partitioned into a sequence of coding tree units (CTUs). Figure 4 An example of a picture being partitioned into CTUs is shown. A CTU may correspond to a coding tree block (CTB). Alternatively, a CTU may include a coding tree block of luma samples and two corresponding coding tree blocks of chroma samples. For example, for a picture containing three sample arrays, a CTU may include an N×N block of luma samples and two corresponding blocks of chroma samples.

[0119] Overview of CTU Segmentation

[0120] As described above, a coding unit (CTU) or a largest coding unit (LCU) may be obtained by recursively partitioning the coding tree unit (CTU) or the largest coding unit (LCU) according to a quadtree / binarytree / ternarytree (QT / BT / TT) structure. For example, the CTU may be first partitioned into a quadtree structure. Thereafter, the leaf nodes of the quadtree structure may be further partitioned using a multi-type tree structure.

[0121] Splitting according to the quadtree means that the current CU (or CTU) is equally split into four. By splitting according to the quadtree, the current CU can be split into four CUs with the same width and the same height. When the current CU is no longer split into the quadtree structure, the current CU corresponds to the leaf node of the quadtree structure. The CU corresponding to the leaf node of the quadtree structure can no longer be split and can be used as the final coding unit mentioned above. Alternatively, the CU corresponding to the leaf node of the quadtree structure can be further split by a multi-type tree structure.

[0122] Figure 5 1 is a diagram illustrating an embodiment of a partition type of a block according to a multi-type tree structure. The partition according to the multi-type tree structure may include two types of partitions according to a binary tree structure and two types of partitions according to a ternary tree structure.

[0123] The two types of splits according to the binary tree structure may include vertical binary split (SPLIT_BT_VER) and horizontal binary split (SPLIT_BT_HOR). Vertical binary split (SPLIT_BT_VER) means that the current CU is equally split into two in the vertical direction. Figure 5 As shown in FIG, by vertical binary splitting, two CUs with the same height as the current CU and half the width of the current CU can be generated. Horizontal binary splitting (SPLIT_BT_HOR) means that the current CU is equally split into two in the horizontal direction. Figure 5 As shown, through horizontal binary partitioning, two CUs with a height half of the height of the current CU and the same width as the current CU can be generated.

[0124] The two types of splits according to the triad structure may include vertical triad split (SPLIT_TT_VER) and horizontal triad split (SPLIT_TT_HOR). In vertical triad split (SPLIT_TT_VER), the current CU is split in a vertical direction at a ratio of 1:2:1. Figure 5As shown, through vertical trifurcated partitioning, two CUs with the same height as the current CU and a width of 1 / 4 of the current CU's width, and a CU with the same height as the current CU and a width of half the current CU's width can be generated. In horizontal trifurcated partitioning (SPLIT_TT_HOR), the current CU is split horizontally at a ratio of 1:2:1. Figure 5 As shown, through horizontal trifurcated partitioning, two CUs with a height of 1 / 4 of the current CU and the same width as the current CU, and a CU with a height of half the current CU and the same width as the current CU can be generated.

[0125] Figure 6 is a diagram illustrating a signaling mechanism of block partitioning information in a quadtree having a nested multi-type tree structure according to the present disclosure.

[0126] Here, the CTU is regarded as the root node of the quadtree and is first split into a quadtree structure. Information (e.g., qt_split_flag) specifying whether quadtree partitioning is performed for the current CU (CTU or node (QT_node) of the quadtree) is signaled. For example, when qt_split_flag has a first value (e.g., "1"), the current CU can be quadtree split. In addition, when qt_split_flag has a second value (e.g., "0"), the current CU is not quadtree split, but becomes a leaf node (QT_leaf_node) of the quadtree. Each quadtree leaf node can then be further split into a multi-type tree structure. That is, the leaf node of the quadtree can become a node (MTT_node) of a multi-type tree. In the multi-type tree structure, a first flag (e.g., Mtt_split_cu_flag) is signaled to indicate whether the current node is additionally split. If the corresponding node is additionally split (for example, if the first flag is 1), the second flag (for example, Mtt_split_cu_vertical_flag) can be signaled to indicate the split direction. For example, the split direction can be a vertical direction when the second flag is 1, and a horizontal direction when the second flag is 0. Then, a third flag (for example, Mtt_split_cu_binary_flag) can be signaled to indicate whether the split type is a binary split type or a ternary split type. For example, the split type can be a binary split type when the third flag is 1, and a ternary split type when the third flag is 0. The nodes of the multi-type tree obtained by binary splitting or ternary splitting can be further split into a multi-type tree structure. However, the nodes of the multi-type tree may not be split into a quadtree structure. If the first flag is 0, the corresponding node of the multi-type tree is no longer split, but becomes a leaf node (MTT_leaf_node) of the multi-type tree. The CU corresponding to the leaf node of the multi-type tree can be used as the above-mentioned final coding unit.

[0127] Based on mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, a multi-type tree partition mode (MttSplitMode) of a CU may be derived as shown in the following Table 1. In the following description, a multi-type tree partition mode may be referred to as a multi-tree partition type or a partition type.

[0128] [Table 1]

[0129] MttSplitMode mtt_split_cu_vertical_flag mtt_split_cu_binary_flag SPLIT_TT_HOR 0 0 SPLIT_BT_HOR 0 1 SPLIT_TT_VER 1 0 SPLIT_BT_VER 1 1

[0130] Figure 7 is a diagram showing an example of partitioning a CTU into a plurality of CUs by applying a multi-type tree after applying a quadtree. Figure 7 In FIG, the bold block edge 710 represents a quadtree partition, while the remaining edges 720 represent a multi-type tree partition. A CU may correspond to a coding block (CB). In an embodiment, a CU may include a coding block of luma samples and two coding blocks of chroma samples corresponding to the luma samples.

[0131] The size of the chroma component (sample) CB or TB can be derived based on the luma component (sample) CB or TB size based on the component ratio according to the color format of the picture / image (chroma format, such as 4:4:4, 4:2:2, 4:2:0, etc.). In the case of a 4:4:4 color format, the size of the chroma component CB / TB can be set to be equal to the size of the luma component CB / TB. In the case of a 4:2:2 color format, the width of the chroma component CB / TB can be set to half the width of the luma component CB / TB, and the height of the chroma component CB / TB can be set to the height of the luma component CB / TB. In the case of a 4:2:0 color format, the width of the chroma component CB / TB can be set to half the width of the luma component CB / TB, and the height of the chroma component CB / TB can be set to half the height of the luma component CB / TB.

[0132] In an embodiment, when the size of the CTU is based on a luma sample unit of 128, the size of the CU may be from 128x128 to 4x4, which is the same size as the CTU. In an embodiment, in the case of a 4:2:0 color format (or chroma format), the chroma CB size may be from 64x64 to 2x2.

[0133] Meanwhile, in an embodiment, the CU size and the TU size may be the same. Alternatively, there may be multiple TUs in a CU region. The TU size generally indicates the luma component (sample) transform block (TB) size.

[0134] The TU size can be derived based on the maximum allowed TB size maxTbSize as a predetermined value. For example, when the CU size is larger than maxTbSize, multiple TUs (TBs) with maxTbSize can be derived from the CU, and transformation / inverse transformation can be performed in units of TUs (TBs). For example, the maximum allowed luma TB size can be 64x64 and the maximum allowed chroma TB size can be 32x32. If the width or height of the CB split according to the tree structure is larger than the maximum transform width or height, the CB can be automatically (or implicitly) split until the TB size limits in the horizontal and vertical directions are met.

[0135] In addition, for example, when intra prediction is applied, the intra prediction mode / type can be derived in units of CU (or CB), and the neighboring reference sample derivation and prediction sample generation process can be performed in units of TU (or TB). In this case, there can be one or more TUs (or TBs) in a CU (or CB) area, and in this case, multiple TUs or (TBs) can share the same intra prediction mode / type.

[0136] Meanwhile, for a quadtree coding tree scheme with nested multi-type trees, the following parameters may be signaled from the encoding apparatus to the decoding apparatus as SPS syntax elements. For example, at least one of the CTU size as a parameter indicating the size of the root node of the quadtree, MinQTSize as a parameter indicating the minimum allowed quadtree leaf node size, MaxBtSize as a parameter indicating the maximum allowed binary tree root node size, MaxTtSize as a parameter indicating the maximum allowed ternary tree root node size, MaxMttDepth as a parameter indicating the maximum allowed hierarchical depth of multi-type tree partitioning starting from the quadtree leaf node, MinBtSize as a parameter indicating the minimum allowed binary tree leaf node size, or MinTtSize as a parameter indicating the minimum allowed ternary tree leaf node size may be signaled.

[0137] As an embodiment using a 4:2:0 chroma format, the CTU size can be set to 128x128 luma blocks and two 64x64 chroma blocks corresponding to these luma blocks. In this case, MinOTSize can be set to 16x16, MaxBtSize can be set to 128x128, MaxTtSzie can be set to 64x64, MinBtSize and MinTtSize can be set to 4x4, and MaxMttDepth can be set to 4. Quadtree partitioning can be applied to the CTU to generate quadtree leaf nodes. Quadtree leaf nodes can be referred to as leaf QT nodes. The size of the quadtree leaf node can range from 16x16 size (e.g., MinOTSize) to 128x128 size (e.g., CTU size). If the leaf QT node is 128x128, it may not be additionally partitioned into a binary tree / ternary tree. This is because, in this case, even if it is partitioned, it exceeds MaxBtsize and MaxTtszie (e.g., 64x64). In other cases, the leaf QT node can be further split into a multi-type tree. Therefore, the leaf QT node is the root node of the multi-type tree, and the leaf QT node can have a multi-type tree depth (mttDepth) value of 0. If the multi-type tree depth reaches MaxMttdepth (for example, 4), further splitting can be ignored. If the width of the multi-type tree node is equal to MinBtSize and is less than or equal to 2xMinTtSize, further horizontal splitting can be ignored. If the height of the multi-type tree node is equal to MinBtSize and is less than or equal to 2xMinTtSize, further vertical splitting can be ignored. When splitting is not considered, the encoding device can skip the signaling of the splitting information. In this case, the decoding device can derive splitting information with a predetermined value.

[0138] At the same time, one CTU may include a coding block of luma samples (hereinafter referred to as "luminance block") and two coding blocks of chroma samples corresponding thereto (hereinafter referred to as "chroma blocks"). The above-mentioned coding tree scheme may be applied equally or separately to the luma blocks and chroma blocks of the current CU. Specifically, the luma blocks and chroma blocks in one CTU may be partitioned into the same block tree structure, and in this case, the tree structure is represented as SINGLE_TREE. Alternatively, the luma blocks and chroma blocks in one CTU may be partitioned into separate block tree structures, and in this case, the tree structure may be represented as DUAL_TREE. That is, when the CTU is divided into dual trees, the block tree structure for the luma block and the block tree structure for the chroma block may exist separately. In this case, the block tree structure for the luma block may be referred to as DUAL_TREE_LUMA, and the block tree structure for the chroma component may be referred to as DUAL_TREE_CHROMA. For P and B slices / tile groups, the luma blocks and chroma blocks in one CTU may be restricted to have the same coding tree structure. However, for I slices / patch groups, luma blocks and chroma blocks may have separate block tree structures. If a separate block tree structure is applied, luma CTBs may be divided into CUs based on a specific coding tree structure, and chroma CTBs may be divided into chroma CUs based on another coding tree structure. That is, this means that a CU in an I slice / patch group to which a separate block tree structure is applied may include coding blocks for the luma component or coding blocks for two chroma components, and a CU in a P or B slice / patch group may include blocks for three color components (one luma component and two chroma components).

[0139] Although a quadtree coding tree structure with nested multi-type trees has been described, the structure for partitioning the CU is not limited thereto. For example, the BT structure and the TT structure may be interpreted as concepts included in the multi-partition tree (MPT) structure, and the CU may be interpreted as being partitioned by the QT structure and the MPT structure. In the example where the CU is partitioned by the QT structure and the MPT structure, a syntax element (e.g., MPT_split_type) including information on how many blocks a leaf node of the QT structure is partitioned into and a syntax element (e.g., MPT_split_mode) including information on whether a leaf node of the QT structure is partitioned into vertical or horizontal directions may be signaled to determine the partition structure.

[0140] In another example, the CU may be split in a manner different from the QT structure, the BT structure, or the TT structure. That is, instead of splitting a CU of a lower depth into 1 / 4 of a CU of a higher depth according to the QT structure, splitting a CU of a lower depth into 1 / 2 of a CU of a higher depth according to the BT structure, or splitting a CU of a lower depth into 1 / 4 or 1 / 2 of a CU of a higher depth according to the TT structure, in some cases the CU of a lower depth may be split into 1 / 5, 1 / 3, 3 / 8, 3 / 5, 2 / 3, or 5 / 8 of a CU of a higher depth, and the method of splitting the CU is not limited thereto.

[0141] A quadtree coding block structure with multiple tree types can provide a very flexible block partitioning structure. Due to the supported partitioning types in the multi-type tree, different partitioning patterns can potentially produce the same coding block structure in some cases. By limiting the occurrence of such redundant partitioning patterns in encoding and decoding devices, the amount of partitioning information data can be reduced.

[0142] Sprite-based image encoding / decoding

[0143] A coding target picture can be divided into multiple CTUs, slices, tiles, or blocks, and a picture can be divided into multiple sub-pictures.

[0144] Within a picture, a sub-picture can be encoded or decoded regardless of whether the previous sub-picture is encoded or decoded. For example, different quantization or different resolutions can be applied to multiple sub-pictures.

[0145] In addition, each sub-picture can be processed like a separate picture. For example, the encoding target picture can be a projected picture or a packed picture in an omnidirectional image / video or a 360-degree image / video.

[0146] In this embodiment, a portion of a screen may be rendered or displayed based on a viewport of a user terminal (e.g., a head-mounted display). Therefore, to achieve low latency, among the sub-screens configuring a screen, at least one sub-screen that covers the viewport may be encoded or decoded prior to or independently of the remaining sub-screens.

[0147] The encoding result of the sub-picture may be referred to as a sub-bitstream, sub-stream, or simply a bitstream. A decoding device may decode the sub-picture from the sub-bitstream, sub-stream, or bitstream. In this case, a high-level syntax (HLS) such as a PPS, SPS, VPS, and / or a decoding parameter set (DPS) may be used to encode / decode the sub-picture.

[0148] In the present disclosure, the high-level syntax (HLS) may include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, or SH syntax. For example, APS (APS syntax) or PPS (PPS syntax) may include information / parameters that may be commonly applied to one or more slices or pictures. SPS (SPS syntax) may include information / parameters that may be commonly applied to one or more sequences. VPS (VPS syntax) may include information / parameters that may be commonly applied to multiple layers. DPS (DPS syntax) may include information / parameters that may be commonly applied to the entire video. For example, DPS may include information / parameters related to the concatenation of coded video sequences (CVS).

[0149] A sub-picture can configure a rectangular area of ​​a coded picture. The size of a sub-picture can be set differently within the picture. The size and position of a particular individual sub-picture can be set equally for all pictures belonging to a sequence. Separate sub-picture sequences can be decoded independently. Blocks and slices (and CTBs) can be restricted to not cross sub-picture boundaries. To this end, the encoding device can perform encoding so that sub-pictures are decoded independently. To this end, semantic restrictions in the bitstream may be required. In addition, the arrangement of blocks, slices, and tiles in a sub-picture can be configured differently for each picture in a sequence.

[0150] The sub-picture design is intended to be an abstraction or encapsulation with a scope smaller than the picture level but larger than the slice or tile group level. Therefore, the VCL NAL units of a subset of the motion constrained tile set (MCTS) can be extracted from one VVC bitstream, and processing such as rearrangement to another VVC bitstream (e.g., modification at the VCL level) can be performed without difficulty. Here, MCTS is a coding technique that allows spatial and temporal independence between tiles. When MCTS is applied, it is impossible to refer to information about tiles that are not included in the MCTS to which the current tile belongs. When an image is split into MCTS and encoded, independent transmission and encoding of the MCTS can be performed.

[0151] The advantage of this sprite design is that it allows changing viewing orientations with mixed-resolution viewports relying on a 360° streaming scheme.

[0152] In the following, reference will be made to Figure 8 and Figure 9 Describes image encoding / decoding methods using slices / tiles.

[0153] Figure 8 3 is a flowchart illustrating a method for encoding an image using slices / tiles by an image encoding apparatus according to an embodiment of the present disclosure.

[0154] The image encoding apparatus may derive slices / tiles in the current picture by dividing the current picture ( S810 ).

[0155] The image encoding apparatus may encode the current picture based on the slices / tiles derived in step S810 ( S820 ).

[0156] Figure 9 3 is a flowchart illustrating a method for decoding an image using slices / tiles by an image decoding apparatus according to an embodiment of the present disclosure.

[0157] The image decoding apparatus may acquire information about a video / image from a bitstream (S910).

[0158] In addition, the image decoding apparatus may derive a slice / patch in the current picture based on the information about the video / image acquired in step S910 (S920). Here, the information about the video / image may include information about the slice / patch.

[0159] Next, the image decoding apparatus may decode the current picture based on the slice / tile derived in step S920 ( S930 ).

[0160] exist Figure 8 and Figure 9 In the present disclosure, the information about slices / tiles may include various information and / or syntax elements described in the present disclosure. The video / image information may include high-level syntax, and the high-level syntax may include information about slices and / or information about tiles. The high-level syntax may include a picture header, and the information about the picture header may be included in the slice header described in the present disclosure. The information about slices may include information specifying one or more slices, and the information about tiles may include information specifying one or more tiles. A slice including one or more tiles may exist in a picture.

[0161] High-Level Syntax (HLS) Signaling

[0162] As described above, high-level syntax may be encoded / signaled for video / image coding.Hereinafter, signaling and syntax elements in a picture header and a slice header according to the present disclosure will be described.

[0163] Picture header and slice header

[0164] A coded picture may consist of one or more slices. The parameters of a coded picture are signaled in a picture header (PH) and the parameters of a slice are signaled in a slice header (SH). The PH is carried in its own NAL unit type. The SH may be present at the beginning of a NAL unit containing the payload of a slice (i.e., slice data). In the following, reference will be made to Figure 10 and Figure 11 Describes the syntax elements of PH and SH and the semantics of the syntax elements.

[0165] Figure 10 is a diagram of an example of the present disclosure showing signaling and syntax elements in a picture header.

[0166] The picture_header_rbsp() contains information common to all slices of the coded picture associated with the picture header (PH). For example, the picture_header_rbsp() may include a reference picture flag (non_reference_picture_flag), GDR picture identification information (gdr_pic_flag), no_output_of_prior_pics_flag, recovery_poc_cnt, ph_pic_parameter_set_id, etc. Here, when gdr_pic_flag is 1, recovery_poc_cnt is signaled in the picture_header_rbsp().

[0167] The first value of non_reference_picture_flag (eg, 1) specifies that the picture associated with the PH is not used as a reference picture. The second value of non_reference_picture_flag (eg, 0) specifies that the picture associated with the PH may or may not be used as a reference picture.

[0168] A first value of gdr_pic_flag (eg, 1) specifies that the picture associated with the PH is a GDR picture. A second value of gdr_pic_flag (eg, 0) specifies that the picture associated with the PH is not a GDR picture.

[0169] no_output_of_prior_pics_flag affects the output of previously decoded pictures in the DPB after decoding of a coded layer video sequence (CLVSS) picture that is not the first picture in the bitstream.

[0170] recovery_poc_cnt specifies the recovery point of the decoded picture in output order.

[0171] ph_pic_parameter_set_id specifies the value of pps_pic_parameter_set_id for the picture parameter set (PPS) in use. pps_pic_parameter_set_id is a value used to identify the PPS to be referenced by another syntax.

[0172] Included in Figure 10The syntax elements in the picture_header_rbsp() syntax structure may be included in the picture_header_structure() syntax structure and signaled. In this case, the picture_header_structure() syntax structure may be included in the picture_header_rbsp() syntax structure and signaled.

[0173] Figure 11 is a diagram illustrating a syntax structure of a slice header according to an embodiment of the present disclosure.

[0174] like Figure 11 As shown, picture_header_in_slice_header_flag, picture_header_structure(), slice_subpic_id, slice_address, num_tiles_in_slice_minus1, etc. can be notified through the slice header signal.

[0175] exist Figure 11 In the example shown, picture_header_in_slice_header_flag specifies whether a picture header syntax structure is present in a slice header syntax structure. A first value of picture_header_in_slice_header_flag (e.g., 1 or true) specifies that a picture header is present in a slice header, and a second value of picture_header_in_slice_header_flag (e.g., 0 or false) specifies that a picture header is not present in a slice header.

[0176] picture_header_structure() may be acquired based on picture_header_in_slice_header_flag. For example, when picture_header_in_slice_header_flag has a first value, picture_header_structure() may be signaled. When picture_header_in_slice_header_flag has a second value, picture_header_structure() may not be included in the slice header, but may be included in a separate NAL unit and signaled.

[0177] slice_subpic_id may be information about a sub-picture identifier for identifying a sub-picture that includes the current slice. slice_subpic_id may be obtained based on subpics_present_flag. For example, when subpics_present_flag is 1, slice_subpic_id may be signaled. subpics_present_flag may specify whether a sub-picture exists in the current picture or whether information about a sub-picture exists in the bitstream. For example, a first value of subpics_present_flag (e.g., 1 or true) may specify that information about a sub-picture exists in the bitstream or that one or more sub-pictures exist in the current picture. A second value of subpics_present_flag (e.g., 0 or false) may specify that information about a sub-picture does not exist in the bitstream or that a sub-picture does not exist in the current picture.

[0178] Slice_address may specify the address of the current slice in the current picture. Slice_address may be obtained based on rect_slice_flag and / or NumTilesInPic. For example, when rect_slice_flag is a first value (e.g., 1 or true) or NumTilesInPic is greater than 1, slice_address may be signaled in the slice header. At this time, rect_slice_flag may indicate an indicator of whether the slice included in the current picture is a rectangular slice. For example, rect_slice_flag may be signaled at the picture level (PPS or picture header). In addition, NumTilesInPic may specify the number of tiles included in the current picture.

[0179] num_tiles_in_slice_minus1 may specify the number of tiles included in the current slice. num_tiles_in_slice_minus1 may be obtained based on rect_slice_flag and NumTilesInPic. For example, when rect_slice_flag is a second value (e.g., 0 or false) and NumTilesInPic is greater than 1, num_tiles_in_slice_minus1 may be signaled in the slice header.

[0180] exist Figure 11 In the illustrated embodiment, the bitstream consistency requirements associated with picture_header_in_slice_header_flag may include the following.

[0181] To meet bitstream conformance, the value of picture_header_in_slice_header_flag is required to be the same in all slices of CLVS.

[0182] In addition, when picture_header_in_slice_header_flag is the first value (eg, 1), in order to meet bitstream consistency, it is required that no NAL unit with the NAL unit type equal to PH_NUT exists in the CLVS.

[0183] In addition, when picture_header_in_slice_header_flag is the second value (eg, 0), in order to satisfy bitstream consistency, a NAL unit with a NAL unit type equal to PH_NUT is required to exist before the first VCL NAL unit of the PU in the PU.

[0184] Figure 12 This example illustrates parsing and decoding Figure 11 Flowchart of the slicing header method.

[0185] First, the image decoding apparatus may acquire a first flag (picture_header_in_slice_header_flag) included in a slice header ( S1210 ).

[0186] The first flag may specify whether a picture header exists in a slice header. In addition, the first flag may specify whether the current picture includes only one slice.

[0187] When the first flag is a first value (e.g., 1 or true) (step S1220-yes), the image decoding apparatus may obtain the picture header from the slice header (S1230). When the first flag is a second value (e.g., 0 or false) (step S1220-no), the picture header may be obtained from the picture header NAL unit instead of the slice header (not shown).

[0188] Thereafter, it may be determined in step S1240 whether subpics_present_flag is a first value (e.g., 1 or true). subpics_present_flag may specify whether the current picture includes a subpicture. Additionally, subpics_present_flag may specify whether information about a subpicture is included in the bitstream. subpics_present_flag may be signaled at a higher level of the slice. For example, subpics_present_flag may be included and signaled in a sequence parameter set.

[0189] When subpics_present_flag is a first value (e.g., 1 or true) (step S1240-yes), the image decoding apparatus may obtain slice_subpic_id from the slice header (S1250). When subpics_present_flag is a second value (e.g., 0 or false) (step S1240-no), the image decoding apparatus may omit (skip) parsing slice_subpic_id from the slice header.

[0190] Thereafter, in step S1260, it may be determined whether rect_slice_flag is a first value (e.g., 1 or true) and / or whether NumTilesInPic is greater than 1. rect_slice_flag may be an indicator indicating whether a slice included in the current picture is a rectangular slice. For example, rect_slice_flag may be signaled at the picture level (PPS or picture header). In addition, NumTilesInPic may specify the number of tiles included in the current picture.

[0191] When rect_slice_flag is a first value (e.g., 1 or true) or NumTilesInPic is greater than 1 (step S1260-yes), the image decoding apparatus may obtain slice_address from the slice header (S1270). When rect_slice_flag is a second value (e.g., 0 or false) and NumTilesInPic is not greater than 1 (step S1260-no), the image decoding apparatus may omit (skip) parsing slice_address from the slice header.

[0192] Thereafter, in step S1280 , it may be determined whether rect_slice_flag is a first value (eg, 1 or true) and / or whether NumTilesInPic is greater than 1.

[0193] When rect_slice_flag is a first value (e.g., 1 or true) or NumTilesInPic is not greater than 1 (step S1280-No), the image decoding device may omit (skip) parsing num_tiles_in_slice_minus1 from the slice header. When rect_slice_flag is a second value (e.g., 0 or false) and NumTilesInPic is greater than 1 (step S1280-Yes), the image decoding device may obtain num_tiles_in_slice_minus1 from the slice header (S1290).

[0194] Thereafter, the image decoding apparatus may decode the slice header by parsing subsequent syntax elements (not shown) from the slice header.

[0195] Figure 13 is an example of Figure 11 Flowchart of a method for encoding a slice header.

[0196] First, the image encoding apparatus may determine a value of a first flag (picture_header_in_slice_header_flag) and encode the first flag in a slice header (S1310).

[0197] When the first flag is a first value (e.g., 1 or true) (step S1320-yes), the image encoding apparatus may encode the picture header in the slice header (S1330). When the first flag is a second value (e.g., 0 or false) (step S1320-no), the picture header is not encoded in the slice header, but may be included in a picture header NAL unit (not shown) and signaled.

[0198] Thereafter, in step S1340, it may be determined whether subpics_present_flag is a first value (eg, 1 or true). The subpics_present_flag may be determined and signaled at a higher level of the slice. For example, the subpics_present_flag may be included in a sequence parameter set and signaled.

[0199] When subpics_present_flag is a first value (e.g., 1 or true) (step S1340-yes), the image encoding apparatus may encode slice_subpic_id in the slice header (S1350). When subpics_present_flag is a second value (e.g., 0 or false) (step S1340-no), the image encoding apparatus may omit (skip) encoding slice_subpic_id in the slice header.

[0200] Thereafter, in step S1360 , it may be determined whether rect_slice_flag is a first value (eg, 1 or true) and / or whether NumTilesInPic is greater than 1.

[0201] When rect_slice_flag is a first value (e.g., 1 or true) or when NumTilesInPic is greater than 1 (step S1360-yes), the image encoding device may encode slice_address in the slice header (S1370). When rect_slice_flag is a second value (e.g., 0 or false) and NumTilesInPic is not greater than 1 (step S1360-no), the image encoding device may omit (skip) encoding slice_address in the slice header.

[0202] Thereafter, in step S1380 , it may be determined whether rect_slice_flag is a first value (eg, 1 or true) and / or whether NumTilesInPic is greater than 1.

[0203] When rect_slice_flag is a first value (e.g., 1 or true) or when NumTilesInPic is not greater than 1 (step S1380-No), the image encoding device may omit (skip) encoding num_tiles_in_slice_minus1 in the slice header. When rect_slice_flag is a second value (e.g., 0 or false) and NumTilesInPic is greater than 1 (step S1380-Yes), the image encoding device may encode num_tiles_in_slice_minus1 in the slice header (S1390).

[0204] Thereafter, the image encoding apparatus may encode the slice header by encoding subsequent syntax elements (not shown) in the slice header.

[0205] In reference Figure 12 and Figure 13 In the described examples, some steps may be changed or omitted. For example, the conditions related to encoding / decoding of slice_address and / or num_tiles_in_slice_minus1 may be changed.

[0206] Hereinafter, the improved reference method considering sub-picture-based image encoding / decoding will be described. Figures 11 to 13 The method of the embodiment is described.

[0207] The image encoding apparatus may encode the current picture based on the sub-picture. Alternatively, the image encoding apparatus may encode at least one sub-picture configuring the current picture and generate a bitstream including encoding information of the encoded at least one sub-picture.

[0208] The image decoding apparatus may decode at least one sub picture included in the current picture based on a bit stream including encoding information of the at least one sub picture.

[0209] As described above, picture_header_in_slice_header_flag can specify whether a picture header is present in a slice header. Additionally, picture_header_in_slice_header_flag can be used to specify whether the current picture includes only one slice or multiple slices. When the current picture includes only one slice, some syntax elements in the slice header have fixed values ​​because the slice is the only slice in the current picture. In this case, it may be efficient not to signal some syntax elements with fixed values.

[0210] Hereinafter, various configurations of the present disclosure for performing efficient signaling will be described. The following configurations can be applied to embodiments of the present disclosure alone or in combination.

[0211] Configuration 1

[0212] When the current picture includes only one slice, the signaling of some syntax elements in the slice header may be skipped (omitted). The values ​​of the syntax elements whose signaling is omitted may be derived or inferred by the image encoding device and / or the image decoding device.

[0213] Whether the current picture includes only one slice may be indicated by a predetermined indicator. Therefore, when the indicator indicates that the current picture includes only one slice, some syntax elements may not be included in the slice header, and their values ​​may be inferred or derived. In this case, the indicator may be used as a condition indicating whether some syntax elements are included in the slice header.

[0214] Configuration 2

[0215] For example, the indicator described in configuration 1 may be picture_header_in_slice_header_flag.

[0216] Configuration 3

[0217] The syntax elements in the slice header whose signaling can be omitted according to the value of picture_header_in_slice_header_flag may include at least one of the following (a) or (b).

[0218] Syntax elements specifying a sub-picture comprising a slice

[0219] The reason why the signaling of the syntax element (a) can be omitted is because when each picture includes only one slice, it is obvious that the sub-picture is not specified. For example, when picture_header_in_slice_header_flag indicates that the current picture includes only one slice, since the current picture is not encoded / decoded based on the sub-picture, the signaling of the information about the sub-picture can be omitted.

[0220] Syntax element for specifying the address of a slice

[0221] The reason why the signaling of syntax element (b) can be omitted is because it is obvious that the slice is the only first slice in the picture. For example, when picture_header_in_slice_header_fla indicates that the current picture includes only one slice, since the current slice is the only slice in the current picture, the signaling of the address of the current slice can be omitted.

[0222] Configuration 4

[0223] When each picture in the sequence has only one slice, sub-pictures are not used. For example, subpics_present_flag or subpic_info_present_flag, which are syntax elements for sub-pictures, can be restricted to a second value (e.g., 0 or false). subpics_present_flag or subpic_info_present_flag can specify whether a sub-picture exists in the current picture or whether information about a sub-picture exists in the bitstream. For example, subpics_present_flag or subpic_info_present_flag can be included in a sequence parameter set and signaled.

[0224] Similarly, when subpics_present_flag or subpic_info_present_flag is a first value (e.g., 1 or true), a flag indicating whether each picture in the sequence includes only one slice or a flag indicating whether a picture header is present in a slice header (e.g., picture_header_in_slice_header_flag) may not indicate that the current picture includes only one slice and may not indicate that a picture header is present in the slice header. Therefore, for example, when subpics_present_flag or subpic_info_present_flag is a first value (e.g., 1 or true), picture_header_in_slice_header_flag may be constrained to have a second value (e.g., 0 or false).

[0225] Configuration 5

[0226] When the picture header is not present in the picture header NAL unit but is present in the slice header, in the CLVS of a specific layer (layer A), the picture headers of all layers that reference layer A (i.e., dependent layers of layer A) and all layers referenced by layer A can be constrained to be present in the slice header instead of the picture header NAL unit. The above constraint is imposed to simplify the detection of picture boundaries within an access unit in the case of a multi-layer bitstream.

[0227] Figure 14 is a diagram illustrating a syntax structure of a slice header according to another embodiment of the present disclosure.

[0228] Because according to Figure 14 Slice header structure and according to the embodiment of Figure 11 The descriptions of the same syntax elements and the same signaling conditions in the slice header structure of the implementation scheme are the same, so the repeated descriptions will be omitted.

[0229] according to Figure 14 In an embodiment of the present invention, the condition for signaling slice_subpic_id may be changed. Specifically, the slice header may include slice_subpic_id based on subpics_present_flag and picture_header_in_slice_header_flag. For example, when subpics_present_flag is a first value (e.g., 1 or true) and picture_header_in_slice_header_flag is a second value (e.g., 0 or false), slice_subpic_id may be signaled in the slice header. This is because, as described above, when picture_header_in_slice_header_flag has a first value, the current picture includes only one slice and sub-picture-based encoding / decoding is not performed, and signaling of information about the sub-picture is unnecessary.

[0230] In addition, according to Figure 14In an embodiment of the present invention, the conditions for signaling slice_address may be changed. Specifically, the slice header may include slice_address based on rect_slice_flag, NumTilesInPic, and picture_header_in_slice_header_flag. For example, when rect_slice_flag is a first value (e.g., 1 or true) or NumTilesInPic is greater than 1, and picture_header_in_slice_header_flag is a second value (e.g., 0 or false), slice_address may be signaled in the slice header. This is because, as described above, when picture_header_in_slice_header_flag has a first value, since the current picture includes only one slice, signaling of information about the address of the slice is unnecessary.

[0231] exist Figure 14 In an embodiment of the present invention, the bitstream conformance requirement of picture_header_in_slice_header_flag can be improved as follows.

[0232] First, the value of picture_header_in_slice_header_flag is required to be the same in all slices in CLVS.

[0233] In addition, when picture_header_in_slice_header_flag is the first value (e.g., 1), NAL units with NAL unit type equal to PH_NUT are required not to be present in CLVS. This is because the picture header is included in the slice header and signaled, so a separate NAL unit for sending the picture header is not required.

[0234] In addition, when picture_header_in_slice_header_flag is the second value (e.g., 0), a NAL unit with the NAL unit type equal to PH_NUT is required to be present in the PU, preceding the first VCL NAL unit of the PU. That is, the current PU is required to have a PH NAL unit. This is because a separate NAL unit is required to send the picture header.

[0235] In addition, when subpics_present_flag or subpic_info_present_flag is the first value (eg, 1), picture_header_in_slice_header_flag is required not to be the first value (eg, 1). In this case, picture_header_in_slice_header_flag may be constrained to have the second value (eg, 0).

[0236] exist Figure 14 In the example of , slice_subpic_id indicates an identifier of a sub-picture including a slice. When slice_subpic_id exists, the variable SubPicIdx is derived such that SubpicIdList[SubPicIdx] is equal to slice_subpic_id. When slice_subpic_id does not exist, the variable SubPicIdx may be derived to be equal to 0.

[0237] exist Figure 14 In the example of , the length (bit length) of slice_subpic_id can be derived as follows.

[0238] If sps_subpic_id_signalling_present_flag is equal to 1, the length of slice_subpic_id is derived to be equal to sps_subpic_id_len_minus1 + 1. Here, sps_subpic_id_signalling_present_flag can specify whether to signal the identifier of the subpicture in the sequence parameter set. sps_subpic_id_len_minus1 is the length information of the subpicture identifier and can be included in the sequence parameter set and signaled.

[0239] Otherwise (if sps_subpic_id_signalling_present_flag is not 1), if ph_subpic_id_signalling_present_flag is 1, the length of slice_subpic_id can be derived to be equal to ph_subpic_id_len_minus1+1. Here, ph_subpic_id_signalling_present_flag can specify whether the identifier of the sub-picture is signaled in the picture header. ph_subpic_id_len_minus1 is the length information of the sub-picture identifier and can be included in the picture header and signaled.

[0240] Otherwise (if both sps_subpic_id_signalling_present_flag and ph_subpic_id_signalling_present_flag are not 1), if pps_subpic_id_signalling_present_flag is 1, the length of slice_subpic_id can be derived to be equal to pps_subpic_id_len_minus1+1. Here, pps_subpic_id_signalling_present_flag can specify whether the identifier of the sub-picture is signaled in the picture parameter set. pps_subpic_id_len_minus1 is the length information of the sub-picture identifier and can be included in the picture parameter set and signaled.

[0241] Otherwise (if all sps_subpic_id_signalling_present_flag, ph_subpic_id_signalling_present_flag, and pps_subpic_id_signalling_present_flag are not 1), the length of slice_subpic_id may be derived to be equal to Ceil(Log2(Sps_num_subpics_minus1+1)). sps_num_subpics_minus1 is the number of subpictures of each picture in the CLVS and may be included and signaled in the sequence parameter set.

[0242] slice_address specifies the slice address of the current slice. When slice_address does not exist, the value of slice_address is inferred to be equal to 0.

[0243] picture_header_structure() may include references Figure 10 At least one syntax element included in the described picture_header_rbsp().

[0244] Figure 15 This example illustrates parsing and decoding Figure 14 Flowchart of the slicing header method.

[0245] Figure 15 Steps S1510 to S1530 are respectively equal to Figure 12 Therefore, repeated descriptions thereof will be omitted.

[0246] Figure 15Steps S1540 to S1570 may correspond to Figure 12 Therefore, repeated description of the common parts will be omitted.

[0247] according to Figure 15 In an embodiment, in step S1540, it can be determined whether subpics_present_flag is a first value (e.g., 1 or true) and whether the first flag is a second value (e.g., 0 or false).

[0248] When subpics_present_flag is the first value and the first flag is the second value (step S1540-Yes), the image decoding apparatus may obtain slice_subpic_id from the slice header (S1550). When subpics_present_flag is the second value (e.g., 0 or false) or the first flag is the first value (e.g., 1 or true) (step S1540-No), the image decoding apparatus may omit (skip) parsing slice_subpic_id from the slice header.

[0249] Thereafter, in step S1560 , it may be determined whether rect_slice_flag is a first value (eg, 1 or true) or whether NumTilesInPic is greater than 1 and the first flag is a second value (eg, 0 or false).

[0250] When rect_slice_flag is the first value or NumTilesInPic is greater than 1 and the first flag is the second value (step S1560-yes), the image decoding device may obtain slice_address from the slice header (S1570). When rect_slice_flag is the second value (for example, 0 or false) and NumTilesInPic is not greater than 1 or the first flag is the first value (for example, 1 or true) (step S1560-no), the image decoding device may omit (skip) parsing slice_address from the slice header.

[0251] Figure 15 Steps S1580 to S1590 are respectively equal to Figure 12 Steps S1280 to S1290 are described above, and thus repeated descriptions thereof will be omitted.

[0252] As reference Figure 12 As described above, the image decoding apparatus may decode the slice header by parsing subsequent syntax elements (not shown) from the slice header.

[0253] Figure 16 is an example of Figure 14 Flowchart of a method for encoding a slice header.

[0254] Figure 16 Steps S1610 to S1630 are respectively equal to Figure 13 The steps S1310 to S1330 are described in detail, and thus repeated descriptions thereof will be omitted.

[0255] Figure 16 Steps S1640 to S1670 may correspond to Figure 16 Therefore, repeated description of the common parts will be omitted.

[0256] according to Figure 16 In an embodiment, in step S1640, it can be determined whether subpics_present_flag is a first value (e.g., 1 or true) and whether the first flag is a second value (e.g., 0 or false).

[0257] When subpics_present_flag is the first value and the first flag is the second value (step S1640-Yes), the image encoding apparatus may encode slice_subpic_id in the slice header (S1650). When subpics_present_flag is the second value (e.g., 0 or false) or the first flag is the first value (e.g., 1 or true) (step S1640-No), the image encoding apparatus may omit (skip) encoding slice_subpic_id in the slice header.

[0258] Thereafter, in step S1660 , it may be determined whether rect_slice_flag is a first value (eg, 1 or true) or whether NumTilesInPic is greater than 1 and the first flag is a second value (eg, 0 or false).

[0259] When rect_slice_flag is the first value or NumTilesInPic is greater than 1 and the first flag is the second value (step S1660-Yes), the image encoding device may encode slice_address in the slice header (S1670). When rect_slice_flag is the second value (for example, 0 or false) and NumTilesInPic is not greater than 1 or the first flag is the first value (for example, 1 or true) (step S1660-No), the image encoding device may omit (skip) encoding slice_address in the slice header.

[0260] Figure 16 Steps S1680 to S1690 are respectively equal to Figure 13 Steps S1380 to S1390 are described above, and thus repeated descriptions thereof will be omitted.

[0261] As reference Figure 13 As described above, the image encoding apparatus may encode the slice header by encoding subsequent syntax elements (not shown) in the slice header.

[0262] In reference Figure 15 and Figure 16 In the described examples, some steps may be changed or omitted. For example, the conditions related to encoding / decoding of slice_address and / or num_tiles_in_slice_minus1 may be changed.

[0263] As a reference Figures 14 to 16 A modified example of the described embodiment, where the improved constraints on picture_header_in_slice_header_flag apply to Figure 11 The embodiment shown. In this case, at least some of the problems of the conventional method can be solved. Specifically, for example, the value of picture_header_in_slice_header_flag can be constrained based on information about the sub-picture (subpics_present_flag or subpic_info_present_flag) signaled at a higher level of the slice header. More specifically, when subpics_present_flag or subpic_info_present_flag is a first value, picture_header_in_slice_header_flag can be constrained to have a second value. Therefore, when subpics_present_flag (or subpic_info_present_flag) is a first value (when information about the sub-picture is present in the bitstream or the current picture includes a sub-picture), picture_header_in_slice_header_flag can indicate that the picture header is not present in the slice header or the current picture includes more than one slice. In Figure 14In the illustrated embodiment, when subpics_present_flag is a first value and picture_header_in_slice_header_flag is a second value, slice_subpic_id can be obtained from the slice header. However, when subpics_present_flag is a first value, since picture_header_in_slice_header_flag is constrained to have a second value, checking subpics_present_flag as a parsing condition for slice_subpic_id may be sufficient. That is, according to this modified example, in steps S1540 and S1640, the determination of whether the first flag has a second value can be omitted. According to this modified example, when subpics_present_flag or subpic_info_present_flag is a first value, the image encoding device may encode picture_header_in_slice_header_flag having a second value. In addition, when subpics_present_flag or subpic_info_present_flag is the first value, the image decoding device can acquire picture_header_in_slice_header_flag having the second value.

[0264] Figure 17 is a diagram illustrating a syntax structure of a slice header according to another embodiment of the present disclosure.

[0265] Because according to Figure 17 Slice header structure and according to the embodiment of Figure 14 The descriptions of the same syntax elements and the same signaling conditions in the slice header structure of the implementation scheme are the same, so the repeated descriptions will be omitted.

[0266] according to Figure 17In an embodiment of the present invention, the condition for signaling picture_header_in_slice_header_flag may be changed. Specifically, the slice header may include picture_header_in_slice_header_flag based on subpics_present_flag. For example, when subpics_present_flag is a first value (e.g., 1 or true), picture_header_in_slice_header_flag may not be signaled in the slice header. For example, when subpics_present_flag is a second value (e.g., 0 or false), picture_header_in_slice_header_flag may be signaled in the slice header. This is because, as described above, when subpics_present_flag has the first value, picture_header_in_slice_header_flag has a fixed value (the second value) because the current picture cannot contain only one slice. Therefore, signaling of picture_header_in_slice_header_flag is unnecessary. In this case, the picture header signaled if picture_header_in_slice_header_flag is the first value may not be signaled through the slice header.

[0267] In addition, according to Figure 17 In an embodiment, when subpics_present_flag is a first value, slice_subpic_id may be signaled in the slice header.

[0268] In the following, for the description of slice_address and num_tiles_in_slice_minus1, refer to Figure 14 .

[0269] exist Figure 17 In the embodiment of the present invention, the bitstream consistency requirements of picture_header_in_slice_header_flag can be consistent with those of reference Figure 14 Same as those described.

[0270] Figure 18 This example illustrates parsing and decoding Figure 17 Flowchart of the slicing header method.

[0271] according to Figure 18 Methods and basis Figure 15The methods differ in some conditions and order of parsing syntax elements, and the descriptions of the commonly disclosed syntax elements may be the same.

[0272] according to Figure 18 In an embodiment, the image decoding apparatus may determine whether the value of subpics_present_flag is a second value (eg, 0 or false) in step S1810.

[0273] When the value of subpics_present_flag is the first value (e.g., 1 or true) in step S1810, the current picture includes more than one slice due to sub-picture-based encoding / decoding. Therefore, in this case, the image decoding apparatus may not obtain picture_header_in_slice_header_flag and the picture header from the slice header, but may obtain slice_subpic_id (S1850).

[0274] When the value of subpics_present_flag is the second value in step S1810, the image decoding device obtains the first flag (picture_header_in_slice_header_flag) from the slice header (S1820). The image decoding device may determine whether the first flag is the first value (S1830), and when the first flag is the first value, obtain the picture header from the slice header (S1840). When the first flag is the second value, the image decoding device does not obtain the picture header from the slice header. In this case, the image decoding device may obtain the picture header through a separate NAL unit. When the value of subpics_present_flag is the second value in step S1810, since sub-picture-based encoding / decoding is not performed, the image decoding device may not obtain information about the sub-picture (slice_subpic_id).

[0275] Figure 18 Steps S1860 to S1890 are respectively equal to Figure 15 Steps S1560 to S1590 are described above, and thus repeated descriptions thereof will be omitted.

[0276] As reference Figure 12 As described above, the image decoding apparatus may decode the slice header line by parsing subsequent syntax elements (not shown) from the slice header.

[0277] Figure 19 is an example of Figure 17 Flowchart of a method for encoding a slice header.

[0278] according to Figure 19 Methods and basis Figure 16The methods differ in some conditions and order of encoding syntax elements, and the descriptions of the commonly disclosed syntax elements may be the same.

[0279] according to Figure 19 In an embodiment, the image encoding apparatus may determine whether the value of subpics_present_flag is a second value (eg, 0 or false) in step S1910.

[0280] When the value of subpics_present_flag is the first value (e.g., 1 or true) in step S1910, the current picture includes more than one slice due to sub-picture-based encoding / decoding. Therefore, in this case, the image encoding apparatus may not encode picture_header_in_slice_header_flag and the picture header in the slice header, but may encode slice_subpic_id (S1950).

[0281] When the value of subpics_present_flag is the second value in step S1910, the image coding apparatus may determine the value of the first flag (picture_header_in_slice_header_flag) and encode the first flag in the slice header (S1920). The image coding apparatus may determine whether the first flag is the first value (S1930), and when the first flag is the first value, encode the picture header in the slice header (S1940). When the first flag is the second value, the image coding apparatus may not encode the picture header in the slice header. In this case, the image coding apparatus may signal the picture header via a separate NAL unit. When the value of subpics_present_flag is the second value in step S1910, since sub-picture-based encoding / decoding is not performed, the image coding apparatus may not encode information about the sub-picture (slice_subpic_id) in the slice header.

[0282] Figure 19 Steps S1960 to S1990 are respectively equal to Figure 16 Steps S1660 to S1690 are described above, and thus repeated descriptions thereof will be omitted.

[0283] As reference Figure 13 As described above, the image encoding apparatus may encode the slice header by encoding subsequent syntax elements (not shown) in the slice header.

[0284] In reference Figure 18 and Figure 19In the example described above, some steps may be changed or omitted. For example, the conditions related to encoding / decoding of slice_address and / or num_tiles_in_slice_minus1 may be changed.

[0285] According to an embodiment of the present disclosure, information on whether a picture header is present in a slice header and / or information on whether a picture includes only one slice may be signaled more efficiently.

[0286] In addition, according to an embodiment of the present disclosure, since information on whether a picture header exists in a slice header is signaled based on whether sub-picture based encoding / decoding is performed, it is possible to prevent unnecessary information from being signaled.

[0287] The names of the syntax elements described in this disclosure may include information about the location of the corresponding syntax element. For example, a syntax element starting with "sps_" may mean that the corresponding syntax element is signaled in a sequence parameter set (SPS). In addition, syntax elements starting with "pps_," "ph_," "sh_," etc. may mean that the corresponding syntax element is signaled in a picture parameter set (PPS), a picture header, and a slice header, respectively.

[0288] Although the exemplary method of the present disclosure is shown as a series of operations for the sake of clarity, it is not intended to limit the order in which the steps are performed, and the steps may be performed simultaneously or in a different order if necessary. To implement the method according to the present disclosure, the steps described may further include other steps, may include the remaining steps except for some steps, or may include other additional steps except for some steps.

[0289] In the present disclosure, an image encoding device or image decoding device that performs a predetermined operation (step) may perform an operation (step) of confirming the execution conditions or circumstances of the corresponding operation (step). For example, if it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or image decoding device may perform the predetermined operation after determining whether the predetermined condition is satisfied.

[0290] The various embodiments of the present disclosure are not a list of all possible combinations and are intended to describe representative aspects of the present disclosure, and matters described in the various embodiments may be applied independently or in combinations of two or more.

[0291] Various embodiments of the present disclosure may be implemented in hardware, firmware, software, or a combination thereof. In the case of implementing the present disclosure in hardware, the present disclosure may be implemented in an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a general purpose processor, a controller, a microcontroller, a microprocessor, or the like.

[0292] In addition, the image decoding device and the image encoding device to which the embodiments of the present disclosure are applied may be included in multimedia broadcast transmission and reception equipment, mobile communication terminals, home theater video equipment, digital theater video equipment, surveillance cameras, video chat equipment, real-time communication equipment such as video communication, mobile streaming equipment, storage media, cameras, video on demand (VoD) service providing equipment, OTT video (over the top video) equipment, Internet streaming service providing equipment, three-dimensional (3D) video equipment, video phone video equipment, medical video equipment, etc., and may be used to process video signals or data signals. For example, OTT video equipment may include game consoles, Blu-ray players, Internet access TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.

[0293] Figure 20 is a diagram illustrating a content streaming system to which an embodiment of the present disclosure can be applied.

[0294] like Figure 20 As shown in , a content streaming system to which the embodiments of the present disclosure are applied may mainly include an encoding server, a streaming server, a network server, a media storage, a user device, and a multimedia input device.

[0295] The encoding server compresses the content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream and sends the bitstream to the streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders directly generate the bitstream, the encoding server can be omitted.

[0296] A bitstream may be generated by applying the image encoding method or the image encoding device according to the embodiment of the present disclosure, and a streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0297] The streaming server transmits multimedia data to user devices based on user requests via a network server, and the network server serves as an intermediary for notifying users of services. When a user requests a desired service from the network server, the network server delivers it to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server serves to control the command and response between devices in the content streaming system.

[0298] The streaming server can receive content from a media storage and / or encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined time.

[0299] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, tablet computers, ultrabooks, wearable devices (e.g., smart watches, smart glasses, head-mounted displays), digital televisions, desktop computers, digital signage, etc.

[0300] Each server in the content streaming system may operate as a distributed server, in which case the data received from each server may be distributed.

[0301] The scope of the present disclosure includes software or executable commands (e.g., operating systems, applications, firmware, programs, etc.) for enabling operations according to various embodiments of the method to be performed on a device or computer, and a non-transitory computer-readable medium having such software or commands stored thereon and executable on a device or computer.

[0302] Industrial Applicability

[0303] The embodiments of the present disclosure may be used to encode or decode an image.

Claims

1. A method for decoding an image, comprising the following steps: Obtain a first flag of whether information about a sub-picture exists in a specified bitstream; Get the second flag of whether the picture header information exists in the specified slice header; obtaining slice address information specifying a slice address of the slice, wherein obtaining the slice address information is associated with a first flag specifying presence of the information about the sub-picture in the bitstream, wherein when the first flag is a first value of true, the second flag is constrained to have a second value of false; as well as The bitstream is decoded based on the first flag, the second flag, and the slice address information.

2. A method for encoding an image, comprising the following steps: encoding a first flag specifying whether information about a sub-picture is present in the bitstream; Encoding a second flag indicating whether picture header information exists in a specified slice header; encoding slice address information specifying a slice address of the slice, encoding the slice address information being associated with a first flag specifying presence of the information about the sub-picture in the bitstream, wherein when the first flag is a first value of true, the second flag is constrained to have a second value of false; as well as The bitstream is encoded based on the first flag, the second flag, and the slice address information.

3. A method for transmitting a bit stream generated by an image encoding method, the image encoding method comprising the following steps: encoding a first flag specifying whether information about a sub-picture is present in the bitstream; Encoding a second flag indicating whether picture header information exists in a specified slice header; encoding slice address information specifying a slice address of the slice, encoding the slice address information being associated with a first flag specifying presence of the information about the sub-picture in the bitstream, wherein when the first flag is a first value of true, the second flag is constrained to have a second value of false; as well as The bitstream is encoded based on the first flag, the second flag, and the slice address information.

4. A non-transitory computer-readable recording medium storing a computer program, wherein when the computer program is executed by a processor, an image encoding method is implemented, the image encoding method comprising the following steps: encoding a first flag specifying whether information about a sub-picture is present in the bitstream; Encoding a second flag indicating whether picture header information exists in a specified slice header; encoding slice address information specifying a slice address of the slice, encoding the slice address information being associated with a first flag specifying presence of the information about the sub-picture in the bitstream, wherein when the first flag is a first value of true, the second flag is constrained to have a second value of false; as well as The bitstream is encoded based on the first flag, the second flag, and the slice address information.