Image encoding / decoding method and device for signaling information related to sub picture and picture header, and method for transmitting bitstream

The image encoding/decoding method and apparatus efficiently signal sub-picture and picture header information to enhance encoding/decoding efficiency, addressing the challenge of high-resolution image compression and reducing transmission and storage costs.

JP2025089525APending Publication Date: 2025-06-12LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025053758
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-01-14
Filing Date
2025-03-27
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality images has led to a need for highly efficient image compression techniques to reduce transmission and storage costs.

Method used

An image encoding/decoding method and apparatus that improves encoding/decoding efficiency by efficiently signaling information related to sub-pictures and picture headers, including the use of flags to indicate the presence of sub-picture information and picture header information in the bitstream.

Benefits of technology

The proposed method and apparatus achieve improved encoding/decoding efficiency, enabling effective transmission, storage, and reproduction of high-resolution, high-quality images while reducing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025089525000001_ABST
    Figure 2025089525000001_ABST
Patent Text Reader

Abstract

To provide an image encoding / decoding method and device for signaling information related to a sub picture and picture header, and a method for transmitting a bitstream.SOLUTION: An image decoding method according to the present disclosure includes the steps of: acquiring a first flag specifying whether information related to a sub picture is present in a bitstream; acquiring a second flag specifying whether picture header information is present in a slice header; and decoding the bitstream on the basis of the first flag and the second flag. When the first flag specifies that the information related to the sub picture is present in the bitstream, the second flag may have a value specifying that the picture header information is not present in the slice header.SELECTED DRAWING: Figure 12
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly, to an image encoding / decoding method for signaling information related to sub-pictures and picture headers, an apparatus, and a method for transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure.

Background Art

[0002] Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has been increasing in various fields. As image data becomes higher in resolution and quality, the amount of information or bits to be transmitted increases relatively compared to conventional image data. The increase in the amount of information or bits to be transmitted results in an increase in transmission costs and storage costs.

[0003] Accordingly, there is a need for a highly efficient image compression technique for effectively transmitting, storing, and reproducing information of high-resolution, high-quality images.

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0005] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus that improve encoding / decoding efficiency by efficiently signaling information related to sub-pictures and picture headers.

[0006] Another object of the present disclosure is to provide a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0007] Furthermore, an object of the present disclosure is to provide a recording medium storing a bitstream generated by an image encoding method or apparatus according to the present disclosure.

[0008] Furthermore, an object of the present disclosure is to provide a recording medium storing a bitstream that is received by an image decoding apparatus according to the present disclosure, decoded, and used for restoring an image.

[0009] The technical problems to be solved by the present disclosure are not limited to the above-described technical problems, and other technical problems not described above will be clearly understood by those having ordinary knowledge in the technical field to which the present disclosure pertains from the following description.

Means for Solving the Problems

[0010] An image decoding method performed by an image decoding apparatus according to an aspect of the present disclosure may include: obtaining a first flag indicating whether information regarding a sub-picture exists in a bitstream; obtaining a second flag indicating whether picture header information exists in a slice header; and decoding the bitstream based on the first flag and the second flag. When the first flag indicates that information regarding the sub-picture exists in the bitstream, the second flag may have a value indicating that the picture header information does not exist in the slice header.

[0011] In the image decoding method according to the present disclosure, when the first flag indicates that information regarding the sub-picture exists in the bitstream, the slice header may include an identifier of a sub-picture including a slice related to the slice header.

[0012] In the image decoding method according to the present disclosure, when the second flag indicates that the picture header information exists in the slice header, the method may further include obtaining the picture header information from the slice header.

[0013] In the image decoding method according to the present disclosure, the second flag may have the same value for all slices within a CLVS (Coded Layer Video Sequence).

[0014] In the image decoding method according to the present disclosure, when the second flag indicates that the picture header information is present in the slice header, the NAL unit that transmits the picture header information may not be present within the CLVS.

[0015] In the image decoding method according to the present disclosure, when the second flag indicates that the picture header information is not present in the slice header, the picture header information can be obtained from a NAL unit whose NAL unit type is PH_NUT.

[0016] In the image decoding method according to the present disclosure, the first flag is signaled at a higher level of the slice, and the second flag can be signaled and included in the slice header.

[0017] An image decoding apparatus according to another aspect of the present disclosure includes a memory and at least one processor. The at least one processor obtains a first flag indicating whether information regarding a sub-picture is present in a bitstream, obtains a second flag indicating whether picture header information is present in a slice header, and can decode the bitstream based on the first flag and the second flag. When the first flag indicates that the information regarding the sub-picture is present in the bitstream, the second flag can have a value indicating that the picture header information is not present in the slice header.

[0018] An image encoding method according to another aspect of the present disclosure may include encoding a first flag indicating whether information regarding a subpicture exists in a bitstream, encoding a second flag indicating whether picture header information exists in a slice header, and encoding the bitstream based on the first flag and the second flag. When the first flag indicates that information regarding the subpicture exists in the bitstream, the second flag may have a value indicating that the picture header information does not exist in the slice header.

[0019] In the image encoding method according to the present disclosure, when the first flag indicates that information regarding the subpicture exists in the bitstream, the slice header may include an identifier of the subpicture including the slice associated with the slice header.

[0020] In the image encoding method according to the present disclosure, when the second flag indicates that the picture header information exists in the slice header, the method may further include encoding the picture header information in the slice header.

[0021] In the image encoding method according to the present disclosure, the second flag may have the same value for all slices in a CLVS (Coded Layer Video Sequence).

[0022] In the image encoding method according to the present disclosure, when the second flag indicates that the picture header information does not exist in the slice header, the picture header information may be signaled via a NAL unit whose NAL unit type is PH_NUT.

[0023] In the image encoding method according to the present disclosure, the first flag is signaled at a higher level of the slice, and the second flag may be included and signaled in the slice header.

[0024] Also, a transmission method according to another aspect of the present disclosure can transmit a bitstream generated by the image encoding apparatus or the image encoding method of the present disclosure.

[0025] A computer-readable recording medium according to another aspect of the present disclosure can store a bitstream generated by the image encoding method or the image encoding apparatus of the present disclosure.

[0026] The features briefly summarized and described above for the present disclosure are merely exemplary aspects of the detailed description of the present disclosure to be described later, and do not limit the scope of the present disclosure.

Advantages of the Invention

[0027] According to the present disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.

[0028] Also, according to the present disclosure, an image encoding / decoding method and apparatus capable of improving the encoding / decoding efficiency by efficiently signaling information regarding sub-pictures and picture headers can be provided.

[0029] Also, according to the present disclosure, a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure can be provided.

[0030] Also, according to the present disclosure, a recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure can be provided.

[0031] Also, according to the present disclosure, a recording medium storing a bitstream received by the image decoding apparatus according to the present disclosure, decoded, and used for image restoration can be provided.

[0032] The effects obtained by the present disclosure are not limited to the effects described above, and other effects not described above will be clearly understood by those having ordinary knowledge in the technical field to which the present disclosure pertains from the following description.

Brief Description of the Drawings

[0033]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Mode for Carrying Out the Invention

[0034] Hereinafter, with reference to the accompanying drawings, embodiments of the present disclosure will be described in detail so that those having ordinary knowledge in the technical field to which the present disclosure pertains can easily implement them. However, the present disclosure can be realized in various different forms and is not limited to the embodiments described herein.

[0035] In describing the embodiments of the present disclosure, if it is determined that a specific description of a known configuration or function may obscure the gist of the present disclosure, the detailed description thereof will be omitted. And in the drawings, parts not related to the description of the present disclosure are omitted, and the same reference numerals are given to the same parts.

[0036] In the present disclosure, when a component is "connected", "coupled" or "joined" to another component, this can include not only a direct connection relationship but also an indirect connection relationship in which another component exists between them. Also, when a component "includes" or "has" another component, this means that, unless otherwise stated to the contrary, it does not exclude other components and can further include other components.

[0037] In the present disclosure, terms such as "first", "second", etc. are used only for the purpose of distinguishing one component from another and do not limit the order or importance, etc. between components unless otherwise specifically mentioned. Therefore, within the scope of the present disclosure, the first component in one embodiment may be referred to as the second component in another embodiment, and similarly, the second component in one embodiment may be referred to as the first component in another embodiment.

[0038] In the present disclosure, components that are distinguished from each other are for clearly explaining their respective features and do not necessarily mean that the components are separated. That is, a plurality of components may be integrated and configured as one hardware or software unit, or one component may be distributed and configured as a plurality of hardware or software units. Therefore, even without separate mention, such integrated or distributed embodiments are also included in the scope of the present disclosure.

[0039] In the present disclosure, the components described in various embodiments do not necessarily mean essential components, and some may be optional components. Therefore, embodiments composed of a subset of the components described in one embodiment are also included in the scope of the present disclosure. Also, embodiments that further include other components in addition to the components described in various embodiments are included in the scope of the present disclosure.

[0040] The present disclosure relates to image encoding and decoding, and the terms used in the present disclosure can have the ordinary meanings in the technical field to which the present disclosure belongs unless newly defined in the present disclosure.

[0041] In the present disclosure, "picture" generally means a unit indicating any one image in a specific time period, and a slice / tile is an encoding unit constituting a part of a picture, and one picture can be composed of one or more slices / tiles. Further, a slice / tile can include one or more CTUs (coding tree units).

[0042] In the present disclosure, "pixel" or "pel" can mean the smallest unit constituting one picture (or image). Further, the term "sample" can be used as a term corresponding to a pixel. A sample can generally indicate a pixel or a pixel value, and can also indicate only a pixel / pixel value of a luma component, or can also indicate only a pixel / pixel value of a chroma component.

[0043] In the present disclosure, "unit" indicates a basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to the region. A unit can be used interchangeably with terms such as "sample array", "block", or "area" as the case may be. In general, an M×N block can include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.

[0044] In the present disclosure, "current block" can mean any one of "current coding block", "current coding unit", "block to be coded", "block to be decoded", or "block to be processed". When prediction is performed, "current block" can mean "current prediction block" or "block to be predicted". When transformation (inverse transformation) / quantization (inverse quantization) is performed, "current block" can mean "current transformation block" or "block to be transformed". When filtering is performed, "current block" can mean "block to be filtered".

[0045] In the present disclosure, "current block" can mean "luma block of the current block" unless explicitly stated as a chroma block. "Chroma block of the current block" can be explicitly expressed including an explicit description of a chroma block such as "chroma block" or "current chroma block".

[0046] In the present disclosure, " / " and "," can be interpreted as "and / or". For example, "A / B" and "A, B" can be interpreted as "A and / or B". Also, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C".

[0047] In the present disclosure, "or" can be interpreted as "and / or". For example, "A or B" can mean 1) only "A", 2) only "B", or 3) "A and B". Alternatively, in the present disclosure, "or" can mean "additionally or alternatively".

[0048] Overview of the video coding system

[0049] FIG. 1 shows a video coding system according to the present disclosure.

[0050] A video coding system according to an embodiment can include an encoding device 10 and a decoding device 20. The encoding device 10 can transmit encoded video and / or image information or data to the decoding device 20 in a file or streaming format via a digital storage medium or a network.

[0051] The encoding device 10 according to an embodiment can include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. The decoding device 20 according to an embodiment can include a reception unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 can be called a video / image encoding unit, and the decoding unit 22 can be called a video / image decoding unit. The transmission unit 13 can be included in the encoding unit 12. The reception unit 21 can be included in the decoding unit 22. The rendering unit 23 can also include a display unit, and the display unit can be configured as a separate device or an external component.

[0052] The video source generation unit 11 can obtain video / images through processes such as capture, synthesis, or generation of video / images. The video source generation unit 11 can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer or the like. In this case, the video / image capture process can be replaced by a process in which related data is generated.

[0053] The symbolization unit 12 can encode the input video / image. For compression and encoding efficiency, the symbolization unit 12 can perform a series of procedures such as prediction, transformation, quantization, etc. The symbolization unit 12 can output the encoded data (encoded video / image information) in the form of a bitstream.

[0054] The transmission unit 13 can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit 21 of the decoding device 20 via a digital storage medium or network in file or streaming format. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray (registered trademark), HDD, SSD, etc. The transmission unit 13 can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit 21 can extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit 22.

[0055] The decoding unit 22 can perform a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operations of the symbolization unit 12 to decode the video / image.

[0056] The rendering unit 23 can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0057] Overview of the image encoding device

[0058] Figure 2 is a diagram schematically showing an image encoding device to which an embodiment according to the present disclosure can be applied.

[0059] As shown in FIG. 2, the image encoding apparatus 100 can include an image dividing unit 110, a subtraction unit 115, a conversion unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse conversion unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190. The inter prediction unit 180 and the intra prediction unit 185 can be collectively referred to as a "prediction unit". The conversion unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse conversion unit 150 can be included in a residual processing unit. The residual processing unit can further include the subtraction unit 115.

[0060] All or at least a part of the plurality of components constituting the image encoding apparatus 100 can be realized by one hardware component (e.g., an encoder or a processor) according to an embodiment. Further, the memory 170 can include a DPB (decoded picture buffer) and can be realized by a digital storage medium.

[0061] The image segmentation unit 110 can divide an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). The coding unit can be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) in a QT / BT / TT (Quad-tree / binary-tree / ternary-tree) structure. For example, one coding unit can be divided into multiple coding units at a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the division of the coding unit, the quadtree structure can be applied first, and the binary tree structure and / or the ternary tree structure can be applied later. Based on the final coding unit that cannot be further divided, the coding procedure according to the present disclosure can be performed. The largest coding unit can be used as the final coding unit, and the coding units at a deeper depth obtained by dividing the largest coding unit can also be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and / or restoration described later. As another example, the processing unit of the coding procedure can be a prediction unit (PU) or a transformation unit (TU). The prediction unit and the transformation unit can be divided or partitioned from the final coding unit respectively. The prediction unit can be a unit of sample prediction, and the transformation unit can be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from the transformation coefficients.

[0062] The prediction unit (inter prediction unit 180 or intra prediction unit 185) can perform prediction on a processing target block (current block) and generate a predicted block that includes prediction samples for the current block. The prediction unit can determine whether intra prediction is applied in units of the current block or CU, or whether inter prediction is applied. The prediction unit can generate various information related to the prediction of the current block and transmit it to the entropy encoding unit 190. The information related to the prediction can be encoded by the entropy encoding unit 190 and output in the form of a bitstream.

[0063] The intra prediction unit 185 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located in the neighborhood of the current block or at a distance according to the intra prediction mode and / or intra prediction technique. The intra prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes according to the degree of fineness of the prediction direction. However, this is only an example, and more or fewer directional prediction modes can be used based on the settings. The intra prediction unit 185 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.

[0064] The inter prediction unit 180 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the peripheral block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different from each other. The temporal neighboring block can be called by names such as a collocated reference block and a collocated CU (colCU). The reference picture including the temporal neighboring block can be called a collocated picture (colPic). For example, the inter prediction unit 180 can configure a motion information candidate list based on the peripheral block and generate information indicating which candidate is used to derive the motion vector and / or the reference picture index of the current block. Based on various prediction modes, inter prediction can be performed. For example, in the case of the skip mode and the merge mode, the inter prediction unit 180 can use the motion information of the peripheral block as the motion information of the current block. In the case of the skip mode, different from the merge mode, the residual signal cannot be transmitted.In the case of the motion information prediction (MVP) mode, the motion vectors of neighboring blocks are used as motion vector predictors, and the motion vector difference and the indicator for the motion vector predictor are encoded to signal the motion vector of the current block. The motion vector difference can mean the difference between the motion vector of the current block and the motion vector predictor.

[0065] The prediction unit can generate a prediction signal based on various prediction methods and / or prediction techniques described below. For example, the prediction unit can apply intra prediction or inter prediction for predicting the current block, and can also apply intra prediction and inter prediction simultaneously. A prediction method that applies intra prediction and inter prediction simultaneously for predicting the current block can be called CIIP (combined inter and intra prediction). Also, the prediction unit can perform intra block copy (IBC) for predicting the current block. Intra block copy can be used for content image / video coding such as games, for example, like SCC (screen content coding). IBC is a method of predicting the current block using a restored reference block within the current picture at a position a predetermined distance away from the current block. When IBC is applied, the position of the reference block within the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but can be performed in the same way as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in the present disclosure.

[0066] The prediction signal generated by the prediction unit can be used to generate a restored signal or can be used to generate a residual signal. The subtraction unit 115 can subtract the prediction signal (predicted block, predicted sample array) output from the prediction unit from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array). The generated residual signal can be transmitted to the conversion unit 120.

[0067] The conversion unit 120 can apply a conversion technique to the residual signal to generate conversion coefficients (transform coefficients). For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen - Loeve Transform), GBT (Graph - Based Transform), or CNT (Conditionally Non - linear Transform). Here, GBT means the conversion obtained from a graph when representing the relationship information between pixels as a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. The conversion process can also be applied to pixel blocks having the same size of a square and can also be applied to blocks of variable size that are not square.

[0068] The quantization unit 130 can quantize the transform coefficients and transmit them to the entropy encoding unit 190. The entropy encoding unit 190 can encode the quantized signal (information regarding the quantized transform coefficients) and output it in the form of a bitstream. The information regarding the quantized transform coefficients can be called residual information. The quantization unit 130 can reorder the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.

[0069] The entropy encoding unit 190 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 190 can encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information can further include information regarding various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. The signaling information, transmitted information, and / or syntax elements referred to in the present disclosure can be encoded via the above-described encoding procedure and included in the bitstream.

[0070] The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoding unit 190 can be provided as internal / external elements of the image encoding apparatus 100, or the transmission unit can also be provided as a component of the entropy encoding unit 190.

[0071] The quantized transform coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 140 and the inverse transformation unit 150, a residual signal (residual block or residual sample) can be restored.

[0072] The addition unit 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the block to be processed as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 155 can be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture and can also be used for inter prediction of the next picture after passing through filtering as described later.

[0073] The filtering unit 160 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 170, specifically in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like. The filtering unit 160 can generate various information related to filtering as described later in the description of each filtering method and transmit it to the entropy encoding unit 190. The information related to filtering can be encoded by the entropy encoding unit 190 and output in the form of a bit stream.

[0074] The modified restored picture transmitted to the memory 170 can be used as a reference picture by the inter prediction unit 180. When inter prediction is applied through this, the image encoding apparatus 100 can avoid prediction mismatches between the image encoding apparatus 100 and the image decoding apparatus, and can also improve the encoding efficiency.

[0075] The DPB in the memory 170 can store the modified restored picture for use as a reference picture by the inter prediction unit 180. The memory 170 can store the motion information of the block in which the motion information in the current picture has been derived (or encoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 180 for utilization as the motion information of the spatial neighboring blocks or the motion information of the temporal neighboring blocks. The memory 170 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 185.

[0076] Overview of the image decoding device

[0077] FIG. 3 is a diagram schematically showing an image decoding apparatus to which an embodiment according to the present disclosure can be applied.

[0078] As shown in FIG. 3, the image decoding apparatus 200 can be configured to include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an addition unit 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 can be collectively referred to as a "prediction unit". The inverse quantization unit 220 and the inverse transform unit 230 can be included in a residual processing unit.

[0079] All or at least a part of the plurality of components constituting the image decoding apparatus 200 can be realized by one hardware component (for example, a decoder or a processor) according to an embodiment. Further, the memory 250 can include a DPB and can be realized by a digital storage medium.

[0080] The image decoding apparatus 200 that has received a bitstream including video / image information can execute a process corresponding to the process performed by the image encoding apparatus 100 of FIG. 2 to restore an image. For example, the image decoding apparatus 200 can perform decoding using the processing unit applied in the image encoding apparatus. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. Then, the restored image signal decoded and output via the image decoding apparatus 200 can be reproduced via a reproducing apparatus (not shown).

[0081] The image decoding device 200 can receive the signal output from the image encoding device of FIG. 2 in the form of a bitstream. The received signal can be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information can further include information regarding various parameter sets such as an Adaptive Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the video / image information can further include general constraint information. The image decoding device can further use the information regarding the parameter set and / or the general constraint information to decode the image. The signaling information, the received information, and / or the syntax elements referred to in the present disclosure can be obtained from the bitstream by being decoded via the decoding procedure. For example, the entropy decoding unit 210 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for image restoration and the quantized value of the conversion coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element from the bitstream, determines a context model using the syntax element information to be decoded, the information of the surrounding blocks and the decoded information of the block to be decoded, or the information of the symbol / bin decoded in the previous step, predicts the occurrence probability of the bin based on the determined context model, and performs arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model.Among the information decoded by the entropy decoding unit 210, the information related to prediction is provided to the prediction units (inter prediction unit 260 and intra prediction unit 265), and the residual values entropy decoded by the entropy decoding unit 210, that is, the quantized transform coefficients and related parameter information, can be input to the inverse quantization unit 220. Also, among the information decoded by the entropy decoding unit 210, the information related to filtering can be provided to the filtering unit 240. On the other hand, a receiving unit (not shown) that receives a signal output from the image encoding device can be further provided as an internal / external element of the image decoding device 200, or the receiving unit can be provided as a component of the entropy decoding unit 210.

[0082] On the other hand, the image decoding device according to the present disclosure can be called a video / image / picture decoding device. The image decoding device can include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoding unit 210, and the sample decoder can include at least one of the inverse quantization unit 220, the inverse transform unit 230, the addition unit 235, the filtering unit 240, the memory 250, the inter prediction unit 260, and the intra prediction unit 265.

[0083] In the inverse quantization unit 220, the quantized transform coefficients can be inverse quantized to output transform coefficients. The inverse quantization unit 220 can reorder the quantized transform coefficients in a two-dimensional block format. In this case, the reordering can be performed based on the coefficient scan order performed by the image encoding device. The inverse quantization unit 220 can perform inverse quantization on the quantized transform coefficients using a quantization parameter (for example, quantization step size information) to obtain transform coefficients.

[0084] In the inverse conversion unit 230, the conversion coefficients can be inversely converted to obtain a residual signal (residual block, residual sample array).

[0085] The prediction unit can perform prediction on the current block and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 210, and can determine a specific intra / inter prediction mode (prediction technique).

[0086] The prediction unit can generate a prediction signal based on various prediction methods (techniques) described later, which is the same as described in the explanation of the prediction unit of the image encoding device 100.

[0087] The intra prediction unit 265 can predict the current block by referring to samples within the current picture. The explanation of the intra prediction unit 185 can also be similarly applied to the intra prediction unit 265.

[0088] The inter prediction unit 260 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the peripheral blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 260 can construct a motion information candidate list based on the peripheral blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes (techniques), and the information related to the prediction can include information indicating the mode (technique) of inter prediction for the current block.

[0089] The adder 235 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter prediction unit 260 and / or the intra prediction unit 265). When there is no residual for the block to be processed as in the case where the skip mode is applied, the predicted block can be used as the restored block. The description of the adder 155 can be similarly applied to the adder 235. The adder 235 can be called a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next block to be processed in the current picture, and can also be used for inter prediction of the next picture after passing through filtering as described later.

[0090] The filtering unit 240 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 240 can apply various filtering methods to the restored picture to generate a modified restored picture, and save the modified restored picture in the memory 250, specifically in the DPB of the memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like.

[0091] The (corrected) restored picture stored in the DPB of the memory 250 can be used as a reference picture in the inter prediction unit 260. The memory 250 can store the motion information of the block from which the motion information in the current picture was derived (or decoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 260 for utilization as the motion information of spatial neighboring blocks or temporal neighboring blocks. The memory 250 can store the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 265.

[0092] In this specification, the embodiments described in the filtering unit 160, inter prediction unit 180, and intra prediction unit 185 of the image encoding apparatus 100 can be similarly or correspondingly applied to the filtering unit 240, inter prediction unit 260, and intra prediction unit 265 of the image decoding apparatus 200, respectively.

[0093] Overview of image segmentation

[0094] The video / image coding method according to the present disclosure can be performed based on the following image division structure. Specifically, procedures such as prediction, residual processing ((inverse) transformation, (inverse) quantization, etc.), syntax element coding, and filtering described later can be performed based on CTUs, CUs (and / or TUs, PUs) derived based on the division structure of the image. The image can be divided in block units, and the block division procedure can be performed in the image division unit 110 of the above-described encoding apparatus. The division-related information can be encoded by the entropy encoding unit 190 and transmitted to the decoding apparatus in the form of a bitstream. The entropy decoding unit 210 of the decoding apparatus can derive the block division structure of the current picture based on the division-related information obtained from the bitstream, and based on this, perform a series of procedures for image decoding (for example, prediction, residual processing, block / picture restoration, in-loop filtering, etc.).

[0095] A picture can be divided into a sequence of coding tree units (CTUs). FIG. 4 shows an example of dividing a picture into CTUs. A CTU can correspond to a coding tree block (CTB). Alternatively, a CTU can include a coding tree block of luma samples and two coding tree blocks of corresponding chroma samples. For example, for a picture including three sample arrays, a CTU can include an N×N block of luma samples and two corresponding blocks of chroma samples.

[0096] Overview of CTU segmentation

[0097] As described above, a coding unit can be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) according to a QT / BT / TT (Quad-tree / binary-tree / ternary-tree) structure. For example, a CTU can first be divided into a quadtree structure. Thereafter, a leaf node of the quadtree structure can be further divided by a multi-type tree structure.

[0098] The division by a quadtree means a division that divides the current CU (or CTU) into four equal parts. By the division by a quadtree, the current CU can be divided into four CUs having the same width and the same height. If the current CU is not further divided into a quadtree structure, the current CU corresponds to a leaf node of the quadtree structure. A CU corresponding to a leaf node of the quadtree structure is not further divided and can be used as the final coding unit described above. Alternatively, a CU corresponding to a leaf node of the quadtree structure can be further divided by a multi-type tree structure.

[0099] FIG. 5 is a diagram showing the division types of blocks by a multi-type tree structure. The division by a multi-type tree structure can include two divisions by a binary tree structure and two divisions by a ternary tree structure.

[0100] The two splits by the binary tree structure can include vertical binary splitting (SPLIT_BT_VER) and horizontal binary splitting (SPLIT_BT_HOR). Vertical binary splitting (SPLIT_BT_VER) means splitting the current CU vertically into two equal parts. As shown in FIG. 4, two CUs having the same height as the current CU and a width that is half the width of the current CU can be generated by vertical binary splitting. Horizontal binary splitting (SPLIT_BT_HOR) means splitting the current CU horizontally into two equal parts. As shown in FIG. 5, two CUs having a height that is half the height of the current CU and the same width as the current CU can be generated by horizontal binary splitting.

[0101] The two splits by the ternary tree structure can include vertical ternary splitting (SPLIT_TT_VER) and horizontal ternary splitting (SPLIT_TT_HOR). Vertical ternary splitting (SPLIT_TT_VER) splits the current CU vertically in a 1:2:1 ratio. As shown in FIG. 5, two CUs having the same height as the current CU and a width that is 1 / 4 of the width of the current CU, and a CU having the same height as the current CU and a width that is half the width of the current CU can be generated by vertical ternary splitting. Horizontal ternary splitting (SPLIT_TT_HOR) splits the current CU horizontally in a 1:2:1 ratio. As shown in FIG. 5, two CUs having a height that is 1 / 4 of the height of the current CU and the same width as the current CU, and one CU having a height that is half the height of the current CU and the same width as the current CU can be generated by horizontal ternary splitting.

[0102] FIG. 6 is a diagram exemplarily showing a signaling mechanism of block division information in a quadtree with nested multi-type tree structure according to the present disclosure.

[0103] Here, the CTU is treated as the root node of the quadtree, and the CTU is first divided into the quadtree structure. Information (e.g., qt_split_flag) indicating whether to perform quadtree splitting on the current CU (CTU or a node (QT_node) of the quadtree) is signaled. For example, if the qt_split_flag is the first value (e.g., "1"), the current CU can be divided into the quadtree. Also, if the qt_split_flag is the second value (e.g., "0"), the current CU is not divided into the quadtree and becomes a leaf node (QT_leaf_node) of the quadtree. Each leaf node of the quadtree can then be further divided into a multi-type tree structure. That is, the leaf node of the quadtree can become a node (MTT_node) of the multi-type tree. In the multi-type tree structure, a first flag (e.g., mtt_split_cu_flag) is signaled to indicate whether the current node is further divided. If the node is further divided (e.g., when the first flag is 1), a second flag (e.g., mtt_split_cu_verticla_flag) is signaled to indicate the splitting direction. For example, when the second flag is 1, the splitting direction is the vertical direction, and when the second flag is 0, the splitting direction can be the horizontal direction. Then, a third flag (e.g., mtt_split_cu_binary_flag) is signaled to indicate whether the splitting type is a binary splitting type or a ternary splitting type. For example, when the third flag is 1, the splitting type is the binary splitting type, and when the third flag is 0, the splitting type can be the ternary splitting type. The nodes of the multi-type tree obtained by binary splitting or ternary splitting can be further partitioned into the multi-type tree structure. However, the nodes of the multi-type tree cannot be partitioned into the quadtree structure.When the first flag is 0, the corresponding node of the multi-type tree is not further split and becomes a leaf node (MTT_leaf_node) of the multi-type tree. The CU corresponding to the leaf node of the multi-type tree can be used as the aforementioned final coding unit.

[0104] Based on the aforementioned mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the multi-type tree splitting mode (MttSplitMode) of the CU can be derived as shown in Table 1. In the following description, the multi-tree splitting mode may be abbreviated as the multi-tree splitting type or splitting type.

[0105] [Table 1]

[0106] Figure 7 shows an example in which a CTU is split into multiple CUs by applying a multi-type tree after applying a quadtree. In Figure 7, the thick block edge 710 indicates quadtree splitting, and the remaining edges 720 indicate multi-type tree splitting. A CU can correspond to a coding block CB. In one embodiment, a CU can include a coding block of luma samples and two coding blocks of chroma samples corresponding to the luma samples.

[0107] The size of the chroma component (sample) CB or TB can be derived based on the size of the luma component (sample) CB or TB according to the component ratio by the color format of the picture / image (chroma format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.). When the color format is 4:4:4, the chroma component CB / TB size can be set to be the same as the luma component CB / TB size. When the color format is 4:2:2, the width of the chroma component CB / TB can be set to half of the width of the luma component CB / TB, and the height of the chroma component CB / TB can be set to the height of the luma component CB / TB. When the color format is 4:2:0, the width of the chroma component CB / TB can be set to half of the width of the luma component CB / TB, and the height of the chroma component CB / TB can be set to half of the height of the luma component CB / TB.

[0108] In one embodiment, when the size of the CTU is 128 based on the luma sample unit, the size of the CU can have sizes ranging from 128×128, which is the same size as the CTU, to 4×4. In one embodiment, in the case of the 4:2:0 color format (or chroma format), the chroma CB size can have sizes ranging from 64×64 to 2×2.

[0109] On the other hand, in one embodiment, the CU size and the TU size can be the same. Or, a plurality of TUs can exist within the CU region. The TU size generally indicates the size of the luma component (sample) TB (Transform Block).

[0110] The TU size can be derived based on a preset maximum allowable TB size (maxTbSize). For example, when the CU size is larger than the maxTbSize, a plurality of TUs (TBs) with the maxTbSize are derived from the CU, and conversion / inverse conversion can be performed in units of the TU (TB). For example, the maximum allowable luma TB size can be 64×64, and the maximum allowable chroma TB size can be 32×32. If the width or height of the CB divided by the tree structure is larger than the maximum conversion width or height, the CB can be automatically (or implicitly) divided until it satisfies the TB size limits in the horizontal and vertical directions.

[0111] Also, for example, when intra prediction is applied, the intra prediction mode / type is derived in units of the CU (or CB), and the peripheral reference sample derivation and prediction sample generation procedures can be performed in units of the TU (or TB). In this case, one or more TUs (or TBs) can exist within one CU (or CB) region, and in this case, the plurality of TUs (or TBs) can share the same intra prediction mode / type.

[0112] On the one hand, for the quadtree coding tree scheme with a multi-type tree, the following parameters can be signaled from the encoder to the decoder as SPS syntax elements. For example, CTUsize, which is a parameter indicating the size of the root node of the quadtree, MinQTSize, which is a parameter indicating the minimum allowable size of the leaf nodes of the quadtree, MaxBTSize, which is a parameter indicating the maximum allowable size of the root node of the binary tree, MaxTTSize, which is a parameter indicating the maximum allowable size of the root node of the ternary tree, MaxMttDepth, which is a parameter indicating the maximum allowed hierarchy depth of the multi-type tree split from the leaf nodes of the quadtree, MinBtSize, which is a parameter indicating the minimum allowable leaf node size of the binary tree, and MinTtSize, which is a parameter indicating the minimum allowable leaf node size of the ternary tree, at least one of them is signaled.

[0113] In one embodiment using a 4:2:0 chroma format, the CTU size can be set to 128×128 luma blocks and two 64×64 chroma blocks corresponding to the luma blocks. In this case, MinQTSize can be set to 16×16, MaxBtSize can be set to 128×128, MaxTtSzie can be set to 64×64, MinBtSize and MinTtSize can be set to 4×4, and MaxMttDepth can be set to 4. Quad-tree splitting can be applied to the CTU to generate leaf nodes of the quad-tree. The leaf nodes of the quad-tree can be called leaf QT nodes. The leaf nodes of the quad-tree can have a size ranging from 16×16 (e.g., the MinQTSize) to 128×128 (e.g., the CTU size). If the leaf QT node is 128×128, it cannot be further split into a binary tree / trinary tree. This is because splitting in this case would exceed MaxBtsize and MaxTtsize (e.g., 64×64). In other cases, the leaf QT node can be further split into a multi-type tree. Thus, the leaf QT node is the root node for the multi-type tree, and the leaf QT node can have a multi-type tree depth (mttDepth) 0 value. If the multi-type tree depth reaches MaxMttdepth (e.g., 4), no further additional splitting can be considered. If the width of the multi-type tree node is the same as MinBtSize and is the same as or smaller than 2xMinTtSize, no further additional horizontal splitting can be considered. If the height of the multi-type tree node is the same as MinBtSize and is the same as or smaller than 2xMinTtSize, no further additional vertical splitting can be considered. When splitting is not considered in this way, the encoding device can omit signaling of the splitting information. In such a case, the decoding device can derive the splitting information to a predetermined value.

[0114] On one hand, one CTU can include a coding block of luma samples (hereinafter referred to as "luma block") and two coding blocks of chroma samples corresponding thereto (hereinafter referred to as "chroma blocks"). The coding tree scheme described above can be similarly applied to the luma block and chroma block of the current CU, or can be applied separately. Specifically, the luma block and chroma block within one CTU can be divided into the same block tree structure, and the tree structure in this case is represented as SINGLE_TREE. Or, the luma block and chroma block within one CTU can be divided into individual block tree structures, and the tree structure in this case is represented as DUAL_TREE. That is, when the CTU is divided into a dual tree, the block tree structure for the luma block and the block tree structure for the chroma block can exist separately. At this time, the block tree structure for the luma block can be called DUAL_TREE_LUMA, and the block tree structure for the chroma block can be called DUAL_TREE_CHROMA. For P and B slices / tile groups, the luma block and chroma block within one CTU can be restricted to have the same coding tree structure. However, for I slices / tile groups, the luma block and chroma block can have individual block tree structures with respect to each other. If the individual block tree structure is applied, the luma CTB (Coding Tree Block) can be divided into CUs based on a specific coding tree structure, and the chroma CTB can be divided into chroma CUs based on another coding tree structure. That is, it can be meant that the CUs within the I slice / tile group to which the individual block tree structure is applied are composed of a coding block of the luma component or coding blocks of two chroma components, and the CUs of the P or B slice / tile group can be composed of blocks of three color components (luma component and two chroma components).

[0115] In the above, the quadtree coding tree structure with a multi-type tree has been described, but the structure in which the CU is divided is not limited to this. For example, the BT structure and the TT structure can be interpreted as concepts included in a Multiple Partitioning Tree (MPT) structure, and the CU can be interpreted as being divided by the QT structure and the MPT structure. In an example where the CU is divided by the QT structure and the MPT structure, a syntax element (e.g., MPT_split_type) including information on how many blocks the leaf node of the QT structure is divided into and a syntax element (e.g., MPT_split_mode) including information on whether the leaf node of the QT structure is divided in the vertical or horizontal direction are signaled, whereby the division structure can be determined.

[0116] In another example, the CU can be divided in a way different from the QT structure, the BT structure, or the TT structure. That is, unlike the case where the CU at a lower depth is divided into 1 / 4 the size of the CU at a higher depth by the QT structure, or the CU at a lower depth is divided into 1 / 2 the size of the CU at a higher depth by the BT structure, or the CU at a lower depth is divided into 1 / 4 or 1 / 2 the size of the CU at a higher depth by the TT structure, the CU at a lower depth can, in some cases, be divided into 1 / 5, 1 / 3, 3 / 8, 3 / 5, 2 / 3, or 5 / 8 the size of the CU at a higher depth, and the way the CU is divided is not limited to this.

[0117] Thus, the quadtree coding block structure with the multi-type tree can provide a very flexible block division structure. On the other hand, due to the division types supported by the multi-type tree, in some cases, different division patterns can potentially lead to the same coding block structure result. The encoding device and the decoding device can reduce the data amount of the division information by restricting the occurrence of such redundant division patterns.

[0118] Encoding / decoding of images based on sub-pictures

[0119] One picture to be coded can be divided into a plurality of CTUs, slices, tiles, or blocks, and further, the picture can also be divided into a plurality of sub-picture units.

[0120] Within a picture, a sub-picture can be coded or decoded regardless of the coding or decoding of the preceding sub-picture. For example, different quantization can be applied to a plurality of sub-pictures, or different resolutions can be applied to them.

[0121] Furthermore, each sub-picture can be processed like an individual picture. For example, the picture to be coded can be a projected picture or a packed picture in a 360-degree image / video or an omnidirectional image / video.

[0122] In such an embodiment, a part of the picture can be rendered or displayed based on the viewport of a user terminal (e.g., a head mount display). Therefore, in order to achieve low latency, at least one sub-picture covering the viewport among the sub-pictures constituting one picture can be coded or decoded preferentially or independently of the remaining sub-pictures.

[0123] The encoding result of a sub-picture may be called a sub-bitstream or a substream, or simply a bitstream. A decoding device can decode a sub-picture from a sub-bitstream or a substream or a bitstream. In such cases, high-level syntaxes such as PPS, SPS, VPS, and / or DPS (decoding parameter sets) can be used to encode / decoder a sub-picture.

[0124] The high-level syntax (HLS) in the present disclosure can include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, and SH syntax. For example, APS (APS syntax) or PPS (PPS syntax) can include information / parameters applicable to one or more slices or pictures in common. SPS (SPS syntax) can include information / parameters applicable to one or more sequences in common. VPS (VPS syntax) can include information / parameters applicable to multiple layers in common. DPS (DPS syntax) can include information / parameters applicable to the entire video in common. For example, DPS can include information / parameters related to the concatenation of CVS (coded video sequence).

[0125] A sub-picture can form a rectangular region of an encoded picture. The sizes of sub-pictures can be set to be different from each other within a picture. For all pictures belonging to one sequence, the sizes and positions of specific individual sub-pictures can be set identically. Individual sub-picture sequences can be decoded independently. Tiles and slices (and CTBs) can be restricted not to span across sub-picture boundaries. For this purpose, an encoding device can perform encoding so that each sub-picture can be decoded independently. For this purpose, semantic restrictions in the bitstream can be required. Also, for each picture in one sequence, the arrangement of tiles, slices, and bricks within a sub-picture can be configured to be different from each other.

[0126] The sub-picture design aims at abstraction or encapsulation in a range smaller than the picture level but larger than the slice or tile group level. Thereby, VCL NAL units of an MCTS (motion constraint tile set) subset can be extracted from one VVC bitstream and processed such as being rearranged into another VVC bitstream without difficulties like modification at the VCL level. Here, MCTS is an encoding technology that enables spatial and temporal independence between tiles. When MCTS is applied, it becomes impossible to refer to the information of tiles not included in the MCTS to which the current tile belongs. When an image is divided into MCTS and encoded, independent transmission and decoding of MCTS become possible.

[0127] Such a sub-picture design has advantages in changing the viewing orientation in mixed resolution viewport dependent 360° streaming schemes.

[0128] Next, with reference to FIGS. 8 and 9, an image encoding / decoding method using slices / tiles will be described.

[0129] FIG. 8 is a flowchart showing a method by which an image encoding apparatus according to an embodiment of the present disclosure encodes an image using slices / tiles.

[0130] The image encoding apparatus can derive slices / tiles in the current picture by dividing the current picture (S810).

[0131] The image encoding apparatus can encode the current picture based on the slices / tiles derived in step S810 (S820).

[0132] FIG. 9 is a flowchart showing a method by which an image decoding apparatus according to an embodiment of the present disclosure decodes an image using slices / tiles.

[0133] The image decoding apparatus can obtain information regarding video / images from the bitstream (S910).

[0134] Then, the image decoding apparatus can derive slices / tiles in the current picture based on the information regarding video / images obtained in step S910 (S920). Here, the information regarding video / images can include information regarding slices / tiles.

[0135] Next, the image decoding apparatus can decode the current picture based on the slices / tiles derived in step S920 (S930).

[0136] In FIGS. 8 and 9, the information for a slice / tile can include various information and / or syntax elements referred to in this disclosure. The video / image information can include a high-level syntax, and the high-level syntax can include information for a slice and / or information for a tile. The high-level syntax can include a picture header, and the information in the picture header can be included in the slice header described in this disclosure. The information for a slice can include information identifying one or more slices in the current picture, and the information for a tile can include information identifying one or more tiles in the current picture. A picture can have slices that include one or more tiles.

[0137] High level syntax (HLS) signaling

[0138] As described above, the high-level syntax can be coded / signaled for video / image coding. Hereinafter, the signaling and syntax elements in the picture header and slice header according to this disclosure will be described.

[0139] Picture header and slice header

[0140] A coded picture can be composed of one or more slices. The parameters for a coded picture are signaled within a picture header (PH), and the parameters for a slice can be signaled within a slice header (SH). The picture header can be transmitted in its own NAL unit form. The slice header can be present at the start of the NAL unit that includes the payload of the slice (i.e., the slice data). Hereinafter, with reference to FIGS. 10 and 11, the syntax elements of the picture header and slice header and the meaning of each syntax element will be described.

[0141] FIG. 10 is a diagram showing an example according to the present disclosure of signaling and syntax elements in a picture header.

[0142] picture_header_rbsp() can include information that is common to all slices within the coded picture related to the picture header. For example, picture_header_rbsp() can include information such as a reference picture flag (non_reference_picture_flag), GDR picture identification information (gdr_pic_flag), no_output_of_prior_pics_flag, recovery_poc_cnt, ph_pic_parameter_set_id, etc. Here, recovery_poc_cnt is signaled in picture_header_rbsp() when gdr_pic_flag is 1.

[0143] The first value of non_reference_picture_flag (e.g., 1) indicates that the picture related to the picture header is not used as a reference picture. The second value of non_reference_picture_flag (e.g., 0) indicates that the picture related to the picture header may or may not be used as a reference picture.

[0144] The first value of gdr_pic_flag (e.g., 1) indicates that the picture related to the picture header is a GDR picture. The second value of gdr_pic_flag (e.g., 0) indicates that the picture related to the picture header is not a GDR picture.

[0145] no_output_of_prior_pics_flag affects the output of previously decoded pictures in the DPB after decoding a CLVS (coded layer video sequence) picture that is not the first picture in the bitstream.

[0146] The recovery_poc_cnt indicates the recovery point of the decoded picture in the output order.

[0147] The ph_pic_parameter_set_id indicates the value of pps_pic_parameter_set_id for the picture parameter set in use. The pps_pic_parameter_set_id is a value for identifying the PPS so that it can be referenced in other syntaxes.

[0148] The syntax elements included in the picture_header_rbsp() syntax structure in FIG. 10 can be signaled included in the picture_header_structure() syntax structure. In this case, the picture_header_structure() syntax structure can be signaled included in the picture_header_rbsp() syntax structure.

[0149] FIG. 11 is a diagram showing the syntax structure of a slice header according to an embodiment of the present disclosure.

[0150] As shown in FIG. 11, picture_header_in_slice_header_flag, picture_header_structure(), slice_subpic_id, slice_address, num_tiles_in_slice_minus1, etc. can be signaled via the slice header.

[0151] In the example shown in FIG. 11, picture_header_in_slice_header_flag can indicate whether the picture header syntax structure exists within the slice header syntax structure. The first value of picture_header_in_slice_header_flag (e.g., 1 or True) indicates that the picture header exists within the slice header, and the second value of picture_header_in_slice_header_flag (e.g., 0 or False) can indicate that the picture header does not exist within the slice header.

[0152] picture_header_structure() can be obtained based on picture_header_in_slice_header_flag. For example, picture_header_structure() can be signaled when picture_header_in_slice_header_flag is the first value. When picture_header_in_slice_header_flag is the second value, picture_header_structure() is not included in the slice header and can be signaled included in a separate NAL unit.

[0153] The slice_subpic_id can be information regarding a subpicture identifier for identifying the subpicture in which the current slice is included. The slice_subpic_id can be obtained based on the subpics_present_flag. For example, the slice_subpic_id can be signaled when the subpics_present_flag is 1. The subpics_present_flag can indicate whether there is a subpicture in the current picture or whether there is information regarding the subpicture in the bitstream. For example, the first value of the subpics_present_flag (e.g., 1 or True) can indicate that there is information regarding the subpicture in the bitstream or there can be one or more subpictures in the current picture. The second value of the subpics_present_flag (e.g., 0 or False) can indicate that there is no information regarding the subpicture in the bitstream or there is no subpicture in the current picture.

[0154] The slice_address can represent the address of the current slice in the current picture. The slice_address can be obtained based on the rect_slice_flag and / or NumTilesInPic. For example, if the rect_slice_flag is the first value (e.g., 1 or True) or if NumTilesInPic is greater than 1, the slice_address can be signaled in the slice header. At this time, the rect_slice_flag can be an indicator indicating whether the slice included in the current picture is a rectangular slice. For example, the rect_slice_flag can be signaled at the picture level (PPS or picture header). Also, NumTilesInPic can represent the number of tiles included in the current picture.

[0155] num_tiles_in_slice_minus1 can represent the number of tiles included in the current slice. num_tiles_in_slice_minus1 can be obtained based on rect_slice_flag and NumTilesInPic. For example, when rect_slice_flag is the second value (e.g., 0 or False) and NumTilesInPic is greater than 1, num_tiles_in_slice_minus1 can be signaled in the slice header.

[0156] In the embodiment shown in FIG. 11, the following content can be included as the bitstream compliance requirements related to picture_header_in_slice_header_flag.

[0157] To meet the bitstream compliance, it is required that the value of picture_header_in_slice_header_flag is the same for all slices within the CLVS.

[0158] Also, when picture_header_in_slice_header_flag is the first value (e.g., 1), to meet the bitstream compliance, it is required that there is no NAL unit of the same NAL unit type as PH_NUT in the CLVS.

[0159] Also, when picture_header_in_slice_header_flag is the second value (e.g., 0), to meet the bitstream compliance, it is required that there is an NAL unit of the same NAL unit type as PH_NUT in the PU prior to the first VCL NAL unit of the PU.

[0160] FIG. 12 is a flowchart showing a method of parsing and decoding the slice header of FIG. 11.

[0161] First, the image decoding device can obtain the first flag (picture_header_in_slice_header_flag) included in the slice header (S1210).

[0162] The first flag can indicate whether the picture header exists within the slice header. Also, the first flag can indicate whether the current picture contains only one slice.

[0163] When the first flag is the first value (e.g., 1 or True) (step S1220 - Yes), the image decoding device can obtain the picture header from the slice header (S1230). When the first flag is the second value (e.g., 0 or False) (step S1220 - No), the picture header can be obtained from the picture header NAL unit instead of the slice header (not shown).

[0164] Thereafter, at step S1240, it can be determined whether the subpics_present_flag is the first value (e.g., 1 or True). The subpics_present_flag can indicate whether the current picture contains sub - pictures. Also, the subpics_present_flag can indicate whether information regarding sub - pictures is included in the bitstream. The subpics_present_flag can be signaled at a higher level of the slice. For example, the subpics_present_flag can be included in the sequence parameter set and signaled.

[0165] When the subpics_present_flag is the first value (e.g., 1 or True) (step S1240 - Yes), the image decoding apparatus can obtain the slice_subpic_id from the slice header (S1250). When the subpics_present_flag is the second value (e.g., 0 or False) (step S1240 - No), the image decoding apparatus can omit (skip) parsing the slice_subpic_id from the slice header.

[0166] Thereafter, in step S1260, it can be determined whether the rect_slice_flag is the first value (e.g., 1 or True) and / or whether NumTilesInPic is greater than 1. The rect_slice_flag can be an indicator indicating whether the slice included in the current picture is a rectangular slice. For example, the rect_slice_flag can be signaled at the picture level (PPS or picture header). Also, NumTilesInPic can represent the number of tiles included in the current picture.

[0167] When the rect_slice_flag is the first value (e.g., 1 or True) or when NumTilesInPic is greater than 1 (step S1260 - Yes), the image decoding apparatus can obtain the slice_address from the slice header (S1270). When the rect_slice_flag is the second value (e.g., 0 or False) and NumTilesInPic is not greater than 1 (step S1260 - No), the image decoding apparatus can omit (skip) parsing the slice_address from the slice header.

[0168] Thereafter, in step S1280, it can be determined whether the rect_slice_flag is the first value (e.g., 1 or True) and / or whether NumTilesInPic is greater than 1.

[0169] If the rect_slice_flag is the first value (e.g., 1 or True), or if NumTilesInPic is not greater than 1 (step S1280 - No), the image decoding device can omit (skip) parsing num_tiles_in_slice_minus1 from the slice header. If the rect_slice_flag is the second value (e.g., 0 or False) and NumTilesInPic is greater than 1 (step S1280 - Yes), the image decoding device can obtain num_tiles_in_slice_minus1 from the slice header (S1290).

[0170] Thereafter, the image decoding device can decode the slice header by parsing subsequent syntax elements (not shown) from the slice header.

[0171] FIG. 13 is a flowchart showing a method for encoding the slice header of FIG. 11.

[0172] First, the image encoding device can determine the value of the first flag (picture_header_in_slice_header_flag) and encode it into the slice header (S1310).

[0173] If the first flag is the first value (e.g., 1 or True) (step S1320 - Yes), the image encoding device can encode the picture header into the slice header (S1330). If the first flag is the second value (e.g., 0 or False) (step S1320 - No), the picture header is not encoded into the slice header and can be signaled included in the picture header NAL unit (not shown).

[0174] Thereafter, at step S1340, it can be determined whether the subpics_present_flag is a first value (e.g., 1 or True). The subpics_present_flag can be determined and signaled at a higher level of the slice. For example, the subpics_present_flag can be included in the sequence parameter set and signaled.

[0175] When the subpics_present_flag is the first value (e.g., 1 or True) (step S1340 - Yes), the image encoding device can encode slice_subpic_id in the slice header (S1350). When the subpics_present_flag is a second value (e.g., 0 or False) (step S1340 - No), the image encoding device can omit (skip) the encoding of slice_subpic_id in the slice header.

[0176] Thereafter, at step S1360, it can be determined whether the rect_slice_flag is a first value (e.g., 1 or True) and / or whether NumTilesInPic is greater than 1.

[0177] When the rect_slice_flag is the first value (e.g., 1 or True) or NumTilesInPic is greater than 1 (step S1360 - Yes), the image encoding device can encode slice_address in the slice header (S1370). When the rect_slice_flag is a second value (e.g., 0 or False) and NumTilesInPic is not greater than 1 (step S1360 - No), the image encoding device can omit (skip) the encoding of slice_address in the slice header.

[0178] Thereafter, at step S1380, it can be determined whether the rect_slice_flag is a first value (e.g., 1 or True) and / or whether NumTilesInPic is greater than 1.

[0179] If the rect_slice_flag is the first value (e.g., 1 or True), or if NumTilesInPic is not greater than 1 (step S1380 - No), the image encoding device can omit (skip) encoding num_tiles_in_slice_minus1 in the slice header. If the rect_slice_flag is the second value (e.g., 0 or False) and NumTilesInPic is greater than 1 (step S1380 - Yes), the image encoding device can encode num_tiles_in_slice_minus1 in the slice header (S1390).

[0180] Thereafter, the image encoding device can perform encoding of the slice header by encoding subsequent syntax elements (not shown) in the slice header.

[0181] In the examples described with reference to FIGS. 12 and 13, some steps can be changed or omitted. For example, the conditions related to the encoding / decoding of slice_address and / or num_tiles_in_slice_minus1 can be changed.

[0182] Hereinafter, a method for improving the embodiments described with reference to FIGS. 11 to 13 will be described in consideration of image encoding / decoding based on sub - pictures.

[0183] The image encoding device can encode the current picture based on sub - pictures. Alternatively, the image encoding device can encode at least one sub - picture constituting the current picture and generate a bitstream including encoding information for the at least one encoded sub - picture.

[0184] The image decoding device can decode at least one sub - picture included in the current picture based on a bitstream including encoding information regarding the at least one sub - picture.

[0185] As described above, the picture_header_in_slice_header_flag can indicate whether a picture header exists within a slice header. Also, the picture_header_in_slice_header_flag may be used to indicate whether the current picture contains only one slice or further more slices. When the current picture contains only one slice, since the slice is the only slice within the current picture, some syntax elements within the slice header have fixed values. In this case, it may be efficient not to signal some syntax elements having fixed values.

[0186] Hereinafter, various configurations of the present disclosure for performing the above-described efficient signaling will be described. The following configurations may be individually applied to embodiments according to the present disclosure, or may be applied in combination.

[0187] Configuration 1

[0188] When the current picture contains only one slice, signaling for some syntax elements within the slice header can be omitted. The values of the syntax elements for which signaling is omitted can be derived or inferred by an image encoding device and / or an image decoding device.

[0189] Whether the current picture contains only one slice or not can be indicated by a predetermined indicator. Thus, for example, when it is indicated by the indicator that the current picture contains only one slice, some of the above-described syntax elements are not included in the slice header and their values can be inferred or derived. At this time, the indicator can be used as a condition indicating whether the some syntax elements are included in the slice header.

[0190] Configuration 2

[0191] The indicator described in Configuration 1 above may be, for example, picture_header_in_slice_header_flag.

[0192] Configuration 3

[0193] The syntax element within the slice header for which signaling can be omitted according to the value of picture_header_in_slice_header_flag can include at least one of the following (a) or (b).

[0194] (a) A syntax element that indicates the subpicture in which the slice is included

[0195] The reason why signaling of the syntax element (a) can be omitted is that when only one slice per picture is included, it is obvious that no subpicture is specified. For example, when picture_header_in_slice_header_flag indicates that the current picture includes only one slice, since the current picture is not encoded / decoded based on subpictures, signaling of information regarding subpictures can be omitted.

[0196] (b) A syntax element that indicates the address of the slice

[0197] The reason why signaling of the syntax element (b) can be omitted is that it is obvious that the said slice is the only first slice within the picture. For example, when picture_header_in_slice_header_flag indicates that the current picture includes only one slice, since the current slice is the only slice within the current picture, signaling of the address for the current slice can be omitted.

[0198] Configuration 4

[0199] If each picture in the sequence has only one slice, sub-pictures are not used either. For example, the syntax elements subpics_present_flag or subpic_info_present_flag regarding sub-pictures can be restricted to a second value (e.g., 0 or False). subpics_present_flag or subpic_info_present_flag can indicate whether there are sub-pictures in the current picture or whether there is information for sub-pictures in the bitstream. subpics_present_flag or subpic_info_present_flag can be signaled, for example, included in the sequence parameter set.

[0200] Similarly, when subpics_present_flag or subpic_info_present_flag is a first value (e.g., 1 or True), a flag indicating whether each picture in the sequence contains only one slice, or a flag indicating whether there is a picture header in the slice header (e.g., picture_header_in_slice_header_flag) cannot indicate that the current picture contains only one slice and cannot indicate that the picture header exists within the slice header. Therefore, for example, when subpics_present_flag or subpic_info_present_flag is a first value (e.g., 1 or True), picture_header_in_slice_header_flag can be restricted to have a second value (e.g., 0 or False).

[0201] Configuration 5

[0202] When the picture header does not exist in the picture header NAL unit but exists in the slice header, within the CLVS of a specific layer (layer A), the picture headers of all layers that reference layer A (i.e., the dependent layers of layer A) and all layers referenced by layer A can be restricted to exist in the slice header rather than in the picture header NAL unit. This restriction can be added to simplify picture boundary detection within an access unit in the case of a multi-layer bitstream.

[0203] FIG. 14 is a diagram showing the syntax structure of a slice header according to another embodiment of the present disclosure.

[0204] Since the descriptions of the same syntax elements and the same signaling conditions in the slice header structure according to the embodiment of FIG. 14 and the slice header structure according to the embodiment of FIG. 11 are the same, duplicate descriptions are omitted.

[0205] According to the embodiment of FIG. 14, the condition for signaling slice_subpic_id can be changed. Specifically, the slice header can include slice_subpic_id based on subpics_present_flag and picture_header_in_slice_header_flag. For example, when subpics_present_flag is a first value (e.g., 1 or True) and picture_header_in_slice_header_flag is a second value (e.g., 0 or False), slice_subpic_id can be signaled in the slice header. This is because, as described above, if picture_header_in_slice_header_flag has the first value, the current picture includes only one slice and encoding / decoding based on sub-pictures is not performed, so signaling of information related to sub-pictures is not necessary.

[0206] Also, according to the embodiment of FIG. 14, the conditions for signaling slice_address can be changed. Specifically, the slice header can include slice_address based on rect_slice_flag, NumTilesInPic, and picture_header_in_slice_header_flag. For example, when rect_slice_flag is a first value (e.g., 1 or True), or NumTilesInPic is greater than 1 and picture_header_in_slice_header_flag is a second value (e.g., 0 or False), slice_address can be signaled in the slice header. This is because, as described above, if picture_header_in_slice_header_flag has a first value, the current picture contains only one slice, so signaling of information regarding the slice address is unnecessary.

[0207] In the embodiment of FIG. 14, the bitstream compliance requirements for picture_header_in_slice_header_flag can be improved as follows.

[0208] First, it is required that the value of picture_header_in_slice_header_flag be the same for all slices within the CLVS.

[0209] Also, when picture_header_in_slice_header_flag is a first value (e.g., 1), it is required that no NAL unit with NAL unit type PH_NUT exists within the CLVS. This is because the picture header is included and signaled in the slice header, so a separate NAL unit for transmitting the picture header is unnecessary.

[0210] Also, when picture_header_in_slice_header_flag is the second value (e.g., 0), it is required that a NAL unit with NAL unit type PH_NUT exists in the PU preceding the first VCL NAL unit of the PU. That is, it is required that the current PU has a PH NAL unit. This is because a separate NAL unit for transmitting the picture header is needed.

[0211] Also, when subpics_present_flag or subpic_info_present_flag is the first value (e.g., 1), it is required that picture_header_in_slice_header_flag is not the first value (e.g., 1). At this time, picture_header_in_slice_header_flag can be restricted to have the second value (e.g., 0).

[0212] In the example of FIG. 14, slice_subpic_id indicates the identifier of the subpicture containing the slice. When slice_subpic_id exists, the variable SubPicIdx is derived such that SubpicIdList[SubPicIdx] is equal to slice_subpic_id. If slice_subpic_id does not exist, the variable SubPicIdx can be derived to be 0.

[0213] In the example of FIG. 14, the length (bit length) of slice_subpic_id can be derived as follows.

[0214] If sps_subpic_id_signalling_present_flag is 1, the length of slice_subpic_id can be derived as sps_subpic_id_len_minus1 + 1. Here, sps_subpic_id_signalling_present_flag can indicate whether the subpicture identifier is signalled in the sequence parameter set. sps_subpic_id_len_minus1 is the length information of the subpicture identifier and can be signalled and included in the sequence parameter set.

[0215] Or (when sps_subpic_id_signalling_present_flag is not 1), if ph_subpic_id_signalling_present_flag is 1, the length of slice_subpic_id can be derived as ph_subpic_id_len_minus1 + 1. Here, ph_subpic_id_signalling_present_flag can indicate whether the subpicture identifier is signalled in the picture header. ph_subpic_id_len_minus1 is the length information of the subpicture identifier and can be signalled and included in the picture header.

[0216] Or (when both sps_subpic_id_signalling_present_flag and ph_subpic_id_signalling_present_flag are not 1), if pps_subpic_id_signalling_present_flag is 1, the length of slice_subpic_id can be derived as pps_subpic_id_len_minus1 + 1. Here, pps_subpic_id_signalling_present_flag can indicate whether the subpicture identifier is signaled in the picture parameter set. pps_subpic_id_len_minus1 is the length information of the subpicture identifier and can be signaled included in the picture parameter set.

[0217] Or (when sps_subpic_id_signalling_present_flag, ph_subpic_id_signalling_present_flag, and pps_subpic_id_signalling_present_flag are all not 1), the length of slice_subpic_id can be derived as Ceil(Log2(sps_num_subpics_minus1 + 1)). sps_num_subpics_minus1 indicates the number of subpictures of each picture within CLVS and can be signaled included in the sequence parameter set.

[0218] slice_address indicates the slice address of the current slice. If slice_address does not exist, the value of slice_address can be inferred as 0.

[0219] picture_header_structure() can include at least one syntax element included in picture_header_rbsp() described with reference to Figure 10.

[0220] FIG. 15 is a flowchart showing a method of parsing and decoding the slice header of FIG. 14.

[0221] Steps S1510 to S1530 in FIG. 15 are the same as steps S1210 to S1230 in FIG. 12 respectively, so duplicate explanations are omitted.

[0222] Steps S1540 to S1570 in FIG. 15 can respectively correspond to steps S1240 to S1270 in FIG. 12. Therefore, duplicate explanations for common content are omitted.

[0223] According to the embodiment of FIG. 15, in step S1540, it can be determined whether the subpics_present_flag is a first value (for example, 1 or True), and whether the first flag is a second value (for example, 0 or False).

[0224] When the subpics_present_flag is the first value and the first flag is the second value (step S1540 - Yes), the image decoding device can obtain the slice_subpic_id from the slice header (S1550). When the subpics_present_flag is the second value (for example, 0 or False), or when the first flag is the first value (for example, 1 or True) (step S1540 - No), the image decoding device can omit (skip) parsing the slice_subpic_id from the slice header.

[0225] Thereafter, in step S1560, it can be determined whether the rect_slice_flag is a first value (for example, 1 or True), or whether NumTilesInPic is greater than 1, and whether the first flag is a second value (for example, 0 or False).

[0226] If the rect_slice_flag is the first value, or if NumTilesInPic is greater than 1 and the first flag is the second value (step S1560 - Yes), the image decoding device can obtain slice_address from the slice header (S1570). If the rect_slice_flag is the second value (for example, 0 or False), or if NumTilesInPic is not greater than 1, or if the first flag is the first value (for example, 1 or True) (step S1560 - No), the image decoding device can omit (skip) parsing slice_address from the slice header.

[0227] Steps S1580 - S1590 in FIG. 15 can respectively correspond to steps S1280 - S1290 in FIG. 12. Therefore, duplicate explanations are omitted.

[0228] As described with reference to FIG. 12, the image decoding device can decode the slice header by parsing subsequent syntax elements (not shown) from the slice header.

[0229] FIG. 16 is a flowchart showing a method for encoding the slice header of FIG. 14.

[0230] Steps S1610 - S1630 in FIG. 16 are the same as steps S1310 - S1330 in FIG. 13 respectively, so duplicate explanations are omitted.

[0231] Steps S1640 - S1670 in FIG. 16 can respectively correspond to steps S1340 - S1370 in FIG. 13. Therefore, duplicate explanations for the common content are omitted.

[0232] According to the embodiment of FIG. 16, in step S1640, it can be determined whether the subpics_present_flag is the first value (for example, 1 or True) and whether the first flag is the second value (for example, 0 or False).

[0233] When the subpics_present_flag is a first value and the first flag is a second value (step S1640 - Yes), the image encoding device can encode slice_subpic_id in the slice header (S1650). When the subpics_present_flag is a second value (e.g., 0 or False), or when the first flag is a first value (e.g., 1 or True) (step S1640 - No), the image encoding device can omit (skip) the encoding of slice_subpic_id in the slice header.

[0234] Thereafter, in step S1660, it can be determined whether the rect_slice_flag is a first value (e.g., 1 or True), or whether NumTilesInPic is greater than 1 and the first flag is a second value (e.g., 0 or False).

[0235] When the rect_slice_flag is a first value, or when NumTilesInPic is greater than 1 and the first flag is a second value (step S1660 - Yes), the image encoding device can encode slice_address in the slice header (S1670). When the rect_slice_flag is a second value (e.g., 0 or False), or when NumTilesInPic is not greater than 1, or when the first flag is a first value (e.g., 1 or True) (step S1660 - No), the image encoding device can omit (skip) the encoding of slice_address in the slice header.

[0236] Steps S1680 - S1690 in FIG. 16 can respectively correspond to steps S1380 - S1390 in FIG. 13. Therefore, duplicate descriptions are omitted.

[0237] As described with reference to FIG. 13, the image encoding apparatus can encode a slice header by encoding subsequent syntax elements (not shown) in the slice header.

[0238] In the example described with reference to FIGS. 15 and 16, some steps can be changed or omitted. For example, the conditions related to the encoding / decoding of slice_address and / or num_tiles_in_slice_minus1 can be changed.

[0239] As a modification of the embodiment described with reference to FIGS. 14 to 16, the improved restriction regarding the picture_header_in_slice_header_flag described above can be applied to the embodiment shown in FIG. 11. In this case, at least some of the problems of the conventional method described above can be solved. Specifically, for example, the value of the picture_header_in_slice_header_flag can be restricted based on the information regarding the sub-picture (subpics_present_flag or subpic_info_present_flag) signaled at a higher level of the slice header. More specifically, when the subpics_present_flag or subpic_info_present_flag is a first value, the picture_header_in_slice_header_flag can be restricted to have a second value. Therefore, when the subpics_present_flag (or subpic_info_present_flag) is the first value (when the information regarding the sub-picture exists in the bitstream or the current picture includes a sub-picture), the picture_header_in_slice_header_flag can indicate that the picture header does not exist in the slice header or that the current picture does not include only one slice. In the embodiment shown in FIG. 14, the slice_subpic_id can be obtained from the slice header when the subpics_present_flag is the first value and the picture_header_in_slice_header_flag is the second value. However, since the picture_header_in_slice_header_flag is restricted to have the second value if the subpics_present_flag is the first value, it is sufficient to check the subpics_present_flag as the parsing condition for the slice_subpic_id. That is, according to this modification, in steps S1540 and S1640, the determination as to whether the first flag is the second value can be omitted.According to this modification example, when subpics_present_flag or subpic_info_present_flag is the first value, the image encoding device can encode picture_header_in_slice_header_flag having the second value. Further, when subpics_present_flag or subpic_info_present_flag is the first value, the image decoding device can obtain picture_header_in_slice_header_flag having the second value.

[0240] FIG. 17 is a diagram showing a syntax structure of a slice header according to another embodiment of the present disclosure.

[0241] Since the description of the same syntax elements and the same signaling conditions in the slice header structure according to the embodiment of FIG. 17 and the slice header structure according to the embodiment of FIG. 14 is the same, duplicate description is omitted.

[0242] According to the embodiment of FIG. 17, the condition for signaling picture_header_in_slice_header_flag can be changed. Specifically, the slice header can include picture_header_in_slice_header_flag based on subpics_present_flag. For example, when subpics_present_flag is the first value (e.g., 1 or True), picture_header_in_slice_header_flag can be not signaled in the slice header. For example, when subpics_present_flag is the second value (e.g., 0 or False), picture_header_in_slice_header_flag can be signaled in the slice header. As described above, if subpics_present_flag has the first value, the current picture cannot include only one slice, so picture_header_in_slice_header_flag has a fixed value (the second value). Therefore, signaling of picture_header_in_slice_header_flag is not required. In this case, if picture_header_in_slice_header_flag is the first value, the signaled picture header can also be not signaled via the slice header.

[0243] Also, according to the embodiment of FIG. 17, when subpics_present_flag is the first value, slice_subpic_id can be signaled in the slice header.

[0244] The following description of slice_address and num_tiles_in_slice_minus1 is the same as that described with reference to FIG. 14.

[0245] In the embodiment of FIG. 17, the bitstream compliance requirements for picture_header_in_slice_header_flag are the same as those described with reference to FIG. 14.

[0246] FIG. 18 is a flowchart showing a method of parsing and decoding the slice header of FIG. 17.

[0247] The method according to FIG. 18 and the method according to FIG. 15 only differ in some conditions and orders for parsing syntax elements, and the descriptions for the commonly disclosed syntax elements can be the same.

[0248] According to the embodiment of FIG. 18, the image decoding apparatus can determine whether the value of subpics_present_flag is a second value (for example, 0 or False) in step S1810.

[0249] In step S1810, when the value of subpics_present_flag is a first value (for example, 1 or True), encoding / decoding based on sub-pictures is performed, so the current picture does not include only one slice. Therefore, in this case, the image decoding apparatus can obtain slice_subpic_id without obtaining picture_header_in_slice_header_flag and the picture header from the slice header (S1850).

[0250] When the value of subpics_present_flag is the second value in step S1810, the image decoder obtains the first flag (picture_header_in_slice_header_flag) from the slice header (S1820). The image decoder determines whether the first flag is the first value (S1830). When the first flag is the first value, the image decoder can obtain the picture header from the slice header (S1840). When the first flag is the second value, the image decoder does not obtain the picture header from the slice header. In this case, the image decoder can obtain the picture header via a separate NAL unit. When the value of subpics_present_flag is the second value in step S1810, since encoding / decoding based on subpictures is not performed, the image decoder may not obtain information (slice_subpic_id) regarding subpictures.

[0251] Steps S1860 to S1890 in FIG. 18 are the same as steps S1560 to S1590 in FIG. 15, respectively, so duplicate explanations are omitted.

[0252] As described with reference to FIG. 12, the image decoder can decode the slice header by parsing subsequent syntax elements (not shown) from the slice header.

[0253] FIG. 19 is a flowchart showing a method for encoding the slice header of FIG. 17.

[0254] The method according to FIG. 19 and the method according to FIG. 16 only differ in some conditions and orders for encoding syntax elements, and the explanations for the commonly disclosed syntax elements can be the same.

[0255] According to the embodiment of FIG. 19, the image encoder can determine whether the value of subpics_present_flag is the second value (for example, 0 or False) in step S1910.

[0256] When the value of subpics_present_flag is the first value (e.g., 1 or True) in step S1910, encoding / decoding based on sub-pictures is performed, so the current picture does not contain only one slice. Therefore, in this case, the image encoding device can encode slice_subpic_id without encoding picture_header_in_slice_header_flag and the picture header in the slice header (S1950).

[0257] When the value of subpics_present_flag is the second value in step S1910, the image encoding device can determine the value of the first flag (picture_header_in_slice_header_flag) and encode the first flag in the slice header (S1920). The image encoding device determines whether the first flag is the first value (S1930). When the first flag is the first value, the image encoding device can encode the picture header in the slice header (S1940). When the first flag is the second value, the image encoding device does not encode the picture header in the slice header. In this case, the image encoding device can signal the picture header via a separate NAL unit. In step S1910, when the value of subpics_present_flag is the second value, encoding / decoding based on sub-pictures is not performed, so the image encoding device may not encode the information (slice_subpic_id) related to the sub-picture in the slice header.

[0258] Steps S1960 to S1990 in FIG. 19 are the same as steps S1660 to S1690 in FIG. 16 respectively, so duplicate explanations are omitted.

[0259] As described with reference to FIG. 13, the image encoding device can perform encoding of the slice header by encoding subsequent syntax elements (not shown) in the slice header.

[0260] In the example described with reference to FIGS. 18 and 19, some steps can be changed or omitted. For example, the conditions related to the encoding / decoding of slice_address and / or num_tiles_in_slice_minus1 can be changed.

[0261] According to an embodiment of the present disclosure, signaling of information regarding whether a picture header exists in a slice header and / or information regarding whether a picture includes only one slice can be performed more efficiently.

[0262] Also, according to an embodiment of the present disclosure, based on whether sub-picture based encoding / decoding is performed, the above information regarding whether a picture header exists in a slice header, etc. is signaled, so that signaling of unnecessary information can be prevented.

[0263] The names of the syntax elements described in the present disclosure can include information regarding the position where the syntax element is signaled. For example, a syntax element starting with "sps_" can be meant to be signaled in an SPS (Sequence Parameter Set). Also, syntax elements starting with "pps_", "ph_", "sh_", etc. can be meant to be signaled in a PPS (Picture Parameter Set), a picture header, a slice header, etc., respectively.

[0264] The exemplary method of the present disclosure is represented as a series of operations for clarity of explanation, but this is not for limiting the order in which the steps are performed. If necessary, each step can also be performed simultaneously or in a different order. To implement the method according to the present disclosure, it can also include additional other steps in the exemplified steps, or include the remaining steps excluding some steps, or include additional other steps excluding some steps.

[0265] In the present disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform an operation (step) of checking the execution conditions and situations of the operation (step). For example, when it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or the image decoding device can perform the predetermined operation after performing an operation of checking whether the predetermined condition is satisfied.

[0266] The various embodiments of the present disclosure do not list all possible combinations, but are for explaining representative aspects of the present disclosure. The matters described in the various embodiments may be applied independently or in combinations of two or more.

[0267] Also, the various embodiments of the present disclosure can be realized by hardware, firmware, software, or combinations thereof. In the case of realization by hardware, it can be realized by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), general processors, controllers, microcontrollers, microprocessors, etc.

[0268] In addition, the image decoding device and the image encoding device to which the embodiments of the present disclosure are applied can be included in a multimedia broadcast transmission / reception device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an over-the-top video (OTT) device, an Internet streaming service providing device, a three-dimensional (3D) video device, an image phone video device, and a medical video device, etc., and can be used to process video signals or data signals. For example, as an over-the-top video (OTT) device, it can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.

[0269] FIG. 20 is a diagram exemplarily showing a content streaming system to which an embodiment according to the present disclosure can be applied.

[0270] As shown in FIG. 20, the content streaming system to which the embodiments of the present disclosure are applied can generally include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0271] The encoding server compresses the content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data to generate a bitstream, and plays a role of transmitting this to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a video camera directly generates a bitstream, the encoding server can be omitted.

[0272] The bitstream can be generated by an image encoding method and / or an image encoding apparatus to which the embodiments of the present disclosure are applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0273] The streaming server transmits multimedia data to a user device based on a user request via a Web server, and the Web server can serve as a medium for informing the user of what services are available. When the user requests a desired service from the Web server, the Web server transmits this to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system can include a separate control server. In this case, the control server can play a role in controlling commands / responses between each device in the content streaming system.

[0274] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0275] Examples of the user device may include mobile phones, smart phones, laptop computers, digital broadcast terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices, for example, smartwatches, smart glasses, HMDs (head mounted displays), digital TVs, desktop computers, digital signage, and the like.

[0276] Each server in the content streaming system can be operated as a distributed server. In this case, the data received from each server can be distributedly processed.

[0277] The scope of the present disclosure includes software or machine-executable commands (for example, operating systems, applications, firmware, programs, etc.) that enable operations according to the methods of various embodiments to be executed on a device or computer, and non-transitory computer-readable media on which such software or commands are stored and can be executed on a device or computer.

Industrial Applicability

[0278] Examples according to the present disclosure can be used for image encoding / decoding.

Claims

1. obtaining a first flag indicating whether information about a sub-picture is present in the bitstream; obtaining a second flag indicating whether picture header information is present in the slice header; and decoding the bitstream based on the first flag and the second flag; The first flag is obtained from a sequence parameter set (SPS); 11. A method for decoding an image, comprising: based on the first flag indicating that the information regarding the sub-picture is present in the bitstream, the second flag having a value indicating that the picture header information is not present in the slice header.

2. 2. The image decoding method of claim 1, wherein, based on the first flag indicating that the information regarding the sub-picture is present in the bitstream, the slice header includes an identifier of a sub-picture that includes a slice associated with the slice header.

3. The image decoding method of claim 1 , further comprising: obtaining the picture header information from the slice header based on the second flag indicating that the picture header information is present in the slice header.

4. The image decoding method of claim 1 , wherein the second flag has the same value for all slices in a coded layer video sequence (CLVS).

5. 2. The image decoding method of claim 1, wherein, based on the second flag indicating that the picture header information is present in the slice header, a network abstraction layer (NAL) unit for transmitting the picture header information is not present in a CLVS.

6. 2. The image decoding method of claim 1, wherein, based on the second flag indicating that the picture header information is not present in the slice header, the picture header information is obtained from a NAL unit whose NAL unit type is the same as PH_NUT.

7. The first flag is signaled at a higher level of a slice; The image decoding method according to claim 1 , wherein the second flag is signaled by being included in the slice header.

8. encoding a first flag indicating whether information about a sub-picture is present in the bitstream; encoding a second flag indicating whether picture header information is present in the slice header; encoding the bitstream based on the first flag and the second flag; The first flag is encoded in a sequence parameter set (SPS); 11. A method for encoding an image, comprising: based on the first flag indicating that the information regarding the sub-picture is present in the bitstream, the second flag having a value indicating that the picture header information is not present in the slice header.

9. The image encoding method of claim 8 , wherein, based on the first flag indicating that the information regarding the sub-picture is present in the bitstream, the slice header includes an identifier of a sub-picture that includes a slice associated with the slice header.

10. The image encoding method of claim 8 , further comprising the step of encoding the picture header information in the slice header based on the second flag indicating that the picture header information is present in the slice header.

11. The image coding method according to claim 8 , wherein the second flag has the same value for all slices in a coded layer video sequence (CLVS).

12. 9. The image coding method of claim 8, wherein, based on the second flag indicating that the picture header information is not present in the slice header, the picture header information is signaled via a network abstraction layer (NAL) unit whose NAL unit type is the same as PH_NUT.

13. The first flag is signaled at a higher level of a slice; The image coding method according to claim 8 , wherein the second flag is signaled by being included in the slice header.

14. 1. A method of transmitting a bitstream for an image, comprising: The method for transmitting the bitstream comprises: generating the bitstream by performing the steps of: encoding a first flag indicating whether information about a sub-picture is present in the bitstream; encoding a second flag indicating whether picture header information is present in a slice header; and encoding the bitstream based on the first flag and the second flag; transmitting the bitstream; The first flag is encoded in a sequence parameter set (SPS); A method of transmitting a bitstream, wherein, based on the first flag indicating that the information regarding the subpicture is present in the bitstream, the second flag has a value indicating that the picture header information is not present in the slice header.