Image Encoding / Decoding Method and Apparatus Based on Hybrid NAL Unit Type, and Method for Transmitting Bitstream

The image encoding/decoding method and apparatus address the challenge of efficiently compressing high-resolution images by employing a hybrid NAL unit type and sub-picture division, resulting in improved efficiency and cost-effectiveness.

JP7697088B2Active Publication Date: 2025-06-23NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024032945
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-05
Filing Date
2024-03-05
Publication Date
2025-06-23
Estimated Expiration
2041-03-05

AI Technical Summary

Technical Problem

There is a need for highly efficient image compression techniques to effectively transmit, store, and reproduce high-resolution and high-quality images, as the increased amount of image data leads to higher transmission and storage costs.

Method used

The proposed solution involves an image encoding/decoding method and apparatus that utilizes a hybrid NAL unit type and divides a picture into sub-pictures with different NAL unit types, allowing for improved encoding/decoding efficiency and flexible hybrid NAL unit types based on sub-picture structures.

Benefits of technology

This approach enhances encoding/decoding efficiency, supports diverse hybrid NAL unit types, and facilitates the transmission and storage of high-resolution images while reducing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007697088000005
    Figure 0007697088000005
  • Figure 0007697088000006
    Figure 0007697088000006
  • Figure 0007697088000007
    Figure 0007697088000007
Patent Text Reader

Abstract

To provide an image encoding / decoding method based on a mixed NAL unit type.SOLUTION: An image decoding method includes the steps of: obtaining VCL NAL (video coding layer network abstraction layer) unit type information of a current picture from a bitstream; determining a NAL unit type of each of a plurality of slices included in the current picture, based on the obtained VCL NAL unit type information; and decoding the plurality of slices based on the determined NAL unit type. The current picture includes a first subpicture and a second subpicture having different NAL unit types, based on that at least some of the plurality of slices have different NAL unit types, and a NAL unit type of the second subpicture is determined based on a NAL unit type of the first subpicture.SELECTED DRAWING: Figure 17
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly, to an image encoding / decoding method and apparatus based on a hybrid NAL unit type, and a method of transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure.

Background Art

[0002] Recently, demands for high-resolution and high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, have been increasing in various fields. As image data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to conventional image data. The increase in the amount of information or bits to be transmitted brings about an increase in transmission costs and storage costs.

[0003] Accordingly, there is a need for a highly efficient image compression technique for effectively transmitting, storing, and reproducing information of high-resolution and high-quality images.

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0005] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus based on a hybrid NAL unit type.

[0006] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus based on two or more sub-pictures having different NAL unit types.

[0007] Another object of the present disclosure is to provide a computer-readable recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0008] Further, an object of the present disclosure is to provide a computer-readable recording medium that stores a bitstream received by an image decoding apparatus according to the present disclosure, decoded, and used for image restoration.

[0009] Further, an object of the present disclosure is to provide a method for transmitting a bitstream generated by an image encoding method or apparatus according to the present disclosure.

[0010] The technical problems to be solved in the present disclosure are not limited to the above-described technical problems, and other technical problems not described above will be clearly understood by those having ordinary knowledge in the technical field to which the present disclosure pertains from the following description.

Means for Solving the Problems

[0011] An image decoding method according to an aspect of the present disclosure includes: obtaining, from a bitstream, VCL (video coding layer) NAL (network abstraction layer) unit type information of a current picture; determining, based on the obtained VCL NAL unit type information, NAL unit types of a plurality of slices in the current picture; and decoding the plurality of slices based on the determined NAL unit types, wherein based on at least some of the plurality of slices having different NAL unit types from each other, the current picture includes a first sub-picture and a second sub-picture having different NAL unit types from each other, and the NAL unit type of the second sub-picture can be determined based on the NAL unit type of the first sub-picture.

[0012] An image decoding apparatus according to another aspect of the present disclosure includes a memory and at least one processor. The at least one processor acquires VCL (video coding layer) NAL (network abstraction layer) unit type information of a current picture from a bitstream, determines the NAL unit type of each of a plurality of slices in the current picture based on the acquired VCL NAL unit type information, decodes the plurality of slices based on the determined NAL unit type, and based on at least some of the plurality of slices having different NAL unit types from each other, the current picture includes a first sub-picture and a second sub-picture having different NAL unit types from each other, and the NAL unit type of the second sub-picture can be determined based on the NAL unit type of the first sub-picture.

[0013] An image encoding method according to another aspect of the present disclosure includes a step of dividing a current picture into one or more sub-pictures, a step of determining the NAL unit type of each of a plurality of slices included in the one or more sub-pictures, and a step of encoding the plurality of slices based on the determined NAL unit type. Based on at least some of the plurality of slices having different NAL unit types from each other, the current picture includes a first sub-picture and a second sub-picture having different NAL unit types from each other, and the NAL unit type of the second sub-picture can be determined based on the NAL unit type of the first sub-picture.

[0014] A computer-readable recording medium according to another aspect of the present disclosure can store a bitstream generated by the image encoding method or the image encoding apparatus of the present disclosure.

[0015] A transmission method according to another aspect of the present disclosure can transmit a bitstream generated by the image encoding apparatus or the image encoding method of the present disclosure.

[0016] The features briefly summarized and described above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure to be described later, and do not limit the scope of the present disclosure.

Advantages of the Invention

[0017] According to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0018] Also, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus based on a hybrid NAL unit type.

[0019] Also, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus based on two or more sub-pictures having different NAL unit types from each other.

[0020] Also, according to the present disclosure, it is possible to provide a computer-readable recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0021] Also, the present disclosure can provide a computer-readable recording medium storing a bitstream received by the image decoding apparatus according to the present disclosure, decoded, and used for restoring an image.

[0022] Also, according to the present disclosure, it is possible to provide a method of transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0023] The effects obtained in the present disclosure are not limited to the effects described above, and other effects not described above will be clearly understood by those of ordinary skill in the technical field to which the present disclosure pertains from the following description.

Brief Description of the Drawings

[0024]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Mode for Carrying Out the Invention

[0025] Hereinafter, with reference to the accompanying drawings, embodiments of the present disclosure will be described in detail so that those having ordinary knowledge in the technical field to which the present disclosure pertains can easily implement them. However, the present disclosure can be realized in various different forms and is not limited to the embodiments described herein.

[0026] When explaining the embodiments of the present disclosure, if it is determined that a specific explanation of a known configuration or function may obscure the gist of the present disclosure, the detailed explanation thereof will be omitted. In the drawings, parts not related to the explanation of the present disclosure are omitted, and the same reference numerals are given to the same parts.

[0027] In the present disclosure, when a certain component is "connected", "coupled" or "connected" to another component, this can include not only a direct connection relationship but also an indirect connection relationship in which another component exists between them. Further, when a certain component "includes" or "has" another component, this means that, unless otherwise stated to the contrary, it does not exclude another component but can further include another component.

[0028] In the present disclosure, terms such as "first" and "second" are used only for the purpose of distinguishing one component from another and do not limit the order or importance between components, etc., unless otherwise specifically mentioned. Therefore, within the scope of the present disclosure, the first component of one embodiment may be called the second component in another embodiment, and similarly, the second component of one embodiment may be called the first component in another embodiment.

[0029] In the present disclosure, the components distinguished from each other are for clearly explaining their respective features, and do not necessarily mean that the components are separated. That is, a plurality of components may be integrated and configured as one hardware or software unit, or one component may be distributed and configured as a plurality of hardware or software units. Therefore, without separate mention, such integrated or distributed embodiments are also included in the scope of the present disclosure.

[0030] In the present disclosure, the components described in various embodiments do not necessarily mean essential components, and some may be optional components. Therefore, embodiments constituted by a subset of the components described in one embodiment are also included in the scope of the present disclosure. In addition, embodiments including further other components in the components described in various embodiments are also included in the scope of the present disclosure.

[0031] The present disclosure relates to image encoding and decoding, and the terms used in the present disclosure can have the ordinary meanings in the technical field to which the present disclosure belongs, unless newly defined in the present disclosure.

[0032] In the present disclosure, "picture" generally means a unit indicating any one image in a specific time period, and a slice / tile is an encoding unit constituting a part of a picture, and one picture can be constituted by one or more slices / tiles. In addition, a slice / tile can include one or more CTUs (coding tree units).

[0033] In the present disclosure, "pixel" or "pel" can mean the smallest unit constituting one picture (or image). In addition, the term "sample" can be used as a term corresponding to a pixel. A sample can generally indicate a pixel or a pixel value, and can also indicate only the pixel / pixel value of the luma component, or can also indicate only the pixel / pixel value of the chroma component.

[0034] In the present disclosure, a "unit" can indicate a basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to the region. A unit can be used interchangeably with terms such as "sample array", "block", or "area" as the case may be. In general, an M×N block can include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.

[0035] In the present disclosure, "current block" can mean any one of "current coding block", "current coding unit", "block to be coded", "block to be decoded", or "block to be processed". When prediction is performed, "current block" can mean "current prediction block" or "block to be predicted". When transformation (inverse transformation) / quantization (inverse quantization) is performed, "current block" can mean "current transformation block" or "block to be transformed". When filtering is performed, "current block" can mean "block to be filtered".

[0036] Also, in the present disclosure, unless explicitly stated as a chroma block, "current block" can mean a block that includes both a luma component block and a chroma component block or the "luma block of the current block". The luma component block of the current block can be expressed explicitly including an explicit description of the luma component block such as "luma block" or "current luma block". Also, the chroma component block of the current block can be expressed explicitly including an explicit description of the chroma component block such as "chroma block" or "current chroma block".

[0037] In the present disclosure, " / " and "," can be interpreted as "and / or". For example, "A / B" and "A, B" can be interpreted as "A and / or B". Also, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C".

[0038] In the present disclosure, "or" can be interpreted as "and / or". For example, "A or B" can mean 1) only "A", 2) only "B", or 3) "A and B". Alternatively, in the present disclosure, "or" can mean "additionally or alternatively".

[0039] Overview of the video coding system

[0040] FIG. 1 is a diagram schematically showing a video coding system to which an embodiment according to the present disclosure can be applied.

[0041] A video coding system according to an embodiment can include an encoding device 10 and a decoding device 20. The encoding device 10 can transmit encoded video and / or image information or data in a file or streaming format to the decoding device 20 via a digital storage medium or a network.

[0042] An encoding device 10 according to an embodiment can include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. A decoding device 20 according to an embodiment can include a reception unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 can be called a video / image encoding unit, and the decoding unit 22 can be called a video / image decoding unit. The transmission unit 13 can be included in the encoding unit 12. The reception unit 21 can be included in the decoding unit 22. The rendering unit 23 can also include a display unit, and the display unit can be configured as a separate device or an external component.

[0043] The video source generation unit 11 can acquire video / images through processes such as video / image capture, synthesis, or generation. The video source generation unit 11 can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated through a computer, etc., and in this case, the video / image capture process can be replaced by the process in which related data is generated.

[0044] The encoding unit 12 can encode the input video / image. The encoding unit 12 can perform a series of procedures such as prediction, transformation, quantization, etc. for compression and encoding efficiency. The encoding unit 12 can output the encoded data (encoded video / image information) in the form of a bitstream.

[0045] The transmission unit 13 can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit 21 of the decoding device 20 via a digital storage medium or a network in the form of a file or a streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit 13 can include elements for generating a media file through a predetermined file format, and can include elements for transmission via a broadcast / communication network. The receiving unit 21 can extract / receive the bitstream from the storage medium or the network and transmit it to the decoding unit 22.

[0046] The decoding unit 22 can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding unit 12.

[0047] The rendering unit 23 can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0048] Overview of the image encoding device

[0049] FIG. 2 is a diagram schematically showing an image encoding apparatus to which an embodiment according to the present disclosure can be applied.

[0050] As shown in FIG. 2, the image encoding apparatus 100 can include an image division unit 110, a subtraction unit 115, a transformation unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transformation unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190. The inter prediction unit 180 and the intra prediction unit 185 can be collectively referred to as a "prediction unit". The transformation unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse transformation unit 150 can be included in a residual processing unit. The residual processing unit can further include the subtraction unit 115.

[0051] All or at least a part of the plurality of components constituting the image encoding apparatus 100 can be realized by one hardware component (for example, an encoder or a processor) according to an embodiment. Also, the memory 170 can include a DPB (decoded picture buffer) and can be realized by a digital storage medium.

[0052] The image segmentation unit 110 can divide an input image (or picture, frame) input to the image encoding apparatus 100 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). The coding unit can be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) according to a QT / BT / TT (Quad-tree / Binary-tree / Ternary-tree) structure. For example, one coding unit can be divided into a plurality of coding units at a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the division of the coding unit, the quadtree structure can be applied first, and the binary tree structure and / or the ternary tree structure can be applied later. Based on the final coding unit that cannot be further divided, the coding procedure according to the present disclosure can be performed. The largest coding unit can be used as the final coding unit, and the coding units at a lower depth obtained by dividing the largest coding unit can also be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, conversion, and / or restoration, which will be described later. As another example, the processing unit of the coding procedure can be a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). The prediction unit and the transform unit can be divided or partitioned from the final coding unit, respectively. The prediction unit can be a unit of sample prediction, and the transform unit can be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0053] The prediction unit (inter prediction unit 180 or intra prediction unit 185) can perform prediction on a processing target block (current block) and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra prediction is applied in units of the current block or CU, or whether inter prediction is applied. The prediction unit can generate various information regarding the prediction of the current block and transmit it to the entropy encoding unit 190. The information regarding the prediction can be encoded by the entropy encoding unit 190 and output in the form of a bitstream.

[0054] The intra prediction unit 185 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located in the neighborhood of the current block or at a distance according to the intra prediction mode and / or intra prediction technique. The intra prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes according to the degree of fineness of the prediction direction. However, this is only an example, and more or fewer directional prediction modes can be used based on the setting. The intra prediction unit 185 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.

[0055] The inter prediction unit 180 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on the reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the peripheral block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different from each other. The temporal neighboring block can be called by names such as a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block can be called a collocated picture (colPic). For example, the inter prediction unit 180 can construct a motion information candidate list based on the peripheral block and generate information indicating which candidate is used to derive the motion vector and / or the reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 180 can use the motion information of the peripheral block as the motion information of the current block. In the case of the skip mode, different from the merge mode, the residual signal cannot be transmitted.In the case of the motion information prediction (MVP) mode, the motion vectors of neighboring blocks are used as motion vector predictors, and the motion vector difference and the indicator for the motion vector predictor are encoded to signal the motion vector of the current block. The motion vector difference can mean the difference between the motion vector of the current block and the motion vector predictor.

[0056] The prediction unit can generate a prediction signal based on various prediction methods and / or prediction techniques described later. For example, the prediction unit can apply intra prediction or inter prediction for the prediction of the current block, and can also apply intra prediction and inter prediction simultaneously. The prediction method of applying intra prediction and inter prediction simultaneously for the prediction of the current block can be called CIIP (combined inter and intra prediction). In addition, the prediction unit can also perform intra block copy (IBC) for the prediction of the current block. Intra block copy can be used for content image / video coding such as games, for example, like SCC (screen content coding). IBC is a method of predicting the current block using a restored reference block within the current picture at a position separated from the current block by a predetermined distance. When IBC is applied, the position of the reference block within the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but can be performed in the same way as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in the present disclosure.

[0057] The prediction signal generated by the prediction unit can be used to generate a restored signal or can be used to generate a residual signal. The subtraction unit 115 can subtract the prediction signal (predicted block, predicted sample array) output from the prediction unit from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array). The generated residual signal can be transmitted to the conversion unit 120.

[0058] The conversion unit 120 can apply a conversion technique to the residual signal to generate transform coefficients. For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen - Loeve Transform), GBT (Graph - Based Transform), or CNT (Conditionally Non - linear Transform). Here, GBT means the transform obtained from a graph when representing the relationship information between pixels as a graph. CNT means the transform obtained based on generating a prediction signal using all previously reconstructed pixels. The conversion process can be applied to pixel blocks having the same size of a square or can be applied to blocks of variable size that are not square.

[0059] The quantization unit 130 can quantize the transform coefficients and transmit them to the entropy encoding unit 190. The entropy encoding unit 190 can encode the quantized signal (information regarding the quantized transform coefficients) and output it in the form of a bitstream. The information regarding the quantized transform coefficients can be called residual information. The quantization unit 130 can reorder the block-form quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients.

[0060] The entropy encoding unit 190 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 190 can also encode, together or separately, information necessary for video / image restoration (such as the values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (such as the encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information can further include information regarding various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. The signaling information, the transmitted information, and / or the syntax elements referred to in the present disclosure can be encoded through the above-described encoding procedure and included in the bitstream.

[0061] The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoding unit 190 can be provided as internal / external elements of the image encoding apparatus 100, or the transmission unit can also be provided as a component of the entropy encoding unit 190.

[0062] The quantized transform coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 140 and the inverse transformation unit 150, a residual signal (residual block or residual sample) can be restored.

[0063] The addition unit 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the block to be processed as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 155 can be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after passing through filtering as described later.

[0064] The filtering unit 160 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and can save the modified restored picture in the memory 170, specifically in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 160 can generate various information related to filtering as described later in the description of each filtering method and transmit it to the entropy encoding unit 190. The information related to filtering can be encoded by the entropy encoding unit 190 and output in the form of a bit stream.

[0065] The modified restored picture transmitted to the memory 170 can be used as a reference picture in the inter prediction unit 180. When inter prediction is applied through this, the image encoding device 100 can avoid prediction mismatches between the image encoding device 100 and the image decoding device, and can also improve the encoding efficiency.

[0066] The DPB in the memory 170 can save the modified restored picture for use as a reference picture in the inter prediction unit 180. The memory 170 can save the motion information of the block where the motion information in the current picture was derived (or encoded) and / or the motion information of the block in the already restored picture. The saved motion information can be transmitted to the inter prediction unit 180 for utilization as the motion information of the spatial neighboring blocks or the motion information of the temporal neighboring blocks. The memory 170 can save the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 185.

[0067] Overview of the image decoding device

[0068] FIG. 3 is a diagram schematically showing an image decoding apparatus to which an embodiment according to the present disclosure can be applied.

[0069] As shown in FIG. 3, the image decoding apparatus 200 can be configured to include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an addition unit 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 can be collectively referred to as a "prediction unit". The inverse quantization unit 220 and the inverse transform unit 230 can be included in a residual processing unit.

[0070] All or at least a part of the plurality of components constituting the image decoding apparatus 200 can be realized by one hardware component (for example, a decoder or a processor) according to an embodiment. Further, the memory 170 can include a DPB and can be realized by a digital storage medium.

[0071] The image decoding apparatus 200 that has received a bitstream including video / image information can execute a process corresponding to the process performed by the image encoding apparatus 100 of FIG. 2 to restore an image. For example, the image decoding apparatus 200 can perform decoding using the processing unit applied in the image encoding apparatus. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. Then, the restored image signal decoded and output via the image decoding apparatus 200 can be reproduced via a reproducing apparatus (not shown).

[0072] The image decoding device 200 can receive the signal output from the image encoding device of FIG. 2 in the form of a bitstream. The received signal can be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information can further include information regarding various parameter sets such as an Adaptive Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the video / image information can further include general constraint information. The image decoding device can further use the information regarding the parameter set and / or the general constraint information for decoding an image. The signaling information, received information, and / or syntax elements referred to in the present disclosure can be obtained from the bitstream by being decoded via the decoding procedure. For example, the entropy decoding unit 210 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for image restoration, the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element from the bitstream, determines a context model using the syntax element information to be decoded, the decoding information of the surrounding blocks and the block to be decoded, or the information of the symbol / bin decoded in the previous step, and predicts the occurrence probability of the bin based on the determined context model to perform arithmetic decoding of the bin, thereby generating a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model.Of the information decoded by the entropy decoding unit 210, the information related to prediction is provided to the prediction units (inter prediction unit 260 and intra prediction unit 265), and the residual values entropy decoded by the entropy decoding unit 210, that is, the quantized transform coefficients and related parameter information, can be input to the inverse quantization unit 220. Also, of the information decoded by the entropy decoding unit 210, the information related to filtering can be provided to the filtering unit 240. On the other hand, a receiving unit (not shown) that receives a signal output from the image encoding device can be further provided as an internal / external element of the image decoding device 200, or the receiving unit can be provided as a component of the entropy decoding unit 210.

[0073] On the other hand, the image decoding device according to the present disclosure can be called a video / image / picture decoding device. The image decoding device can include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoding unit 210, and the sample decoder can include at least one of the inverse quantization unit 220, the inverse transform unit 230, the addition unit 235, the filtering unit 240, the memory 250, the inter prediction unit 260, and the intra prediction unit 265.

[0074] In the inverse quantization unit 220, the quantized transform coefficients can be inverse quantized to output transform coefficients. The inverse quantization unit 220 can reorder the quantized transform coefficients in a two-dimensional block format. In this case, the reordering can be performed based on the coefficient scan order performed in the image encoding device. The inverse quantization unit 220 can perform inverse quantization on the quantized transform coefficients using a quantization parameter (for example, quantization step size information) to obtain transform coefficients.

[0075] In the inverse conversion unit 230, the conversion coefficients can be inversely converted to obtain a residual signal (residual block, residual sample array).

[0076] The prediction unit can perform prediction on the current block and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 210, and can determine a specific intra / inter prediction mode (prediction technique).

[0077] The prediction unit can generate a prediction signal based on various prediction methods (techniques) described later, which is the same as described in the explanation of the prediction unit of the image encoding device 100.

[0078] The intra prediction unit 265 can predict the current block by referring to samples within the current picture. The explanation of the intra prediction unit 185 can also be similarly applied to the intra prediction unit 265.

[0079] The inter prediction unit 260 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 260 can configure a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes (techniques), and the information regarding the prediction can include information indicating the mode (technique) of inter prediction for the current block.

[0080] The adder 235 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to a predicted signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 260 and / or the intra prediction unit 265). When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the restored block. The description of the adder 155 can be similarly applied to the adder 235. The adder 235 may also be referred to as a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next processing target block in the current picture, and can also be used for inter prediction of the next picture through filtering as described later.

[0081] The filtering unit 240 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 240 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 250, specifically, in the DPB of the memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like.

[0082] The (modified) restored picture stored in the DPB of the memory 250 can be used as a reference picture in the inter prediction unit 260. The memory 250 can store the motion information of the block where the motion information in the current picture has been derived (or decoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 260 for utilization as the motion information of the spatial neighboring blocks or the motion information of the temporal neighboring blocks. The memory 250 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 265.

[0083] In this specification, the embodiments described in the filtering unit 160, inter prediction unit 180, and intra prediction unit 185 of the image encoding apparatus 100 can be similarly or correspondingly applied to the filtering unit 240, inter prediction unit 260, and intra prediction unit 265 of the image decoding apparatus 200, respectively.

[0084] General image / video coding procedure

[0085] In image / video coding, pictures constituting an image / video can be encoded / decoded according to a series of decoding orders. The picture order corresponding to the output order of the decoded pictures can be set to be different from the decoding order. Based on this, during inter prediction, not only forward prediction but also backward prediction can be performed.

[0086] FIG. 4 is a flowchart schematically showing an image decoding procedure to which an embodiment according to the present disclosure can be applied.

[0087] Each procedure shown in FIG. 4 can be performed by the image encoding apparatus of FIG. 3. For example, in FIG. 4, step S410 can be performed by the entropy decoding unit 210, step S420 can be performed by a prediction unit including the intra prediction unit 265 and the inter prediction unit 260, step S430 can be performed by a residual processing unit including the inverse quantization unit 220 and the inverse transform unit 230, step S440 can be performed by the addition unit 235, and step S450 can be performed by the filtering unit 240. Step S410 can include the information decoding procedure described in the present disclosure, step S420 can include the inter / intra prediction procedure described in the present disclosure, step S430 can include the residual processing procedure described in the present disclosure, step S440 can include the block / picture restoration procedure described in the present disclosure, and step S450 can include the in-loop filtering procedure described in the present disclosure.

[0088] Referring to FIG. 4, the picture decoding procedure can generally include, as shown in the description of FIG. 3, an image / video information acquisition procedure (S410) for obtaining from a bitstream (by decoding), a picture restoration procedure (S420 - S440), and an in-loop filtering procedure (S450) for the restored picture. The picture restoration procedure can be performed based on the predicted samples and residual samples obtained through the inter / intra prediction (S420) and residual processing (S430, inverse quantization and inverse transformation for the quantized transform coefficients) processes described in the present disclosure. Through the in-loop filtering procedure for the restored picture generated by the picture restoration procedure, a modified restored picture can be generated, and the modified restored picture can be output as a decoded picture, and can also be stored in the decoded picture buffer or memory 250 of the decoding device and used as a reference picture in the inter prediction procedure during the decoding of subsequent pictures. In some cases, the in-loop filtering procedure can be omitted, and in this case, the restored picture can be output as a decoded picture, and can also be stored in the decoded picture buffer or memory 250 of the decoding device and used as a reference picture in the inter prediction procedure during the decoding of subsequent pictures. The in-loop filtering procedure (S450) can include, as described above, a deblocking filtering procedure, a SAO (sample adaptive offset) procedure, an ALF (adaptive loop filter) procedure, and / or a bilateral filter procedure, etc., and some or all of them can be omitted. Also, one or some of the deblocking filtering procedure, SAO (sample adaptive offset) procedure, ALF (adaptive loop filter) procedure, and bilateral filter procedure can be sequentially applied, or all of them can be sequentially applied. For example, after the deblocking filtering procedure is applied to the restored picture, the SAO procedure can be performed.Alternatively, for example, after the deblocking filtering procedure is applied to the reconstructed picture, the ALF procedure can be performed. This can be similarly performed in the encoding apparatus.

[0089] FIG. 5 is a flowchart schematically showing an image encoding procedure to which an embodiment according to the present disclosure can be applied.

[0090] Each procedure shown in FIG. 5 can be performed by the image encoding apparatus of FIG. 2. For example, step S510 can be performed by a prediction unit including the intra prediction unit 185 or the inter prediction unit 180, step S520 can be performed by a residual processing unit including the conversion unit 120 and / or the quantization unit 130, and step S530 can be performed by the entropy encoding unit 190. Step S510 can include the inter / intra prediction procedure described in the present disclosure, step S520 can include the residual processing procedure described in the present disclosure, and step S530 can include the information encoding procedure described in the present disclosure.

[0091] Referring to FIG. 5, the picture encoding procedure can include not only a procedure of roughly encoding information for picture restoration (e.g., prediction information, residual information, partitioning information, etc.) and outputting it in the form of a bitstream, as described with respect to FIG. 2, but also a procedure of generating a restored picture for the current picture and a procedure (optional) of applying in-loop filtering to the restored picture. The encoding apparatus can derive (modified) residual samples from the quantized transform coefficients via the inverse quantization unit 140 and the inverse transform unit 150, and generate a restored picture based on the prediction samples that are the output of step S510 and the (modified) residual samples. The restored picture generated in this way can be the same as the restored picture generated by the decoding apparatus described above. Through the in-loop filtering procedure for the restored picture, a modified restored picture can be generated, which can be stored in the decoded picture buffer or memory 170 and can be used as a reference picture in the inter prediction procedure during the encoding of subsequent pictures, in the same way as in the case of the decoding apparatus. As described above, in some cases, part or all of the in-loop filtering procedure can be omitted. When the in-loop filtering procedure is performed, (in-loop) filtering-related information (parameters) can be encoded by the entropy encoding unit 190 and output in the form of a bitstream, and the decoding apparatus can perform the in-loop filtering procedure in the same way as the encoding apparatus based on the filtering-related information.

[0092] Through such in-loop filtering procedures, noise generated during image / video coding, such as blocking artifacts and ringing artifacts, can be reduced, and subjective / objective visual quality can be improved. Also, by performing the in-loop filtering procedures in both the encoding device and the decoding device, the encoding device and the decoding device can derive the same prediction results, enhance the reliability of picture coding, and reduce the amount of data to be transmitted for picture coding.

[0093] As described above, picture restoration procedures can be performed not only in the decoding device but also in the encoding device. Restored blocks can be generated based on intra prediction / inter prediction for each block unit, and a restored picture including the restored blocks can be generated. When the current picture / slice / tile group is an I picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based only on intra prediction. On the other hand, when the current picture / slice / tile group is a P or B picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based on intra prediction or inter prediction. In this case, inter prediction can be applied to some blocks within the current picture / slice / tile group, and intra prediction can also be applied to some of the remaining blocks. The color components of a picture can include a luma component and a chroma component, and unless explicitly limited in the present disclosure, the methods and examples proposed in the present disclosure can be applied to the luma component and the chroma component.

[0094] Examples of coding hierarchy and structure

[0095] The coded video / image according to the present disclosure can be processed, for example, according to the coding hierarchy and structure described later.

[0096] FIG. 6 is a diagram showing an example of a hierarchical structure for a coded image / video.

[0097] The coded image / video can be divided into a VCL (video coding layer) that performs decoding processing of the image / video and handles itself, a lower-level system that transmits and stores the encoded information, and a NAL (network abstraction layer) that exists between the VCL and the lower-level system and is responsible for network adaptation functions.

[0098] In the VCL, it is possible to generate VCL data including compressed image data (slice data), or to generate a parameter set including information such as a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), a Video Parameter Set (VPS), or a SEI (Supplemental Enhancement Information) message that is additionally required for the decoding process of the image.

[0099] In the NAL, a NAL unit can be generated by adding header information (NAL unit header) to the RBSP (Raw Byte Sequence Payload) generated in the VCL. At this time, the RBSP refers to slice data, parameter sets, SEI messages, etc. generated in the VCL. The NAL unit header can include NAL unit type information specified by the RBSP data included in the corresponding NAL unit.

[0100] As shown in FIG. 6, NAL units can be classified into VCL NAL units and Non-VCL NAL units according to the type of RBSP generated by VCL. A VCL NAL unit can mean a NAL unit that contains information (slice data) for an image, and a Non-VCL NAL unit can mean a NAL unit that contains information (parameter set or SEI message) necessary for decoding an image.

[0101] The above-mentioned VCL NAL units and Non-VCL NAL units can be transmitted via a network with header information attached according to the data standard of the lower system. For example, the NAL unit can be transformed into a data format of a predetermined standard such as the H.266 / VVC file format, RTP (Real-time Transport Protocol), TS (Transport Stream), etc., and transmitted via various networks.

[0102] As described above, the NAL unit type can be specified according to the RBSP data structure (structure) included in the NAL unit, and information regarding such NAL unit type can be stored in the NAL unit header and signaled. For example, it can be roughly classified into a VCL NAL unit type and a Non-VCL NAL unit type according to whether the NAL unit contains information (slice data) for an image. The VCL NAL unit type can be classified according to the nature and type of the picture contained in the VCL NAL unit, etc., and the Non-VCL NAL unit type can be classified according to the type of parameter set, etc.

[0103] An example of the NAL unit type specified by the type of parameter set / information included in the Non-VCL NAL unit type, etc. is listed below.

[0104] -DCI (Decoding capability information) NAL unit type (NUT): Type for NAL units containing DCI

[0105] -VPS (Video Parameter Set) NUT: Type for NAL units containing VPS

[0106] -SPS (Sequence Parameter Set) NUT: Type for NAL units containing SPS

[0107] -PPS (Picture Parameter Set) NUT: Type for NAL units containing PPS

[0108] -APS (Adaptation Parameter Set) NUT: Type for NAL units containing APS

[0109] -PH (Picture header) NUT: Type for NUL units containing a picture header

[0110] The above NAL unit types have syntax information for the NAL unit type, and the syntax information can be stored in the NAL unit header and signaled. For example, the syntax information is nal_unit_type, and the NAL unit type can be specified using the value of nal_unit_type.

[0111] On one hand, a picture can include a plurality of slices, and one slice can include a slice header and slice data. In this case, one picture header can be further added for a plurality of slices (slice header and slice data set) within one picture. The picture header (picture header syntax) can include information / parameters that are commonly applicable to the picture. The slice header (slice header syntax) can include information / parameters that are commonly applicable to the slice. The APS (APS syntax) or PPS (PPS syntax) can include information / parameters that are commonly applicable to one or more slices or pictures. The SPS (SPS syntax) can include information / parameters that are commonly applicable to one or more sequences. The VPS (VPS syntax) can include information / parameters that are commonly applicable to multiple layers. The DCI can include information / parameters related to decoding capability.

[0112] In the present disclosure, a high level syntax (HLS) can include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DCI syntax, picture header syntax, and slice header syntax. Also, in the present disclosure, a low level syntax (LLS) can include, for example, slice data syntax, CTU syntax, coding unit syntax, transform unit syntax, and the like.

[0113] On the one hand, in the present disclosure, the image / video information encoded by the encoding device and signaled in the form of a bitstream to the decoding device not only includes information related to partitioning within a picture, intra / inter prediction information, residual information, in-loop filtering information, etc., but may also include the information of the slice header, the information of the picture header, the information of the APS, the information of the PPS, the information of the SPS, the information of the VPS, and / or the information of the DCI. Further, the image / video information may further include general constraint information and / or the information of the NAL unit header.

[0114] Overview of entry points signaling

[0115] As described above, the VCL NAL unit can include slice data as RBSP (Raw Byte Sequence Payload). The slice data is aligned in byte units within the VCL NAL unit and can include one or more subsets. At least one entry point for Random Access (RA) can be defined for the subset, and parallel processing can be performed based on the entry point.

[0116] The VVC standard supports WPP (wavefront parallel processing), which is one of various parallel processing techniques. Multiple slices within a picture can be encoded / decoded in parallel based on WPP.

[0117] To activate the parallel processing capability, the entry point information can be signaled. The image decoding device can directly access the starting point of the data segment included in the NAL unit based on the entry point information. Here, the starting point of the data segment can mean the starting point of the tile within the slice or the starting point of the CTU row within the slice.

[0118] The entry point information can be signaled within a higher level syntax, such as in a Picture Parameter Set (PPS) and / or a slice header.

[0119] FIG. 7 is a diagram showing a Picture Parameter Set (PPS) according to an embodiment of the present disclosure, and FIG. 8 is a diagram showing a slice header according to an embodiment of the present disclosure.

[0120] First, referring to FIG. 7, the Picture Parameter Set (PPS) can include an entry_point_offsets_present_flag as a syntax element indicating the presence or absence of signaling of the entry point information.

[0121] The entry_point_offsets_present_flag can indicate whether there is signaling of the entry point information within a slice header that refers to the Picture Parameter Set (PPS). For example, an entry_point_offsets_present_flag having a first value (e.g., 0) can indicate that there is no signaling of the entry point information for a tile or a specific CTU row(s) within the tile in the slice header. In contrast, an entry_point_offsets_present_flag having a second value (e.g., 1) can indicate that there is signaling of the entry point information for a tile or a specific CTU row in the slice header.

[0122] On the other hand, FIG. 7 shows a case where the entry_point_offsets_present_flag is included in the Picture Parameter Set (PPS), but this is exemplary and the embodiments of the present disclosure are not limited thereto. For example, the entry_point_offsets_present_flag may be included in the Sequence Parameter Set (SPS).

[0123] Next, referring to FIG. 8, the slice header can include offset_len_minus1 and entry_point_offset_minus1[i] as syntax elements for identifying the entry point.

[0124] offset_len_minus1 can indicate a value obtained by subtracting 1 from the bit length of entry_point_offset_minus1[i]. The value of offset_len_minus1 can have a range of 0 or more and 31 or less. In one example, offset_len_minus1 can be signaled based on a variable NumEntryPoints representing the total number of entry points. For example, offset_len_minus1 can be signaled only when NumEntryPoints is greater than 0. Also, offset_len_minus1 can be signaled based on the entry_point_offsets_present_flag described above with reference to FIG. 7. For example, offset_len_minus1 can be signaled only when the entry_point_offsets_present_flag has a second value (e.g., 1) (i.e., when there is signaling of entry point information in the slice header).

[0125] entry_point_offset_minus1[i] can represent the i-th entry point offset in bytes and can be represented by adding 1 bit to offset_len_minus1. The slice data in the NAL unit can include the same number of subsets as the value obtained by adding 1 to NumEntryPoints, and the index value indicating each of the subsets can have a range of 0 or more and NumEntryPoints or less. The first byte of the slice data in the NAL unit can be represented as byte 0.

[0126] When entry_point_offset_minus1[i] is signaled, the emulation prevention bytes included in the slice data within the NAL unit can be counted as part of the slice data for subset identification. Subset 0, which is the first subset of the slice data, can have a composition from byte 0 to entry_point_offset_minus1[0]. Similarly, subset k, which is the k-th subset of the slice data, can have a composition from firstByte[k] to lastByte[k]. Here, firstByte[k] can be derived as shown in Equation 1 below, and lastByte[k] can be derived as shown in Equation 2 below.

[0127]

Number

[0128]

Number

[0129] In Equation 1 and Equation 2, k is 1 or greater and can have a range of values from NumEntryPoints minus 1.

[0130] The last subset of the slice data (i.e., the NumEntryPoints-th subset) can be composed of the remaining bytes of the slice data.

[0131] On the one hand, before decoding a CTU including the first CTB of the CTB rows in each tile, if a predetermined synchronization process for context variables is not performed (e.g., sps_entropy_coding_sync_enabled_flag == 0), and if a slice contains one or more complete tiles, each subset of the slice data can be composed of all the encoded bits for all the CTUs in the same tile. In this case, the total number of subsets of the slice data can be the same as the total number of tiles in the slice.

[0132] In contrast, if the predetermined synchronization process is not performed and a slice contains one subset for the CTB rows within a single tile, NumEntryPoints can be 0. In this case, one subset of the slice data can be composed of all the encoded bits for all the CTUs in the slice.

[0133] In contrast, if the predetermined synchronization process is performed (e.g., sps_entropy_coding_sync_enabled_flag == 1), each subset can be composed of all the encoded bits for all the CTUs of one CTB row within one tile. In this case, the total number of subsets of the slice data can be the same as the total number of CTB rows per tile in the slice.

[0134] Overview of mixed NAL unit type

[0135] Generally, one NAL unit type can be set for one picture. As described above, the syntax information representing the NAL unit type can be stored and signaled in the NAL unit header within the NAL unit. For example, the syntax information is nal_unit_type, and the NAL unit type can be specified using the nal_unit_type value.

[0136] An example of the NAL unit type to which the embodiments according to the present disclosure can be applied is as shown in Table 1 below.

[0137]

Table 1-1

[0138]

Table 1-2

[0139] Referring to Table 1, the VCL NAL unit type can be classified into NAL unit types numbered from 0 to 12 according to the nature and type of the picture, etc. Also, the non-VCL NAL unit type can be classified into NAL unit types numbered from 13 to 31 according to the type of the parameter set, etc.

[0140] Specific examples of the VCL NAL unit type are as follows.

[0141] - IRAP (Intra Random Access Point) NAL unit type (NUT): The type for the NAL unit of the IRAP picture, which is set in the range of IDR_W_RADL to CRA_NUT.

[0142] - IDR (Instantaneous Decoding Refresh) NUT: The type for the NAL unit of the IDR picture, which is set to IDR_W_RADL or IDR_N_LP.

[0143] - CRA (Clean Random Access) NUT: The type for the NAL unit of the CRA picture, which is set to CRA_NUT.

[0144] - RADL (Random Access Decodable Leading) NUT: The type for the NAL unit of the RADL picture, which is set to RADL_NUT.

[0145] - RASL (Random Access Skipped Leading) NUT: The type for the NAL unit of the RASL picture, which is set to RASL_NUT.

[0146] - Trailing NUT: The type for the NAL unit of the trailing picture, which is set to TRAIL_NUT.

[0147] - GDR (Gradual Decoding Refresh) NUT: The type for the NAL unit of the GDR picture, which is set to GDR_NUT.

[0148] - STSA (Step-wise Temporal Sublayer Access) NUT: The type for the NAL unit of the STSA picture, which is set to STSA_NUT.

[0149] On the other hand, the VVC standard allows one picture to include multiple slices with different NAL unit types from each other. For example, one picture can include at least one first slice having a first NAL unit type and at least one second slice having a second NAL unit type different from the first NAL unit type. In this case, the NAL unit type of the picture may also be called a mixed NAL unit type. In this way, by the VVC standard supporting the mixed NAL unit type, in processes such as the content synthesis process and the encoding / decoding process, multiple pictures can be more easily reconstructed / synthesized.

[0150] However, according to the existing scheme regarding the hybrid NAL unit type, there is a limit that only the hybrid between two NAL unit types is allowed. Also, when a picture has a hybrid NAL unit type, one or more VCL NAL units of the picture must have an NAL unit type in the range from IDR_W_RADL to CRA_NUT, and the remaining VCL NAL units of the picture must have an NAL unit type in the range from TRAIL_NUT to RSV_VCL_6. As a result, although the hybrid NAL unit type is useful in the image processing process, there is a problem that it cannot be generally utilized.

[0151] To solve such a problem, according to an embodiment of the present disclosure, the hybrid between two or more NAL unit types is allowed, and more diverse types of hybrid NAL unit types can be provided.

[0152] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.

[0153] If one picture includes two or more slices and the NAL unit types of each of the slices are different from each other, the picture may be restricted to include two or more sub-pictures. That is, when one picture has a mixed NAL unit type, the picture can include two or more sub-pictures.

[0154] A sub-picture can include one or more slices and can form a rectangular area within the picture. The sizes of each of the sub-pictures within the picture can be set to be different from each other. In contrast, for all pictures belonging to one sequence, the sizes and positions of a specific individual sub-picture can be set to be the same as each other.

[0155] FIG. 9 is a diagram showing an example of a sub-picture.

[0156] Referring to FIG. 9, one picture can be divided into 18 tiles. 12 tiles can be arranged on the left side of the picture, and each of the tiles can include one slice consisting of a 4×4 CTU. Also, 6 tiles can be arranged on the right side of the picture, and each of the tiles can be respectively composed of a 2×2 CTU and include two slices stacked vertically. As a result, the picture includes 24 sub-pictures and 24 slices, and each of the sub-pictures can include one slice.

[0157] In one embodiment, each sub-picture within one picture can be treated as one picture to support the hybrid NAL unit type. When a sub-picture is treated as one picture, the sub-picture can be encoded / decoded independently regardless of the encoding / decoding results of other sub-pictures. Here, independent encoding / decoding means that the block division structure (such as a single-tree structure, a dual-tree structure, etc.), prediction mode type (such as intra prediction, inter prediction, etc.), decoding procedure, etc. of the sub-picture can be set differently from other sub-pictures. For example, when the first sub-picture is encoded / decoded based on the intra prediction mode, the second sub-picture adjacent to the first sub-picture and treated as one picture can be encoded / decoded based on the inter prediction mode.

[0158] Thus, when one picture includes two or more independent sub-pictures and each of the sub-pictures has a different NAL unit type from each other, the picture can have a hybrid NAL unit type.

[0159] FIG. 10 is a diagram showing an example of a picture having a hybrid NAL unit type.

[0160] Referring to FIG. 10, one picture 1000 can include first to third sub-pictures 1010 to 1030. The first and third sub-pictures 1010 and 1030 can each include two slices. In contrast, the second sub-picture 1020 can include four slices.

[0161] When the first to third sub-pictures 1010 to 1030 are each treated as one picture, the first to third sub-pictures 1010 to 1030 can be independently encoded to form different bitstreams. For example, the encoded slice data of the first sub-picture 1010 can be encapsulated into one or more NAL units having an NAL unit type such as RASL_NUT to form a first bitstream (Bitstream1). Also, the encoded slice data of the second sub-picture 1020 can be encapsulated into one or more NAL units having an NAL unit type such as RADL_NUT to form a second bitstream (Bitstream2). Also, the encoded slice data of the third sub-picture 1030 can be encapsulated into one or more NAL units having an NAL unit type such as RASL_NUT to form a third bitstream (Bitstream3). As a result, one picture 1000 can have a hybrid NAL unit type in which RASL_NUT and RADL_NUT are mixed.

[0162] In one embodiment, all slices included in each sub-picture within a picture can be restricted to have the same NAL unit type. For example, the two slices included in the first sub-picture 1010 can both have an NAL unit type such as RASL_NUT. Also, the four slices included in the second sub-picture 1020 can both have an NAL unit type such as RADL_NUT. Also, the two slices included in the third sub-picture 1030 can both have an NAL unit type such as RASL_NUT.

[0163] Information about sub - pictures can be signaled within a higher - level syntax, such as within a Picture Parameter Set (PPS) and a Sequence Parameter Set (SPS). Also, information regarding the application of the hybrid NAL unit type can be signaled within a higher - level syntax, such as within a Picture Parameter Set (PPS).

[0164] FIG. 11 is a diagram showing an example of a Picture Parameter Set (PPS) according to an embodiment of the present disclosure, and FIG. 12 is a diagram showing an example of a Sequence Parameter Set (SPS) according to an embodiment of the present disclosure.

[0165] First, referring to FIG. 11, the Picture Parameter Set (PPS) can include pps_mixed_nalu_types_in_pic_flag as a syntax element regarding the application of the hybrid NAL unit type.

[0166] pps_mixed_nalu_types_in_pic_flag (or mixed_nalu_types_in_pic_flag) can indicate whether the current picture has a hybrid NAL unit type. For example, pps_mixed_nalu_types_in_pic_flag having a first value (e.g., 0) can indicate that the current picture does not have a hybrid NAL unit type. In this case, the current picture can have the same NAL unit type for all VCL NAL units, for example, the same NAL unit type as an encoded slice NAL unit. In contrast, pps_mixed_nalu_types_in_pic_flag having a second value (e.g., 1) can indicate that the current picture has a hybrid NAL unit type.

[0167] In one embodiment, when the current picture has a mixed NAL unit type (e.g., pps_mixed_nalu_types_in_pic_flag == 1), the VCL NAL units of the current picture can be restricted from having NAL unit types such as GDR_NUT.

[0168] In one embodiment, when the current picture has a mixed NAL unit type (e.g., pps_mixed_nalu_types_in_pic_flag == 1), if any VCL NAL unit of the current picture has a NAL unit type (NAL unit type A) such as IDR_W_RADL, IDR_N_LP, or CRA_NUT, all other VCL NAL units of the current picture can be restricted to have a NAL unit type such as the NAL unit type A or a NAL unit type such as TRAIL_NUT. For example, if any VCL NAL unit of the current picture has a NAL unit type such as IDR_W_RADL, all other NAL units of the current picture can have a NAL unit type such as IDR_W_RADL or TRAIL_NUT.

[0169] When the current picture has a mixed NAL unit type (e.g., pps_mixed_nalu_types_in_pic_flag == 1), each sub-picture within the current picture can have any of the VCL NAL unit types described above with reference to Table 1. For example, if the sub-picture within the current picture is an IDR sub-picture, the sub-picture can have a NAL unit type such as IDR_W_RADL or IDR_N_LP. Alternatively, if the sub-picture within the current picture is a trailing sub-picture, the sub-picture can have a NAL unit type such as TRAIL_NUT.

[0170] Therefore, the pps_mixed_nalu_types_in_pic_flag having a second value (e.g., 1) can indicate that a picture referring to a Picture Parameter Set (PPS) can include slices having different NAL unit types from each other. Here, the picture can originate from a sub-picture bitstream merge operation where the encoder must ensure matching of the bitstream structure and alignment between the parameters of the original bitstream. As an example of the alignment, for a slice having an NAL unit type such as IDR_W_RADL or IDR_N_LP, if the reference picture list (RPL) syntax element does not exist in the slice header (e.g., sps_idr_rpl_present_flag == 0), and the current picture including the slice has a mixed NAL unit type (e.g., pps_mixed_nalu_types_in_pic_flag == 1), the current picture can be restricted not to include a slice having an NAL unit type such as IDR_W_RADL or IDR_N_LP.

[0171] On the other hand, when restricted such that the mixed NAL unit type is not applied to all pictures within an Output Layer Set (OLS) (e.g., gci_no_mixed_nalu_types_in_pic_constraint_flag == 1), the pps_mixed_nalu_types_in_pic_flag can have a first value (e.g., 0).

[0172] Also, the Picture Parameter Set (PPS) can include pps_no_pic_partition_flag (or no_pic_partion_flag) as a syntax element indicating how a picture is partitioned.

[0173] The pps_no_pic_partition_flag can indicate whether picture partitioning can be applied to the current picture. For example, a pps_no_pic_partition_flag having a first value (e.g., 0) can indicate that the current picture cannot be partitioned. In contrast, a pps_no_pic_partition_flag having a second value (e.g., 1) can indicate that the current picture can be partitioned into two or more tiles or slices. When the current picture has a mixed NAL unit type (e.g., pps_mixed_nalu_types_in_pic_flag == 1), the current picture can be restricted to be partitioned into two or more tiles or slices (e.g., pps_no_pic_partition_flag = 1).

[0174] Also, the picture parameter set (PPS) can include pps_num_subpics_minus1 (or num_subpics_minus1) as a syntax element representing the number of sub-pictures.

[0175] pps_num_subpics_minus1 can represent a value obtained by subtracting 1 from the number of sub-pictures included in the current picture. pps_num_subpics_minus1 can be signaled only when picture partitioning can be applied to the current picture (e.g., pps_no_pic_partition_flag == 1). When pps_num_subpics_minus1 is not signaled, the value of pps_num_subpics_minus1 can be inferred to be 0. On the other hand, the syntax element representing the number of sub-pictures can also be signaled in a higher-level syntax different from the picture parameter set (PPS), for example, within the sequence parameter set (SPS).

[0176] In one embodiment, when the current picture contains only one sub-picture (e.g., pps_num_subpics_minus1 == 0), the current picture can be restricted from having a mixed NAL unit type (e.g., pps_mixed_nalu_types_in_pic_flag = 0). That is, when the current picture has a mixed NAL unit type (e.g., pps_mixed_nalu_types_in_pic_flag == 1), the current picture can be restricted to contain two or more sub-pictures (e.g., pps_num_subpics_minus1 > 0).

[0177] Next, referring to FIG. 12, the sequence parameter set (SPS) can include sps_subpic_treated_as_pic_flag[i] (or subpic_treated_as_pic_flag[i]) as a syntax element related to the handling of sub-pictures during encoding / decoding.

[0178] sps_subpic_treated_as_pic_flag[i] can indicate whether each sub-picture in the current picture is treated as one picture. For example, sps_subpic_treated_as_pic_flag[i] with a first value (e.g., 0) can indicate that the i-th sub-picture in the current picture is not treated as one picture. In contrast, sps_subpic_treated_as_pic_flag[i] with a second value (e.g., 1) can indicate that the i-th sub-picture in the current picture is treated as one picture during the encoding / decoding process excluding the in-loop filtering operation. When sps_subpic_treated_as_pic_flag[i] is not signaled, it can be inferred that sps_subpic_treated_as_pic_flag[i] has the second value (e.g., 1).

[0179] In one embodiment, when the current picture includes two or more sub-pictures (e.g., pps_num_subpics_minus1>0) and at least one of the sub-pictures is not treated as one picture (e.g., sps_subpic_treated_as_pic_flag[i]==0), the current picture can be restricted from having a mixed NAL unit type (e.g., pps_mixed_nalu_types_in_pic_flag=0). That is, when the current picture has a mixed NAL unit type (e.g., pps_mixed_nalu_types_in_pic_flag==1), all the sub-pictures in the current picture can be restricted to be treated as one picture (e.g., sps_subpic_treated_as_pic_flag[i]=1).

[0180] Hereinafter, the NAL unit types according to the embodiments of the present disclosure will be described in detail according to picture types.

[0181] (1) IRAP (Intra Random Access Point) picture

[0182] An IRAP picture is a randomly accessible picture and can have a NAL unit type such as IDR_W_RADL, IDR_N_LP, or CRA_NUT as described above with reference to Table 1. Other pictures than the IRAP picture do not need to be referred to for inter prediction during the decoding process. The IRAP picture can include an IDR (Instantaneous decoding refresh) picture and a CRA (Clean random access) picture.

[0183] The first picture in the bitstream in the decoding order may be limited to an IRAP picture or a GDR (Gradual Decoding Refresh) picture. For a single-layer bitstream, if the essential parameter set to be referred to is available, even if no picture preceding the IRAP picture in the decoding order is decoded at all, the IRAP picture and all non-RASL (non-RASL) pictures following the IRAP picture in the decoding order can be decoded correctly.

[0184] In one embodiment, the IRAP picture can have no hybrid NAL unit type. That is, the pps_mixed_nalu_types_in_pic_flag described above for the IRAP picture can have a first value (for example, 0), and all slices in the IRAP picture can have the same NAL unit type as each other within the range of IDR_W_RADL to CRA_NUT. As a result, if the first slice decoded within the picture has an NAL unit type in the range of IDR_W_RADL to CRA_NUT, the picture can be determined to be an IRAP picture.

[0185] (2) CRA (Clean Random Access) picture

[0186] The CRA picture is one of the IRAP pictures and can have an NAL unit type such as CRA_NUT as described above with reference to Table 1. The CRA picture does not need to be referred to by other pictures other than the CRA picture for inter prediction in the decoding process.

[0187] The CRA picture may be the first picture in the bitstream in the decoding order, or may be a picture after the first. The CRA picture can be related to a RADL or RASL picture.

[0188] When NoIncorrectPicOutputFlag has a second value (e.g., 1) for a CRA picture, the RASL picture related to the CRA picture cannot be decoded because it refers to a picture that does not exist in the bitstream, and as a result, it cannot be output by the image decoder. Here, NoIncorrectPicOutputFlag can indicate whether a picture preceding a recovery point picture in the decoding order is output earlier than the recovery point picture. For example, NoIncorrectPicOutputFlag having a first value (e.g., 0) can indicate that a picture preceding a recovery point picture in the decoding order can be output earlier than the recovery point picture. In this case, the CRA picture may not be the first picture in the bitstream or the first picture following an EOS (End Of Sequence) NAL unit in the decoding order. This can mean that random access has not occurred. In contrast, NoIncorrectPicOutputFlag having a second value (e.g., 1) can indicate that a picture preceding a recovery point picture in the decoding order cannot be output earlier than the recovery point picture. In this case, the CRA picture may be the first picture in the bitstream or the first picture following an EOS NAL unit in the decoding order. This can mean that random access has occurred. On the other hand, NoIncorrectPicOutputFlag may also be called NoOutputBeforeRecoveryFlag depending on the embodiment.

[0189] For all picture units (PUs) that follow the current picture in the decoding order within the CLVS (coded layer video sequence), the reference picture list 0 (e.g., RefPicList[0]) and the reference picture list 1 (e.g., RefPicList[1]) for one slice included in the CRA sub-picture belonging to the picture unit (PUs) can be restricted so as not to include any picture that precedes the picture including the CRA sub-picture in the decoding order within the active entry. Here, a picture unit (PU) can mean a set of NAL units that are mutually related according to a predetermined classification rule for one coded picture and include a plurality of consecutive NAL units in the decoding order.

[0190] (3) IDR (Instantaneous Decoding Refresh) picture

[0191] An IDR picture is one of the IRAP pictures and can have a NAL unit type such as IDR_W_RADL or IDR_N_LP as described above with reference to Table 1. An IDR picture does not need to be referenced by other pictures other than the IDR picture for inter prediction during the decoding process.

[0192] An IDR picture may be the first picture in the bitstream in the decoding order, or may be a picture after the first. Each IDR picture can be the first picture of the CVS (Coded Video Sequence) in the decoding order.

[0193] When an IDR picture has a NAL unit type such as IDR_W_RADL for each NAL unit, the IDR picture can have associated RADL pictures. In contrast, when an IDR picture has a NAL unit type such as IDR_N_LP for each NAL unit, the IDR picture may not have associated leading pictures. On the other hand, the IDR picture may not be associated with RASL pictures.

[0194] For all picture units (PUs) that follow the current picture in decoding order within a CLVS (Coded Layer Video Sequence), the reference picture list 0 (e.g., RefPicList[0]) and the reference picture list 1 (e.g., RefPicList[1]) for one slice included in the IDR sub-picture belonging to the picture unit (PUs) can be restricted so as not to include any picture that precedes the picture including the IDR sub-picture in decoding order within an active entry.

[0195] (4) RADL (Random Access Decodable Leading) picture

[0196] A RADL picture is one of the leading pictures and can have a NAL unit type such as RADL_NUT as described above with reference to Table 1.

[0197] The RADL picture may not be used as a reference picture in the decoding process of the trailing picture related to the same IRAP picture as the RADL picture. When the field_seq_flag has a first value (e.g., 0) for the RADL picture, the RADL picture can precede in the decoding order all non-leading pictures with the same related IRAP picture. Here, the field_seq_flag can indicate whether the CLVS (Coded Layer Video Sequence) conveys a picture representing fields or a picture representing frames. For example, the field_seq_flag having a first value (e.g., 0) can indicate that the CLVS conveys a picture representing frames. In contrast, the field_seq_flag having a second value (e.g., 1) can indicate that the CLVS conveys a picture representing fields.

[0198] (5)RASL (Random Access Skipped Leading) picture

[0199] The RASL picture is one of the leading pictures and can have a NAL unit type such as RASL_NUT as described above with reference to Table 1.

[0200] In one example, all RASL pictures can be leading pictures of the related CRA picture. When the NoIncorrectPicOutputFlag has a second value (e.g., 1) for the CRA picture, the RASL picture cannot be decoded because it references a picture that does not exist in the bitstream, and as a result, it cannot be output by the image decoding device.

[0201] A RASL picture may not be used as a reference picture in the decoding process of a non-RASL picture. However, if there is a RADL picture belonging to the same layer as the RASL picture and related to the same CRA picture, the RASL picture can be used as a collocated reference picture for inter-prediction of RADL sub-pictures included in the RADL picture.

[0202] When the field_seq_flag has a first value (e.g., 0) for a RASL picture, the RASL picture can precede in the decoding order with respect to all non-leading pictures of the CRA picture related to the RASL picture.

[0203] (6) Trailing Picture

[0204] A trailing picture is a non-IRAP picture that follows in the output order with respect to the related IRAP picture or GDR picture, and may not be a STSA picture. Also, a trailing picture can follow in the decoding order with respect to the related IRAP picture. That is, a trailing picture that follows in the output order but precedes in the decoding order with respect to the related IRAP picture is not allowed.

[0205] (7) GDR (Gradual Decoding Refresh) Picture

[0206] A GDR picture is a randomly accessible picture and can have a NAL unit type such as GDR_NUT as described above with reference to Table 1.

[0207] (8) STSA (Step-wise Temporal Sublayer Access) Picture

[0208] An STSA picture is a randomly accessible picture that can have a NAL unit type such as STSA_NUT as described above with reference to Table 1.

[0209] The STSA picture may not refer to a picture having the same TemporalId as the STSA picture for inter prediction. Here, the TemporalId can be an identifier indicating temporal hierarchy, for example, a temporal sublayer in scalable video coding. As an example, the STSA picture can be restricted to have a TemporalId greater than 0.

[0210] A picture having the same TemporalID as the STSA picture and following the STSA picture in the decoding order may not refer to a picture having the same TemporalID as the STSA picture and preceding the STSA picture in the decoding order for inter prediction. The STSA picture can activate up-switching from the immediately lower sublayer of the current sublayer to which the STSA picture belongs to the current sublayer.

[0211] FIG. 13 is a diagram for explaining the decoding order and output order for each picture type.

[0212] Multiple pictures can be classified into I pictures, P pictures, or B pictures according to a prediction method. An I picture means a picture to which only intra prediction can be applied and can be decoded without referring to other pictures. An I picture can be called an intra picture and can include the aforementioned IRAP pictures. A P picture means a picture to which intra prediction and one-way inter prediction can be applied and can be decoded by referring to one other picture. A B picture means a picture to which intra prediction and bi-directional / one-way inter prediction can be applied and can be decoded by referring to one or two other pictures. P pictures and B pictures may also be called inter pictures and can include the aforementioned RADL pictures, RASL pictures, and trailing pictures.

[0213] Inter pictures can be further classified into leading pictures (LP) or non-leading pictures (NLP) according to the decoding order and the output order. A leading picture means a picture that follows an IRAP picture in the decoding order and precedes an IRAP picture in the output order and can include the aforementioned RADL pictures and RASL pictures. A non-leading picture means a picture that follows an IRAP picture in both the decoding order and the output order and can include the aforementioned trailing pictures.

[0214] In FIG. 13, each picture name can indicate the picture type described above. For example, the I5 picture can be an I picture, the B0, B2, B3, B4, and B6 pictures can be B pictures, and the P1 and P7 pictures can be P pictures. Further, in FIG. 13, each arrow can indicate the reference direction between pictures. For example, the B0 picture can be decoded by referring to the P1 picture.

[0215] Referring to FIG. 13, the I5 picture can be an IRAP picture, for example, a CRA picture. When random access is performed on the I5 picture, the I5 picture can be the first picture in the decoding order.

[0216] The B0 and P1 pictures are pictures that precede the I5 picture in the decoding order and can form a video sequence separate from the I5 picture. The B2, B3, B4, B6, and P7 pictures are pictures that follow the I5 picture in the decoding order and can form one video sequence with the I5 picture.

[0217] Since the B2, B3, and B4 pictures follow the I5 picture in the decoding order and precede the I5 picture in the output order, they can be classified as leading pictures. The B2 picture can be decoded by referring to the P1 picture that precedes the I5 picture in the decoding order. Therefore, when random access occurs to the I5 picture, the B2 picture cannot be correctly decoded because it refers to the P1 picture that does not exist in the bitstream. Picture types like the B2 picture can be called RASL pictures. In contrast, the B3 picture can be decoded by referring to the I5 picture that precedes the B3 picture in the decoding order. Therefore, when random access occurs to the I5 picture, the B3 picture can be correctly decoded by referring to the already decoded I5 picture. Also, the B4 picture can be decoded by referring to the I5 picture and the B3 picture that precede the B4 picture in the decoding order. Therefore, when random access occurs to the I5 picture, the B4 picture can be correctly decoded by referring to the already decoded I5 picture and B3 picture. Picture types like the B3 and B4 pictures can be called RADL pictures.

[0218] On the one hand, since pictures B6 and P7 follow picture I5 in the decoding order and the output order, they can be classified as non-reference pictures. Picture B6 can be decoded by referring to the preceding picture I5 and picture P7 in the decoding order. Therefore, when a random access to picture I5 occurs, picture B6 can be correctly decoded by referring to the already decoded pictures I5 and P7.

[0219] Therefore, within one video sequence, the decoding process and the output process can be performed in different orders based on the picture type. For example, if one video sequence includes IRAP pictures, reference pictures, and non-reference pictures, the decoding process is performed in the order of IRAP pictures, reference pictures, and non-reference pictures, while the output process can be performed in the order of reference pictures, IRAP pictures, and non-reference pictures.

[0220] Hereinafter, the hybrid constraint conditions for different types of hybrid NAL unit types will be described in detail.

[0221] (1) Hybrid NAL unit type 1

[0222] The hybrid between NAL unit types related to IRAP / GDR pictures can be partially restricted.

[0223] In one embodiment, when one or more slices in a picture have an NAL unit type such as IDR_N_LP, all other slices in the picture can have no NAL unit type such as CRA_NUT, IDR_W_RADL, or GDR_NUT. In this case, there may be no picture that follows the picture in the decoding order and precedes the picture in the output order.

[0224] In another embodiment, if one or more slices within a picture have a NAL unit type such as STSA_NUT, all other slices within the picture can have no NAL unit type such as CRA_NUT, IDR_W_RADL, IDR_N_LP or GDR_NUT.

[0225] In another embodiment, if one or more slices within a picture have a NAL unit type such as IDR_W_RADL, IDR_N_LP or CRA_NUT, all other slices within the picture can have a NAL unit type such as IDR_W_RADL, IDR_N_LP or CRA_NUT. For example, a hybrid NAL unit type based on CRA_NUT and IDR_W_RADL is acceptable.

[0226] On the other hand, all pictures belonging to an IRAP or GDR access unit (AU) can be restricted to have the same NAL unit type. For example, all pictures belonging to an IRAP access unit can have any NAL unit type among CRA_NUT, IDR_W_RADL and IDR_N_LP. Also, all pictures belonging to a GDR access unit can have a NAL unit type such as GDR_NUT.

[0227] (2) Hybrid NAL unit type 2

[0228] The hybrid of the NAL unit type for an IRAP picture and the NAL unit type for a leading picture can be partially restricted.

[0229] In one embodiment, if at least one sub-picture within a picture has a NAL unit type such as IDR_W_RADL or IDR_N_LP, all other sub-pictures within the picture can have no NAL unit type such as RASL_NUT.

[0230] In other embodiments, if at least one sub-picture in a picture has a NAL unit type such as IDR_N_LP, all other sub-pictures in said picture can have no NAL unit type such as RADL_NUT.

[0231] (3) Hybrid NAL unit type 3

[0232] Hybridization between the NAL unit type for an IRAP picture and the NAL unit type for a trailing picture is acceptable. For example, if at least one slice in a picture has a NAL unit type such as CRA_NUT, IDR_W_RADL or IDR_N_LP, at least one slice in said picture can have a NAL unit type such as TRAIL_NUT.

[0233] (4) Hybrid NAL unit type 4

[0234] Hybridization between NAL unit types for leading pictures is acceptable. For example, if at least one slice in a picture has a NAL unit type such as RADL_NUT (or RASL_NUT), all other slices in said picture can have a NAL unit type such as RASL_NUT (or RADL_NUT).

[0235] In one embodiment, a picture having a hybrid NAL unit type of RASL_NUT and RADL_NUT can be treated as a RASL picture. For an IRAP picture related to said picture, if NoIncorrectPicOutputFlag (or NoOutputBeforeRecorveryFlag) has a second value (e.g., 1), said picture is marked as not requiring output and the output process can be skipped.

[0236] FIG. 14 is a flowchart showing a method for determining the NAL unit type of a current picture according to an embodiment of the present disclosure.

[0237] Referring to FIG. 14, the image encoding apparatus can determine whether the current picture includes two or more sub-pictures (S1410). The segmentation information of the current picture can be signaled using one or more syntax elements within the upper level syntax. For example, via the picture parameter set (PPS) described above with reference to FIG. 11, a pps_no_pic_partition_flag indicating whether the current picture is segmented or not, and a pps_num_subpics_minus1 indicating the number of sub-pictures included in the current picture can be signaled. When the current picture includes two or more sub-pictures, the pps_no_pic_partition_flag can have a first value (e.g., 0), and the pps_num_subpics_minus can have a value greater than 0.

[0238] When the current picture does not include two or more sub-pictures (``NO'' in S1410), the image encoding apparatus can determine that the current picture has a single NAL unit type (S1440). That is, all slices within the current picture can have any one of the plurality of NAL unit types described above with reference to Table 1.

[0239] In contrast, when the current picture includes two or more sub-pictures (\"YES\" in S1410), the image coding device can determine whether each sub-picture in the current picture is to be treated as one picture (S1420). Information regarding whether a sub-picture is to be treated as one picture can be signaled using a predetermined syntax element within the upper-level syntax. For example, sps_subpic_treated_as_pic_flag[i] indicating whether a sub-picture is to be treated as one picture can be signaled via the sequence parameter set (SPS) described above with reference to FIG. 12. When each sub-picture in the current picture is to be treated as one picture, sps_subpic_treated_as_pic_flag[i] can have a second value (e.g., 1).

[0240] When each sub-picture in the current picture is not to be treated as one picture (\"NO\" in S1420), the image coding device can determine that the current picture has a single NAL unit type (S1440).

[0241] In contrast, when each sub-picture in the current picture is to be treated as one picture (\"YES\" in S1420), the image coding device can determine whether a predetermined hybrid constraint condition is satisfied for the current picture (S1430). The hybrid constraint condition can be determined based on the picture (or sub-picture) type as described above. For example, when at least one sub-picture in the current picture has an NAL unit type such as IDR_W_RADL or IDR_N_LP, all other sub-pictures in the current picture do not have to have an NAL unit type such as RASL_NUT. Alternatively, when at least one sub-picture in the current picture has an NAL unit type such as IDR_N_LP, all other sub-pictures in the current picture do not have to have an NAL unit type such as RADL_NUT.

[0242] When the hybrid constraint condition is satisfied for the current picture (\"YES\" in S1430), the image encoding device can determine that the current picture has a single NAL unit type (S1440).

[0243] In contrast, when the hybrid constraint condition is not satisfied for the current picture (NO in S1430), the image encoding device can determine that the current picture has a hybrid NAL unit type (S1450).

[0244] On the other hand, the image encoding device can generate a bitstream based on the encoding information of the current picture and transmit the generated bitstream to the image decoding device. When the current picture has a hybrid NAL unit type, the image encoding device can generate sub-bitstreams for each sub-picture in the current picture. At this time, a plurality of sub-bitstreams generated for the current picture can constitute one bitstream. In one embodiment, the bitstream composed of a plurality of sub-bitstreams is a single-layer bitstream, and in order to satisfy bitstream conformance, the following constraints can be applied.

[0245] -(Constraint 1) Each picture other than the first picture in decoding order within the bitstream is considered to be related to a previous IRAP picture in decoding order.

[0246] -(Constraint 2) When a picture is a leading picture of an IRAP picture, the picture must be a RADL or RASL picture.

[0247] -(Constraint 3) When a picture is a trailing picture of an IRAP picture, the picture must not be a RADL or RASL picture.

[0248] -(Constraint 4) There shall be no RASL pictures related to an IDR picture in the bitstream.

[0249] -(Constraint 5) There shall be no RADL pictures related to an IDR picture having a NAL unit type such as IDR_N_LP in the bitstream. In this case, if the essential parameter set to be referenced is available (either within the bitstream or via external means), by discarding all picture units (PUs) before the IRAP picture unit (PU), random access (and correct decoding of the IRAP picture and all subsequent non-RASL pictures) is possible at the position of the said IRAP picture unit (PU).

[0250] -(Constraint 6) All pictures preceding an IRAP picture in the decoding order shall precede the said IRAP picture in the output order and shall precede all RADL pictures related to the said IRAP picture in the output order.

[0251] -(Constraint 7) All RASL pictures related to a CRA picture shall precede all RADL pictures related to the said CRA picture in the output order.

[0252] -(Constraint 8) All RASL pictures related to a CRA picture shall follow all IRAP pictures preceding the said CRA picture in the output order.

[0253] -(Constraint 9) When the field_seq_flag has the first value (for example, 0) and the current picture is a leading picture associated with an IRAP picture, the current picture must precede all non-leading pictures associated with the IRAP picture in the decoding order. Or, for the first leading picture picA and the last leading picture picB associated with the IRAP picture, there must be one non-leading picture that precedes picA in the decoding order, and there must be no non-leading pictures between picA and picB in the decoding order.

[0254] On the other hand, as described above, information regarding whether the current picture has a mixed NAL unit type can be signaled using a predetermined syntax element within the upper-level syntax. For example, through the picture parameter set (PPS) described above with reference to FIG. 11, a pps_mixed_nalu_types_in_pic_flag indicating whether the current picture has a mixed NAL unit type can be signaled. In this case, the image decoding device can determine whether the current picture has a mixed NAL unit type based on the pps_mixed_nalu_types_in_pic_flag. For example, when the pps_mixed_nalu_types_in_pic_flag has the first value (for example, 0), the image decoding device can determine that the current picture has a single NAL unit type. In contrast, when the pps_mixed_nalu_types_in_pic_flag has the second value (for example, 1), the image decoding device can determine that the current picture has a mixed NAL unit type.

[0255] In one embodiment, when the current picture has a mixed NAL unit type, the type of the current picture can be determined based on the NUT value that the current picture has.

[0256] FIG. 15 is a flowchart showing a method for determining the type of a current picture having a hybrid NAL unit type according to an embodiment of the present disclosure.

[0257] Referring to FIG. 15, an image decoding apparatus can determine whether a current picture includes a slice having NAL unit types of CRA_NUT and IDR_W_RADL (first condition) (S1510).

[0258] When the first condition is satisfied (``YES'' in S1510), the type of the current picture can be determined as an IRAP picture (S1520).

[0259] When the first condition is not satisfied (``NO'' in S1510), the image decoding apparatus can determine whether the current picture includes a slice having NAL unit types of RASL_NUT and RADL_NUT (second condition) (S1530).

[0260] When the second condition is satisfied (``YES'' in S1530), the type of the current picture can be determined as a RASL picture (S1540). At this time, when NoIncorrectPicOutputFlag (or NoOutputBeforeRecorveryFlag) has a second value (for example, 1) with respect to the IRAP picture related to the current picture, the current picture can be marked as not requiring output.

[0261] When the second condition is not satisfied (``NO'' in S1530), the type of the current picture can be determined as a trailing picture (S1550).

[0262] On the other hand, in FIG. 15, step S1530 is shown to be performed after step S1510, but this can be variously modified according to the embodiment. For example, step S1530 may be performed simultaneously with step S1510, or may be performed before step S1510.

[0263] As described above, according to an embodiment of the present disclosure, when the current picture includes two or more sub-pictures and each of the sub-pictures is treated as one picture, the current picture can have a hybrid NAL unit type. Further, according to an embodiment of the present disclosure, various types of hybrid NAL unit types can be allowed based on a predetermined hybrid constraint condition. Thereby, in various applications and use cases, the hybrid NAL unit type can be applied more flexibly.

[0264] Hereinafter, with reference to FIGS. 16 and 17, an image encoding / decoding method according to an embodiment of the present disclosure will be described in detail.

[0265] FIG. 16 is a flowchart showing an image encoding method according to an embodiment of the present disclosure.

[0266] The image encoding method of FIG. 16 can be performed by the image encoding apparatus of FIG. 2. For example, step S1610 can be performed by the image division unit 110, and steps S1620 and S1630 can be performed by the entropy encoding unit 190.

[0267] Referring to FIG. 16, the image encoding apparatus can divide the current picture into one or more sub-pictures (S1610). The division information of the current picture can be signaled using one or more syntax elements in the upper-level syntax. For example, via the picture parameter set (PPS) described above with reference to FIG. 11, a pps_no_pic_partition_flag indicating whether the current picture is divided or not, and a pps_num_subpics_minus1 indicating the number of sub-pictures included in the current picture can be signaled. When the current picture is divided into two or more sub-pictures, the pps_no_pic_partition_flag can have a first value (for example, 0), and the pps_num_subpics_minus can have a value greater than 0.

[0268] Each sub-picture within the current picture can be treated as one picture. When a sub-picture is treated as one picture, the sub-picture can be encoded / decoded independently regardless of the encoding / decoding results of other sub-pictures. During encoding / decoding, information regarding the handling of the sub-picture can be signaled using syntax elements within the upper-level syntax. For example, via the sequence parameter set (SPS) described above with reference to FIG. 12, sps_subpic_treated_as_pic_flag[i] indicating whether each sub-picture within the current picture is treated as one picture can be signaled. If the i-th sub-picture within the current picture is treated as one picture during the encoding / decoding process excluding the in-loop filtering operation, sps_subpic_treated_as_pic_flag[i] can have a second value (e.g., 1). On the other hand, if sps_subpic_treated_as_pic_flag[i] is not signaled, it can be inferred that sps_subpic_treated_as_pic_flag[i] has the second value (e.g., 1).

[0269] At least some of the sub-pictures within the current picture can have different NAL unit types from each other. For example, if the current picture includes a first sub-picture and a second sub-picture, the first sub-picture can have a first NAL unit type, and the second sub-picture can have a second NAL unit type different from the first NAL unit type. An example of the NAL unit types that a sub-picture can have is as described above with reference to Table 1.

[0270] All slices included in each sub-picture within the current picture can have the same NAL unit type. In the above-described example, all slices included in the first sub-picture can have the first NAL unit type, and all slices included in the second sub-picture can have the second NAL unit type.

[0271] The image encoding device can determine the NAL unit type of each of a plurality of slices included in one or more sub-pictures of the current picture (S1620).

[0272] In one embodiment, when at least a part of a plurality of slices included in the current picture have different NAL unit types from each other, the current picture can include a first sub-picture and a second sub-picture (i.e., two or more sub-pictures) having different NAL unit types from each other. In this case, all slices included in each sub-picture within the current picture can have the same NAL unit type.

[0273] In one embodiment, the NAL unit type of the sub-picture can be determined based on the sub-picture type. For example, when the sub-picture is an IRAP sub-picture, the NAL unit type of the sub-picture can be determined as IDR_W_RADL, IDR_N_LP, or CRA_NUT. Alternatively, when the sub-picture is a RASL sub-picture, the NAL unit type of the sub-picture can be determined as RASL_NUT.

[0274] On the other hand, the NAL unit type that the second sub-picture within the current picture has can be determined based on the NAL unit type that the first sub-picture within the current picture has.

[0275] In one embodiment, when the first sub-picture in the current picture has a NAL unit type such as IDR_N_LP, the second sub-picture in the current picture can have a NAL unit type different from IDR_W_RADL and CRA_NUT. At this time, when the second sub-picture has a NAL unit type such as TRAIL_NUT, the second sub-picture can follow the first sub-picture in decoding order and output order.

[0276] In one embodiment, the plurality of slices included in the current picture can have NAL unit types different from GDR_NUT. For example, when the first sub-picture in the current picture has a NAL unit type such as STSA_NUT, the second sub-picture in the current picture can have a NAL unit type different from GDR_NUT.

[0277] In one embodiment, when the first sub-picture in the current picture has a NAL unit type such as IDR_W_RADL, IDR_N_LP or CRA_NUT, the second sub-picture in the current picture can have a NAL unit type different from RADL_NUT and RASL_NUT.

[0278] In one embodiment, the IRAP picture and the GDR picture can be restricted from having hybrid NAL unit types. Thereby, when at least some of the plurality of slices included in the current picture have different NAL unit types from each other (that is, when the current picture has a hybrid NAL unit type), the current picture can have a picture type different from the IRAP (intra random access point) picture and the GDR (gradual decoding refresh) picture.

[0279] The image encoding device can encode a plurality of slices included in the current picture based on the NAL unit type determined in step S1620 (S1630). At this time, as described above, the encoding process for each slice can be performed in units of CUs (coding units) based on a predetermined prediction mode. On the other hand, each sub-picture can be independently encoded and can constitute different sub-bitstreams. For example, a first sub-bitstream including the encoding information of the first sub-picture can be constituted, and a second sub-bitstream including the encoding information of the second sub-picture can be constituted.

[0280] As described above, according to an embodiment of the present disclosure, the current picture can have two or more NAL unit types based on the sub-picture structure. Also, the current picture can have various types of hybrid NAL unit types based on a predetermined hybrid constraint condition.

[0281] FIG. 17 is a flowchart showing an image decoding method according to an embodiment of the present disclosure.

[0282] The image decoding method of FIG. 17 can be performed by the image decoding device of FIG. 3. For example, steps S1710 and S1720 can be performed by the entropy decoding unit 210, and step S1730 can be performed by the inverse quantization unit 220 to the intra prediction unit 265.

[0283] Referring to FIG. 17, the image decoding device can obtain the VCL NAL unit type information of the current picture from the bitstream (S1710).

[0284] The VCL NAL unit type information of the current picture can include the NAL unit type value of the VCL NAL unit containing the encoded picture data (e.g., slice data) of the current picture. The NAL unit type value can be obtained by parsing the syntax element nal_unit_type included in the NAL unit header of the VCL NAL unit.

[0285] The image decoding device can determine the NAL unit type of each of a plurality of slices in the current picture based on the VCL NAL unit type information of the current picture (S1720).

[0286] In one embodiment, at least some of the plurality of slices in the current picture can have different NAL unit types from each other. In this case, the current picture can include a first sub-picture and a second sub-picture (i.e., two or more sub-pictures) having different NAL unit types from each other. All slices included in each sub-picture in the current picture can have the same NAL unit type. Also, the NAL unit type of the second sub-picture in the current picture can be determined based on the NAL unit type of the first sub-picture in the current picture.

[0287] In one embodiment, when the first sub-picture in the current picture has an NAL unit type such as IDR_N_LP, the second sub-picture in the current picture can have an NAL unit type different from IDR_W_RADL and CRA_NUT. At this time, when the second sub-picture has an NAL unit type such as TRAIL_NUT, the second sub-picture can follow the first sub-picture in decoding order and output order.

[0288] In one embodiment, the plurality of slices included in the current picture can have NAL unit types different from GDR_NUT. For example, if the first sub-picture in the current picture has an NAL unit type such as STSA_NUT, the second sub-picture in the current picture can have an NAL unit type different from GDR_NUT.

[0289] In one embodiment, if the first sub-picture in the current picture has an NAL unit type such as IDR_W_RADL, IDR_N_LP, or CRA_NUT, the second sub-picture in the current picture can have an NAL unit type different from RADL_NUT and RASL_NUT.

[0290] In one embodiment, IRAP pictures and GDR pictures can be restricted from having hybrid NAL unit types. Thereby, when at least some of the plurality of slices included in the current picture have different NAL unit types from each other (i.e., when the current picture has a hybrid NAL unit type), the current picture can have a picture type different from an IRAP (intra random access point) picture and a GDR (gradual decoding refresh) picture.

[0291] On the one hand, when the current picture has a hybrid NAL unit type, the type of the current picture can be determined based on the NUT value of the slice. In one embodiment, when the slices in the current picture have NAL unit types of CRA_NUT and IDR_W_RADL, the type of the current picture can be determined as an IRAP picture. Alternatively, when the slices in the current picture have NAL unit types of RASL_NUT and RADL_NUT, the type of the current picture can be determined as a RASL picture. At this time, if NoIncorrectPicOutputFlag (or NoOutputBeforeRecoveryFlag) has a second value (for example, 1) for the IRAP picture related to the current picture, the current picture is marked as not requiring output and the output process can be skipped. Except for the cases described above, the type of the current picture can be determined as a trailing picture.

[0292] The image decoding device can decode each slice in the current picture based on the NAL unit type determined in step S1720 (S1730). At this time, the decoding process for each slice can be performed in units of CUs (coding units) based on a predetermined prediction mode, as described above.

[0293] As described above, according to an embodiment of the present disclosure, the current picture can have two or more NAL unit types based on the sub-picture structure. Also, the current picture can have various types of hybrid NAL unit types based on a given hybrid constraint condition.

[0294] The names of the syntax elements described in this disclosure can include information regarding the position where the syntax element is signaled. For example, a syntax element starting with "sps_" can be meant to be signaled in a sequence parameter set (SPS). Further, syntax elements starting with "pps_", "ph_", "sh_", etc. can be meant to be signaled in a picture parameter set (PPS), a picture header, a slice header, etc., respectively.

[0295] The exemplary methods of this disclosure are presented in a series of operations for clarity of explanation, but this is not for restricting the order in which the steps are performed. If necessary, each step can also be performed simultaneously or in a different order. To implement the method according to this disclosure, it can also include further other steps in the exemplified steps, or include the remaining steps excluding some steps, or include additional other steps excluding some steps.

[0296] In this disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform an operation (step) of checking the execution conditions and situations of the operation (step). For example, when it is described that a predetermined operation is performed if a predetermined condition is satisfied, the image encoding device or the image decoding device can perform the operation of checking whether the predetermined condition is satisfied and then perform the predetermined operation.

[0297] The various embodiments of this disclosure do not list all possible combinations, but are for explaining representative aspects of this disclosure. The matters described in the various embodiments can be applied independently or in combinations of two or more.

[0298] In addition, various embodiments of the present disclosure can be implemented by hardware, firmware, software, or a combination thereof. In the case of implementation by hardware, it can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general processors, controllers, microcontrollers, microprocessors, etc.

[0299] In addition, the image decoding device and the image encoding device to which the embodiments of the present disclosure are applied can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an over-the-top video (OTT) device, an Internet streaming service providing device, a three-dimensional (3D) video device, an image phone video device, and a medical video device, etc., and can be used for processing video signals or data signals. For example, the over-the-top video (OTT) device can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.

[0300] FIG. 18 is a diagram illustrating a content streaming system to which an embodiment of the present disclosure can be applied.

[0301] As shown in FIG. 18, the content streaming system to which the embodiment of the present disclosure is applied can generally include an encoding server, a streaming server, a Web server, a media storage, a user device, and a multimedia input device.

[0302] The encoding server compresses the content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data to generate a bitstream and transmits this to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a video camera directly generates a bitstream, the encoding server can be omitted.

[0303] The bitstream can be generated by the image encoding method and / or the image encoding device to which the embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0304] The streaming server transmits multimedia data to the user device based on the user's request via the Web server, and the Web server can serve as a medium to inform the user of what services are available. When the user requests a desired service from the Web server, the Web server transmits this to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system can include a separate control server. In this case, the control server can play a role in controlling commands / responses between each device in the content streaming system.

[0305] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0306] Examples of the user device may include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device, for example, a smartwatch, smart glass, an HMD (head mounted display), a digital TV, a desktop computer, a digital signage, and the like.

[0307] Each server in the content streaming system can be operated as a distributed server, and in this case, the data received from each server can be distributedly processed.

[0308] The scope of the present disclosure includes software or machine-executable commands (for example, an operating system, an application, firmware, a program, etc.) that enable the operations according to the methods of various embodiments to be executed on a device or a computer, and a non-transitory computer-readable medium on which such software or commands are stored and can be executed on the device or the computer.

Industrial Applicability

[0309] Examples according to the present disclosure are available for encoding / decoding images.

Claims

1. An image decoding method performed by an image decoding device, comprising: The image decoding method includes: Obtaining video coding layer (VCL) network abstraction layer (NAL) unit type information of a current picture from a bitstream; determining a NAL unit type for each of a plurality of slices included in the current picture based on the obtained VCL NAL unit type information; and decoding the plurality of slices based on the determined NAL unit type. Based on at least some of the slices having different NAL unit types, the current picture includes a first sub-picture and a second sub-picture having different NAL unit types; a NAL unit type of the second sub-picture is determined based on a NAL unit type of the first sub-picture; A method for decoding an image, wherein all slices included in a sub-picture in the current picture have the same NAL unit type.

2. The image decoding method of claim 1 , wherein the second sub-picture has a NAL unit type other than IDR_W_RADL and CRA_NUT based on the NAL unit type of the first sub-picture being IDR_N_LP.

3. The image decoding method according to claim 2 , wherein the second sub-picture follows the first sub-picture in decoding order and output order.

4. The image decoding method of claim 1 , wherein the slices have a NAL unit type other than GDR_NUT.

5. The image decoding method of claim 1 , wherein the second sub-picture has a NAL unit type other than RADL_NUT and RASL_NUT based on the NAL unit type of the first sub-picture being IDR_W_RADL, IDR_N_LP, or CRA_NUT.

6. The image decoding method of claim 1 , wherein the current picture has a picture type other than an intra random access point (IRAP) picture and a gradual decoding refresh (GDR) picture.

7. An image coding method performed by an image coding device, comprising: The image encoding method includes: dividing a current picture into one or more sub-pictures; determining a NAL unit type for each of a plurality of slices included in the one or more sub-pictures; and encoding the plurality of slices based on the determined NAL unit type; the current picture includes a first sub-picture and a second sub-picture having different NAL unit types based on at least some of the slices having different NAL unit types; a NAL unit type of the second sub-picture is determined based on a NAL unit type of the first sub-picture; An image coding method, wherein all slices included in a sub-picture in the current picture have the same NAL unit type.

8. The image coding method of claim 7 , wherein the second sub-picture has a NAL unit type other than IDR_W_RADL and CRA_NUT based on the NAL unit type of the first sub-picture being IDR_N_LP.

9. 9. The image coding method of claim 8, wherein the second sub-picture follows the first sub-picture in decoding order and output order.

10. The image coding method of claim 7 , wherein the slices have a NAL unit type other than GDR_NUT.

11. The image coding method of claim 7 , wherein the second sub-picture has a NAL unit type other than RADL_NUT and RASL_NUT based on the NAL unit type of the first sub-picture being IDR_W_RADL, IDR_N_LP, or CRA_NUT.

12. The image encoding method of claim 7 , wherein the current picture has a picture type other than an intra random access point (IRAP) picture and a gradual decoding refresh (GDR) picture.

13. A method for transmitting a bitstream generated by an image coding method, comprising the steps of: The image encoding method includes: dividing a current picture into one or more sub-pictures; determining a network abstraction layer (NAL) unit type for each of a plurality of slices included in the one or more sub-pictures; encoding the plurality of slices based on the determined NAL unit type; the current picture includes a first sub-picture and a second sub-picture having different NAL unit types based on at least some of the slices having different NAL unit types; a NAL unit type of the second sub-picture is determined based on a NAL unit type of the first sub-picture; A method for transmitting a bitstream, in which all slices contained in a subpicture in the current picture have the same NAL unit type.